Cert Notes/ Commute Study Notes
Roadmap
KOEN
CLF-C02 · FoundationalCloud Practitioner - Foundational
DVA-C02 · AssociateDeveloper - Associate
SAA-C03 · AssociateSolutions Architect - Associate
SOA-C02 · AssociateCloudOps Engineer - Associate
SAP-C02 · ProfessionalSolutions Architect - Professional
DOP-C02 · ProfessionalDevOps Engineer - Professional
SCS-C03 · SpecialtySecurity - Specialty
MLA-C01 · AssociateMachine Learning Engineer - Associate
AIF-C01 · FoundationalAI Practitioner - Foundational
DEA-C01 · AssociateData Engineer - Associate
  • Week 1
    • 1.What Is Data Engineering
    • 2.Batch vs Streaming
    • 3.A Bird's-Eye View of AWS Data Services
    • 4.Data Formats and Modeling
    • 5.Week 1 Comprehensive Review
  • Week 2
    • 1.Batch Ingestion: S3 Upload, DataSync, Transfer Family, Snow
    • 2.Kinesis Data Streams: Shards, Partition Keys, and Throughput
    • 3.Kinesis Data Firehose: Delivery Streams and Loading
    • 4.Amazon MSK (Kafka): Topics, Partitions, and When to Use What
    • 5.Week 2 Synthesis: Data Ingestion Part 1 Review
  • Week 3
    • 1.Streaming Processing: Managed Service for Apache Flink and Window Aggregation
    • 2.Ingestion Reliability: Idempotency, Ordering, Retries, Deduplication, DLQ
    • 3.CDC and Data Replication: Database Migration Service (DMS)
    • 4.Ingestion Architecture Patterns: Lambda Architecture and Event-Driven Ingestion
    • 5.Week 3 Synthesis: Data Ingestion Part 2 Review
  • Week 4
    • 1.Glue Data Catalog and Crawlers: Adding Metadata to Data
    • 2.Glue ETL Job: Transform Data on Spark
    • 3.Glue Studio and DataBrew: Transform Without Code
    • 4.Schema Management and Data Quality: Tolerate Evolution, Guarantee Trust
    • 5.Week 4 Synthesis: The Big Picture of AWS Glue Transformation
  • Week 5
    • 1.Amazon EMR: Spark, Hive and Cluster Operations, Plus EMR Serverless
    • 2.Lambda Transformation and Lightweight Processing: Event-Driven ETL's Limits and Fit
    • 3.Orchestration: Step Functions, MWAA, and Glue Workflows Selection Criteria
    • 4.Performance and Cost Optimization: File Format, Compression, Partitioning, and Small File Problem
    • 5.Week 5 Synthesis: Data Transformation 2 — Engines, Orchestration, Optimization Integration Review
  • Week 6
    • 1.S3 Data Lake Layout and Partitioning Strategy
    • 2.AWS Lake Formation Central Permission Management
    • 3.Open Table Formats: Iceberg, Hudi, Delta Lake
    • 4.S3 Storage Management and Cost Optimization
    • 5.Week 6 Comprehensive Review: Data Lake Recap
  • Week 7
    • 1.Amazon Redshift: Distribution and Sort Keys and Workload Optimization
    • 2.Amazon Athena: Serverless Queries and Cost Optimization
    • 3.DynamoDB (Analytics Perspective): Key Design and Stream-Based Pipelines
    • 4.RDS/Aurora and Store Selection: OLTP, Zero-ETL, Workload-to-Store Decision
    • 5.Week 7 Comprehensive Review: Analytics Stores Recap
  • Week 8
    • 1.Pipeline Monitoring: CloudWatch Metrics, Logs, and Alarms
    • 2.Data Quality and Validation: Glue Data Quality and Quality Gates
    • 3.Logging, Audit, and Troubleshooting: CloudTrail and Failure Recovery
    • 4.Cost and Performance Operations: Monitoring, Sizing, Auto Scaling
    • 5.Week 8 Comprehensive Review: Data Operations and Support Recap
  • Week 9
    • 1.Access Control: IAM and Lake Formation Permissions
    • 2.Encryption: KMS and Service-Specific Encryption
    • 3.Sensitive Data Protection: Macie and Masking
    • 4.Data Governance: Catalog, Lineage, Sharing, Auditing
    • 5.Week 9 Synthesis: Security and Governance Review
  • Week 10
    • 1.Integrated Review of Domains 1 & 2: Ingestion, Transformation & Storage Management
    • 2.Integrated Review of Domains 3 & 4: Operations & Support, Security & Governance
    • 3.Full-Length Practice Exam Pace: 8 Integrated Scenarios
    • 4.Common Traps & Keywords: "Requirement → Service" Translation Guide
    • 5.Final D-Day Prep: Exam Structure, Time Management & Scenario Breakdown Strategy
MLS-C01 · SpecialtyMachine Learning - Specialty
← DEA-C01/Week 1/Day 5
DEA-C01· AssociateWeek 1 · Day 5~16 min read

Day 5 - Week 1 Comprehensive Review

Over the past four days we laid the foundation of data engineering. Day 1 covered the role and the pipeline, OLTP/OLAP, and lake/warehouse; Day 2 split batch from streaming; Day 3 unfolded the map of AWS data services; and Day 4 dealt with data formats and schema evolution. They may look like scattered pieces, but in fact they thread together into a single pipeline.

Today, instead of re-listing individual concepts, we follow one scenario from beginning to end to see how the four days of knowledge interlock. The exam doesn't test fragmented memorization — it asks "what would you choose in this situation" — so reviewing through these connections is the most effective approach.

Week 1 on One Page

[Day1] Roles & concepts   [Day2] Processing  [Day3] Services      [Day4] Formats
 Data engineer             Batch              Ingest Kinesis/DMS   CSV/JSON (raw)
 = pipeline owner          vs                 Store S3/Redshift    Parquet/ORC (analytics)
 OLTP vs OLAP              Streaming          Process Glue/EMR     Avro (streaming)
 Lake vs Warehouse         (latency/          Analyze Athena/QS    Schema evolution
                            throughput)       Governance LakeFmt

The keywords of the four days are ultimately different facets of one question: "how do you ingest, store, process, analyze, and govern data."

💡 Related theory: The single principle running through all these judgments is "the workload determines the design." OLTP or OLAP, batch or stream, row-based or columnar — the answers all come from the workload characteristic of "how is the data read and written." Data engineering's way of thinking is not to pick the technology first and force the workload to fit, but to analyze the workload and then choose the technology.

Threading It Together with a Scenario

Let's follow the requirements of an e-commerce company. Each requirement is tagged with which Week 1 concept it touches.

Requirement 1. "Make the order data analyzable without putting load on the production order DB."

  • The production DB is OLTP (Day 1) → heavy aggregation must not run directly on it
  • Replicate only the changes to the analytics environment with DMS CDC (Day 3)
  • The analytics environment is OLAP (Day 1) → Redshift or S3+Athena

Requirement 2. "I want to catch fraudulent transactions the moment a payment happens."

  • Immediacy required → streaming (Day 2)
  • Ingest payment events with Kinesis Data Streams (Day 3)
  • The schema may change, so Avro + Schema Registry (Day 4)

Requirement 3. "Yesterday's revenue report, automatically, every morning."

  • Periodic, bounded data → batch (Day 2)
  • Process a day's worth with Glue (Day 3), load as Parquet (Day 4)
  • Query and visualize with Athena/QuickSight (Day 3)
-- The analytics query for Requirement 3 — fast and cheap because it's stored as Parquet
SELECT region, SUM(amount) AS revenue
FROM orders          -- Parquet in S3, registered in the Glue Catalog
WHERE order_date = DATE '2026-06-25'   -- partition pruning
GROUP BY region;

A single company can have all three requirements at once, each using different paradigms, services, and formats. The key insight is that there is no "one right answer" — there is a fitting combination for each requirement.

💡 Related theory: Merge Requirements 1–3 into one picture and you naturally get a lakehouse. All raw data (order CDC, payment streams, logs) gathers in the S3 data lake, gets refined into Parquet through batch and streaming processing, and Athena, Redshift, and QuickSight analyze on top of it. This is why Day 1's "lake vs warehouse" gets "unified into the lakehouse" in the modern era.

Sorting Out the Commonly Confused Pairs

These are the pairs that show up as exam traps.

Confusing PairKey Distinction
OLTP vs OLAPTransactions (few rows, fast) vs analytics (massive rows, aggregation)
Lake vs Warehouseschema-on-read (flexible) vs schema-on-write (strict)
Batch vs StreamingBounded, periodic vs unbounded, continuous
Kinesis Streams vs FirehoseLow latency, multiple consumers vs auto-loading delivery service
Glue vs EMRManaged serverless ETL vs control, large-scale clusters
Row-based vs ColumnarWrites, whole rows (Avro) vs reads, aggregation (Parquet)
Athena vs RedshiftAd-hoc, serverless, pay-per-scan vs always-on, large-scale, loaded
IAM vs Lake FormationResource level vs data (table/column/row) level

Getting the distinction criteria for just these eight pairs right lets you solve a good share of the scenario questions in Week 1's scope.

💡 Related theory: Nearly every pair in this table shares the same tension: "flexibility and generality vs efficiency and specialization." The lake is flexible but needs management; the warehouse is efficient but rigid. Serverless (Glue/Athena) is convenient but weaker on fine-grained control; clusters (EMR/Redshift) are powerful but carry operational burden. Once you are conscious of the trade-off, memorization turns into understanding.

The Exam Perspective in One Line

DEA-C01 ultimately asks about these four domains (reconfirming Day 1).

1. Ingestion & Transformation (34%)  ← Day2 batch/stream + Day3 Kinesis/Glue/EMR
2. Store Management (26%)            ← Day1 lake/warehouse + Day3 S3/Redshift + Day4 formats
3. Operations & Support (22%)        ← (Week 2 onward) orchestration, monitoring
4. Security & Governance (18%)       ← Day3 LakeFormation/IAM/KMS

Week 1 laid nearly all the groundwork for domains 1 and 2. From next week we go into the deep details of each service (partitioning strategies, Glue job tuning, Kinesis shards, and so on), but remember that all of it rests on the big picture we drew today.

Wrapping Up

Week 1's conclusion is simple. The data engineer is the person who builds pipelines, and that pipeline flows through ingestion → storage → processing → analytics, wrapped by governance. Every design judgment (OLTP/OLAP, batch/stream, lake/warehouse, row-based/columnar) ultimately comes down to "what does the workload demand."

If you have this way of thinking in hand, Week 1 is a success. From Week 2 we put flesh on this skeleton and get into actually working with each service.

📝 Practice Questions

Click a choice to reveal the answer and explanation.

Question 1

What is the single design principle that runs through all of Week 1, determining every choice of OLTP/OLAP, batch/stream, and row-based/columnar?

Question 2

Which combination best satisfies the requirement "I want to run large-scale aggregation analysis on the production order DB's data without putting load on it"?

Question 3

Which combination of processing paradigm, ingestion service, and format best fits the requirement "detect fraudulent transactions the moment a payment happens"?

Question 4

Which statement most accurately describes the key difference between Kinesis Data Streams and Kinesis Data Firehose?

Question 5

What is the modern standard pattern that combines the strengths of the data lake and the data warehouse, layering a table format and catalog on top of an S3 data lake to provide warehouse-grade queries?

PreviousData Formats and ModelingWeek 1 · Day 4Next Batch Ingestion: S3 Upload, DataSync, Transfer Family, SnowWeek 2 · Day 1

On this page

  • Week 1 on One Page
  • Threading It Together with a Scenario
  • Sorting Out the Commonly Confused Pairs
  • The Exam Perspective in One Line
  • Wrapping Up
  • Practice Questions