Cert Notes/ Commute Study Notes
Roadmap
KOEN
CLF-C02 · FoundationalCloud Practitioner - Foundational
DVA-C02 · AssociateDeveloper - Associate
SAA-C03 · AssociateSolutions Architect - Associate
SOA-C02 · AssociateCloudOps Engineer - Associate
SAP-C02 · ProfessionalSolutions Architect - Professional
DOP-C02 · ProfessionalDevOps Engineer - Professional
  • Week 1
    • 1.DevOps as an Operating Model: The Five Axes of CALMS and the Truth DORA Proved Through Measurement
    • 2.Well-Architected Framework: Rereading It Through the Six Lenses of DevOps
    • 3.The AWS DevOps Tool Map: The Code* Series and the Real Picture Beyond It
    • 4.Multi-Account Strategy: The Real Picture of Organizations, Control Tower, and IAM Identity Center
    • 5.Week 1 Wrap-Up: Cementing the DevOps Thinking Frame with Scenarios
  • Week 2
    • 1.CodeCommit Deep Dive: What Changes When Git Hosting is Integrated with IAM
    • 2.GitHub Actions ↔ AWS OIDC: Eliminating Static Keys Permanently Through Federation
    • 3.CodeArtifact and Supply Chain Security: When Dependencies Become Attack Surface
    • 4.DevSecOps Shift Left: Automated Security Gates Built with Code Signing, CodeGuru, and Inspector
    • 5.Week 2 Synthesis: DevSecOps Thinking Framework From Source Control to Code Signing
  • Week 3
    • 1.The Real Meaning of buildspec.yml: The Moment a Pipeline Specification Becomes Code
    • 2.The Physics of Build Speed: Trade-offs Among Cache, Parallelism, and Compute
    • 3.Design Principles of Secret Management: The Criteria That Separate Secrets Manager from Parameter Store
    • 4.VPC CodeBuild, Custom Images, ARM/Graviton: Expanding the Boundary of the Build Environment
    • 5.Week 3 Review: Integrated CodeBuild Scenarios and Practical Judgment
  • Week 4
    • 1.In-place vs Blue/Green, AppSpec: The Physics of Deployment Strategies
    • 2.EC2/On-Prem Deployment + Auto Scaling Integration: The Intersection of Instance Lifecycle and Deployment
    • 3.Lambda Deployment: Linear/Canary/AllAtOnce and the Math of Aliases
    • 4.ECS Blue/Green + CodeDeploy Traffic Shift: The Logic of Two Target Groups
    • 5.Week 4 Review: Integrated Scenarios in CodeDeploy Deployment Strategies
  • Week 5
    • 1.CodePipeline Architecture: Understanding Why Stage, Action, and Artifact Were Designed This Way
    • 2.Multi-Account Pipeline: Why Cross-Account IAM Is Required
    • 3.Action Providers: How Lambda, Step Functions, and Manual Approval Extend the Pipeline
    • 4.Dynamic Pipeline: V2 Variable System, Trigger Filters, and Execution Mode Design
    • 5.Week 5 Review: CodePipeline Integration Scenarios
  • Week 6
    • 1.ECR: Solving the Problems that Container Image Registries Must Solve
    • 2.ECS Automatic Deployment: From Task Definition Update to Auto Scaling
    • 3.EKS CI/CD: Why GitOps Was Born, and How ArgoCD and Flux Changed Kubernetes
    • 4.App Runner and ECS Copilot: The Spectrum of Container Operations Abstraction
    • 5.Week 6 Review + 12 Scenario Practice Problems
  • Week 7
    • 1.Serverless CI/CD: Lambda, SAM, and Canary Deployments
    • 2.Lambda Permissions, Layers, and Container Images
    • 3.API Gateway and Serverless Integrations
    • 4.X-Ray and CloudWatch for Serverless Observability
    • 5.Week 7 Review: Serverless CI/CD Summary
  • Week 8
    • 1.CloudFormation Advanced: Nested·Cross-Stack and the Deep Story of Modularization
    • 2.StackSets: The Deep Story of Deploying IaC to Thousands of Accounts
    • 3.Custom Resource·Hooks·Change Set: CloudFormation's Extension and Validation Mechanisms
    • 4.CDK·CDK Pipelines·Terraform: Deep Comparison of Modern IaC Tools and Self-Evolving Pipelines
    • 5.Week 8 Integrated Scenario: IaC Tools Meeting Within One Incident
  • Week 9
    • 1.Systems Manager: Run Command·Session Manager·Patch Manager's Deep Story
    • 2.State Manager·Inventory·Compliance: The Abstraction of Desired State
    • 3.AppConfig: Feature Flags and Progressive Deployment
    • 4.Secrets Manager and Parameter Store: Lifecycle and Rotation
    • 5.Week 9 Integrated Scenario: Configuration Management Across Fleet Scale
  • Week 10
    • 1.CloudWatch Metrics: Time Series, Dimensions, and Alarm Evaluation Model Deep Dive
    • 2.CloudWatch Logs: Groups, Streams, Subscriptions, and Insights Deep Dive
    • 3.Container Insights·Lambda Insights·EMF: Workload-Specific Observability Deep Dive
    • 4.Synthetics·RUM·Evidently: Three Lenses Measuring User Experience
    • 5.Week 10 Synthesis: Tying Observability into Incidents
  • Week 11
    • 1.X-Ray: Causal Graphs of Distributed Tracing and the Deep Story of the Trace Model
    • 2.X-Ray Sampling: Reservoir Algorithm and the Economics of Tracing at Operational Scale
    • 3.ADOT: The Deep Story of OpenTelemetry Ending the Tracing Tool Wars
    • 4.OpenSearch · AMP · AMG: The Two Worlds of Inverted Indices and Time-Series Databases
    • 5.Week 11 Synthesis: Real-World Decision-Making in Tracing and Telemetry Observability
  • Week 12
    • 1.EventBridge: Event Bus Routing Model and the Nervous System of Asynchronous Automation
    • 2.Systems Manager Automation: Codifying Runbooks and the Operator's Disappearance
    • 3.Auto-Healing and Control Theory: When Systems Repair Themselves
    • 4.ChatOps and Incident Manager: The Coordination Layer
    • 5.Week 12 Synthesis: The Five-Stage Pipeline and Decision Trees
  • Week 13
    • 1.Multi-AZ High Availability: Distributed Principles Underlying Replication Consistency, Quorum, and Failover
    • 2.Multi-Region Resilience: Distributed Principles of DNS Routing, Global Replication, and Encryption Boundaries
    • 3.Four DR Strategies: Tradeoffs of RTO, RPO, Cost and Their Economics
    • 4.Validating Resilience: Chaos Engineering with Resilience Hub and FIS
    • 5.Week 13 Comprehensive Review: Tying Together High Availability, Multi-Region, DR, and Resilience Validation
  • Week 14
    • 1.GuardDuty and Automatic Isolation: Signal Processing, Statistics, and Auto-Response Principles
    • 2.Security Hub: SIEM Principles of Normalization, Aggregation, and Auto-Remediation
    • 3.AWS Config: Closed-Loop Control Principles for State Recording, Drift Detection, and Automated Remediation
    • 4.Audit Manager, Macie, Inspector: Evidence Automation, Data Classification, Vulnerability Scanning Principles
    • 5.Week 14 Comprehensive Review: The Big Picture of Security Automation Stack and Practical Scenarios
  • Week 15
    • 1.Multi-Account Enterprise CI/CD: Governance and Platform Engineering Principles for 50+ Accounts
    • 2.Hybrid CI/CD: Bridging On-Premises and AWS into One Deployment Model
    • 3.Large-Scale ECS/EKS Operations: Scheduling, GitOps, Cost Principles for 100+ Microservices
    • 4.Serverless Large-Scale Incident Auto-Response: Recovery Without People, Safety Rails
    • 5.Week 15 Synthesis: Reading Signals and Trade-Off Judgment
  • Week 16
    • 1.Domain 1+2 Integrated Review: SDLC Automation and IaC as One Thread
    • 2.Domain 3+4 Integrated: Resilience and Observability as Failure Prevention
    • 3.Incident Response and Security Compliance Woven Through Everything
    • 4.Full Exam Scenarios (Domains 1-6 Integrated)
    • 5.D-Day Exam Prep: Mental State, Time Management, Last-Minute Do's and Don'ts
SCS-C03 · SpecialtySecurity - Specialty
MLA-C01 · AssociateMachine Learning Engineer - Associate
AIF-C01 · FoundationalAI Practitioner - Foundational
DEA-C01 · AssociateData Engineer - Associate
MLS-C01 · SpecialtyMachine Learning - Specialty
← DOP-C02/Week 1/Day 1
DOP-C02· ProWeek 1 · Day 1~50 min read

Day 1 - DevOps as an Operating Model: The Five Axes of CALMS and the Truth DORA Proved Through Measurement

When people first hear the word DevOps, most understand it as "tools like Jenkins and Ansible." If you start there and stay at that level all the way to the exam, you'll freeze in front of Professional-level scenario questions. What DOP-C02 asks is not the names of tools, but "why this combination, why this order, why measure with this metric." The thinking framework that produces those answers is exactly CALMS and the DORA 4 metrics.

This article covers why DevOps emerged, why it became an operating model rather than a mere collection of tools, and how the DORA research that made this operating model measurable interlocks with AWS tool selection. If you can picture which automation should have kicked in when the alarm went off at 3 AM, the exam scenarios solve themselves naturally.

The Birth of DevOps: The Moment the Silos Had to Break

The first DevOpsDays was held on June 23, 2009, in Ghent, Belgium. The word "DevOps," coined there by Patrick Debois, didn't appear out of nowhere. The previous year, at the 2008 Agile conference in Toronto, Andrew Shafer had opened a BoF session on "Agile Infrastructure," but hardly anyone showed up. Debois came alone, and their conversation raised the question "why is Dev doing Agile while Ops can't keep up" — which led to DevOpsDays a year later.

The industry context of that moment matters. In 2007-2008, John Allspaw and Paul Hammond of Flickr dropped a bombshell with their talk "10+ Deploys Per Day" (Velocity 2009). At the time, enterprises considered one deployment per quarter normal, yet Flickr deployed 10 times a day while remaining stable. How was that possible? The answer was "it's possible when Dev and Ops work like one team, use the same tools, and look at the same metrics." That's the starting point of the DevOps movement.

💡 Related theory: Conway's Law (1968) — "A system's architecture reflects the communication structure of the organization that built it." In other words, when Dev and Ops teams are separated, the code and operational infrastructure ossify in separated forms, and the "throw it over the wall" anti-pattern naturally emerges between them. DevOps applies the Inverse Conway Maneuver — "redesign the organization first to match the desired architecture." Team Topologies (Skelton & Pais, 2019) formalized this idea into four team types (Stream-aligned, Platform, Enabling, Complicated-subsystem).

📚 Case study: In 2010, when John Allspaw joined Etsy as CTO, he established the "blameless postmortem" culture. This is not mere kindness — it's a core SRE principle. In a culture that blames people, engineers hide incidents, system flaws stay hidden, and the same incidents repeat. Etsy shared a "Postmortem of the Week" every week and built a learning-organization culture, which soon became a standard industry SRE practice. AWS's Correction of Errors (COE) process borrows exactly this pattern.

CALMS: DevOps Maturity Decomposed Into Five Axes

The CALMS framework, organized by Jez Humble and Damon Edwards around 2010, decomposes what "having achieved DevOps" means into five measurable axes.

LetterMeaningKey questionAWS mapping
C CultureCollaboration, shared responsibility, blameless culture"Who gets scolded when an incident happens?"Cross-account IAM, ChatOps (Chatbot + Slack), AWS Incident Manager
A AutomationEliminate manual procedures, Pipeline-as-Code"How many steps do humans still do by hand?"CodePipeline, CodeBuild, CDK, SSM Automation Runbook
L LeanSmall batches, WIP limits, waste elimination"How many days does a single PR stay open?"Trunk-based dev, CodeCommit + Feature flags (AppConfig)
M MeasurementMeasure everything, data-driven decisions"How do you know this change is good?"CloudWatch Metrics, Container Insights, DORA dashboards
S SharingShare knowledge, tools, and failures"Do one team's insights flow to other teams?"AWS Service Catalog, Internal Developer Platform (IDP), Wiki/Confluence

These five axes are not independent. Lean is impossible without Automation (with too much manual work you can't make batches small), and Culture doesn't change without Measurement (without data you end up arguing based on "gut feeling" alone). So when a DOP-C02 scenario asks "where is this company stuck," it's usually a situation where a deficit in one axis has dragged the others down like dominoes.

🔍 Going deeper: CALMS's "Lean" is a direct descendant of the Toyota Production System (TPS). The two pillars of TPS — Jidoka (automation with a human touch) and Just-in-Time — translate directly into DevOps's "automation + small batch." In particular, TPS's "Andon cord" concept (workers have the authority to stop the line) carries over into DevOps's principle that "anyone can stop a deployment." This got formalized as Google's SRE Error Budget, producing mechanisms like "if the team-defined SLO is broken, deployments are automatically frozen."

💡 Related theory: WIP (Work In Progress) limits are the core of Kanban. According to Little's Law (average processing time = WIP / Throughput), as WIP grows, the processing time (lead time) of any single work item increases linearly. So a team with 10 PRs open simultaneously has a longer lead time than a team that merges PRs one at a time. What trunk-based development pursues is the arithmetic consequence of this Little's Law.

[ DevOps Maturity Surface — the 5 CALMS axes ]

       Culture
         /\
        /  \
   Sharing  Automation
       \    /
        \  /
   Measurement — Lean

The shortest edge of each axis is the organization's true DevOps maturity
(the weakest axis is the ceiling for the whole).
This is why "doing Automation well alone" is not DevOps.

DORA 4 Metrics: The Decisive Blow of Measurable DevOps

Starting in 2013, the DORA (DevOps Research and Assessment) research led by Nicole Forsgren, Jez Humble, and Gene Kim surveyed over 32,000 engineers over six years, statistically proving "how DevOps connects to business outcomes." The results were compiled into the 2018 book "Accelerate" and the annually published "State of DevOps Report." And the measurement indicators were condensed into just four.

MetricDefinitionElite thresholdHighMediumLow
Deployment FrequencyFrequency of production deploymentsOn demand (multiple times per day)Once per day ~ once per weekOnce per week ~ once per monthLess than once per month
Lead Time for ChangesTime from commit → prodUnder 1 hourUnder 1 day1 day ~ 1 week1 week ~ 1 month
Change Failure Rate (CFR)Percentage of deployments causing incidents0-15%16-30%16-30%16-30%
MTTR (Time to Restore)Time to recover from an incidentUnder 1 hourUnder 1 day1 day ~ 1 weekOver 1 week

The biggest statistical finding of DORA is that speed (Deployment Frequency, Lead Time) and stability (CFR, MTTR) are not a trade-off but positively correlated. Teams that deploy frequently also have fewer incidents and recover faster. It's counter-intuitive, but deploying small changes frequently means ① the scope of each change is small, making debugging easier, ② automation and rollback are well established, making recovery fast, and ③ deployment is routine, so the sense of risk doesn't dull.

📚 Case study: Amazon disclosed in 2011 that it deployed on average once every 11.6 seconds, i.e., about 2.7 million deployments per year (Jon Jenkins, Velocity 2011). This was possible because of the combination of microservices + 2-Pizza Teams (8 people or fewer) + fully automated pipelines. At the same time, the average enterprise deployed once per quarter. That gap became the justification for the DevOps movement.

🔍 Going deeper: Starting with the 2021 DORA report, a fifth metric — Reliability (SLO achievement rate) — was added. Initially there was pushback: "isn't MTTR similar to SLO?" But MTTR only looks at "recovery after an incident," while Reliability looks at "the state of no incidents occurring." They are different dimensions. This connects to the "SLO-based automatic rollback" scenarios frequently seen in DOP-C02 (e.g., CloudWatch Alarm → CodeDeploy auto-rollback).

⚠️ Pitfall: If you force DORA metrics as KPIs, gaming will inevitably occur. Anti-patterns like "splitting meaningless commits into small units to inflate Deployment Frequency" or "not calling an incident an incident to make MTTR look good." DORA themselves explicitly state that "metrics are diagnostic tools, not evaluation tools." The moment a company ties DORA to performance reviews, that data becomes a lie.

Mapping DORA → AWS Tools — What the Exam Really Asks

DOP-C02 scenarios come in the form "a company is experiencing problem X. What is the most suitable AWS solution?" Translate that X into a DORA metric and the answer becomes visible.

Scenario 1: "Deployments happen once a quarter. We want to increase frequency" → Improve Deployment Frequency

  • Build a trunk-based automated pipeline with CodePipeline
  • Separate deployment from release with feature flags (AppConfig) (dark launch)
  • Consider Lambda-level microservice decomposition to enable smaller batches

Scenario 2: "It takes 2 weeks from commit to prod" → Improve Lead Time

  • Remove manual approval gates (minimize CodePipeline Manual Approval)
  • Shorten build time with CodeBuild caching (S3/Local)
  • Remove friction between environments with multi-account Cross-Account IAM

Scenario 3: "Incidents occur in 30% of deployments" → Improve Change Failure Rate

  • Canary deployments (Lambda alias + traffic shifting, CodeDeploy Blue/Green)
  • Automatic rollback based on CloudWatch Alarms
  • Enforce tests with pre-deploy hooks (CodeDeploy AppSpec hooks)

Scenario 4: "Recovery takes days after an incident" → Improve MTTR

  • Automatic remediation via EventBridge → Lambda → SSM Automation
  • Unify alerts and runbooks with AWS Incident Manager + Chatbot (Slack)
  • Quickly identify root cause with X-Ray + Container Insights

With this mapping in your head, the "obviously this one" answer becomes visible on the exam, and at the same time you can distinguish among the 2-3 options that all technically work "which one is most precisely what the question is asking."

🎯 Scenario: A fintech company reported "one deployment per month, a 30% chance of an outage per deployment, and average recovery of 2 days." Its DORA rating is Low on every axis. Where do you start? The answer is Automation first. Why? Without automation, small batches are impossible (no Lean), fast rollback is impossible (MTTR won't shrink), and data collection is hard (no Measurement). When CALMS's A axis collapses, the other four axes all get dragged down. On AWS, the trio of CodePipeline + CodeBuild + CodeDeploy is the starting point.

Werner Vogels's "You Build It, You Run It" — Redefining the Boundary of Responsibility

This single line, delivered by Werner Vogels (AWS CTO) in a 2006 ACM Queue interview, is the compressed form of AWS's organizational operating model. It means the development team takes on operational responsibility for its own code — and it's not a mere slogan but a structural enforcement mechanism. Amazon's Two-Pizza Team (a single team of 8 or fewer owns a service's entire lifecycle) is the embodiment of this principle, and AWS's microservice architecture is its product.

Why is this a decisive element of DevOps? When operational responsibility is separated, developers don't know "whether my code triggers alarms at 3 AM," and operators don't know "who wrote this code and why." When neither knows, incidents repeat and the system stays perpetually broken. You Build It, You Run It eliminates this information asymmetry.

AWS tools are precisely aligned with this philosophy. CodePipeline lets developers define their own pipelines directly, CloudWatch lets them put their own service's metrics on their own dashboards, and X-Ray lets them view their own code's traces in their own console. The model where "a platform team manages the infrastructure and the dev team just writes code" is not AWS's default design assumption. When "establish a dedicated operations team" appears as an exam option, it's almost always a trap.

💡 Related theory: Combine Conway's Law with You Build It, You Run It and you get a service-oriented organization. When Amazon broke apart its monolith in the late 1990s, it enforced a 1:1 mapping of "one service = one team, one team = one service," and as a result, microservice architecture emerged naturally. In other words, AWS's microservices didn't appear "because a distributed-systems book said they were good" but as a "byproduct of organizational structure."

📚 Case study: After a 2008 data center failure prevented Netflix from shipping DVDs for 3 days, Netflix migrated to AWS and simultaneously created Chaos Monkey (2010). This later evolved into the field of Chaos Engineering and is the direct ancestor of AWS Fault Injection Simulator (FIS). The idea of "deliberately breaking things before an incident happens" is an extreme form of You Build It, You Run It — only a team that bears operational responsibility can resolve to cause incidents in advance.

SRE and DevOps — Different Facets of the Same Thing

SRE (Site Reliability Engineering), created at Google in 2003 under Ben Treynor Sloss, is a twin that grew up elsewhere at nearly the same time as DevOps. From the outside they look similar, but their internal definitions differ.

DimensionDevOpsSRE
OriginIndustry movement (2009 Ghent)Single company (Google, 2003)
Core principlesCALMSError Budget, SLO/SLI, toil elimination
MeasurementDORA 4 metricsSLO achievement rate, Error Budget burn rate
Organizational model"Dev + Ops unified""Separate SRE team alongside Dev teams (50% coding, 50% operations)"
Responsibility boundaryYou build it, you run itSLO-based responsibility sharing
Definition of automationPipeline-as-CodeToil < 50% enforced

The key insight is Ben Treynor's one-liner: "Class SRE implements DevOps." That is, DevOps is the abstract interface (principles), and SRE is its concrete implementation. The two movements don't conflict, and the AWS tool ecosystem supports both paradigms.

The DOP-C02 exam doesn't explicitly distinguish the two, but when keywords like "Error Budget," "SLO-based automatic rollback," or "toil elimination" appear, the question has a stronger SRE flavor; keywords like "CI/CD pipeline design" or "cross-account automation" have a stronger DevOps flavor.

🔍 Going deeper: SRE's Error Budget is intuitive. "With a monthly SLO of 99.9%, you have a downtime budget of 43.2 minutes per month. If you spend the whole budget, no new feature deployments that month — only stability work." What happens when this is automated? Detect SLO violations with CloudWatch Alarms → EventBridge → a Lambda disables the CodePipeline deployment stage → notify Slack. This pattern appears frequently on the exam in AWS contexts.

💡 Related theory: An SLO (Service Level Objective) is different from an SLA (an external contract). An SLA is a legal contract (usually 99.9%); an SLO is an internal target (usually 99.95%, set slightly higher than the SLA). The 0.05% buffer between them is the "operational safety margin." And an SLI (Indicator) is the actual metric that measures the SLO — for example, "5xx response ratio < 0.1%." On AWS, you compute SLIs with CloudWatch Metric Math + Composite Alarms to monitor SLOs.

Pipeline-as-Code and GitOps — The Second Evolution of Automation

When the Automation axis of CALMS first took hold, GUI pipelines based on the Jenkins UI were the standard. But that was a problem — because the pipeline itself wasn't code, it couldn't be version-controlled, couldn't be reviewed, and one accidental click in the GUI by one person was game over. What broke through this limitation was the Jenkinsfile (2016), and its extension is GitOps (2017, Weaveworks).

The definition of Pipeline-as-Code is simple: the pipeline definition itself goes into a Git repository as code. CodePipeline may look like a GUI tool when created in the console, but internally it exports as a JSON definition and can be codified with CDK Pipelines or Terraform. GitHub Actions' .github/workflows/*.yml and GitLab CI's .gitlab-ci.yml are the same idea.

GitOps goes one step further: the desired state of the operating environment also lives in Git as the Single Source of Truth (SSOT). Who changed what in prod is all recorded as git commits, and when drift occurs it is automatically reconciled to the desired state. ArgoCD and Flux in the Kubernetes ecosystem are the representative tools, and on AWS the EKS + ArgoCD combination is the most common.

💡 Related theory: GitOps grew naturally out of Kubernetes's declarative API + control loop paradigm. K8s controllers continuously compare "desired state" and "current state" and reconcile them; put the desired state in Git and a git push becomes a deployment. Pull-based GitOps (ArgoCD) has the cluster polling git, so it has a clearer security boundary than push-based (cluster credentials are not exposed externally).

🔍 Going deeper: The AWS Code* series is a push-based model. CodePipeline watches Git/ECR/S3 and pushes when it detects changes. Using ArgoCD with EKS switches you to pull-based. The two models have different security trade-offs — push requires the central pipeline to hold credentials for every cluster, while pull only requires each cluster to have read access to Git. In multi-cluster, multi-account environments, pull-based almost always wins.

The Four Axes of the AWS DevOps Tool Map

AWS's DevOps tools are not just the Code* series. Classify them along the following four axes and it becomes clear on the exam what must interlock with what.

[ AWS DevOps tools — 4 axes ]

  [Source/Build]           [Deploy]
   CodeCommit               CodeDeploy
   CodeArtifact             CodePipeline (orchestration)
   CodeBuild                Elastic Beanstalk
   GitHub/GitLab(OIDC)      AppRunner
        |                       |
        +-----[ IaC ]-----------+
        |   CloudFormation      |
        |   CDK / SAM           |
        |   Terraform           |
        +-----[ Operate ]-------+
            CloudWatch
            X-Ray / ADOT
            SSM (Automation, Patch, AppConfig)
            EventBridge / Chatbot / Incident Manager
            Config / GuardDuty / Security Hub

These four axes map exactly onto "Domains 1-6." Source/Build/Deploy is Domain 1 (SDLC automation), IaC is Domain 2 (configuration management), monitoring within Operate is Domain 4, incidents are Domain 5, security is Domain 6, and Domain 3 (resilience) spans all four axes.

The Starting Point of DORA Measurement, Seen Through the CLI

Theory alone is abstract, so let's see where you actually pull data from to measure DORA's "Deployment Frequency."

# Count SUCCEEDED executions in CodePipeline's execution history — last 30 days
aws codepipeline list-pipeline-executions \
  --pipeline-name prod-pipeline \
  --max-items 1000 \
  --query "pipelineExecutionSummaries[?status=='Succeeded' && startTime>=\`$(date -u -d '30 days ago' +%Y-%m-%dT%H:%M:%SZ)\`]" \
  | jq 'length'
 
# CodeDeploy deployment history — needed for measuring Lead Time
aws deploy list-deployments \
  --application-name prod-app \
  --include-only-statuses Succeeded \
  --create-time-range start=2026-04-01,end=2026-04-30 \
  | jq '.deployments | length'
 
# Record directly as a CloudWatch metric (custom)
aws cloudwatch put-metric-data \
  --namespace "DevOps/DORA" \
  --metric-name DeploymentFrequency \
  --value 1 \
  --dimensions Pipeline=prod,Env=prod

One insight here: AWS does not provide a "dashboard that measures DORA directly." DORA is a measurement framework, not a tool, so companies must pull data from their own sources (CodePipeline, GitHub, Jira, PagerDuty) and synthesize it themselves. This is the correct answer when the exam asks "How would you measure DevOps maturity?" — build your own dashboard with CloudWatch Metrics + Custom Metrics + QuickSight or Grafana.

📚 Case study: The Four Keys project open-sourced by Google in 2018 (github.com/GoogleCloudPlatform/fourkeys) is a reference implementation that collects GitHub/Jira/PagerDuty data into BigQuery and automatically computes DORA metrics. In an AWS environment, you build the same pattern as an EventBridge → Firehose → S3 → Athena → QuickSight pipeline. The exam doesn't ask about it directly, but the components of such pipelines frequently appear as scenario options.

Wrapping Up — The Same Picture for the Exam and Real Work

Etch today's three pictures into memory. First, DevOps is an operating model, not tools. If any one of CALMS's five axes is weak, the rest get dragged down, and exam scenarios almost always ask "which axis has collapsed and how do you recover it." Second, the DORA 4 metrics are measurable DevOps. Speed and stability are not a trade-off but move together, and the AWS Code* + CloudWatch combination is the foundation of that measurement. Third, AWS's tool philosophy is "You Build It, You Run It" — consistently designed around the model where the dev team is responsible for operating its own code.

In the next article, we reinterpret the Well-Architected Framework from a DevOps perspective. We'll confirm together that the 6 Pillars are not a simple checklist but a thinking framework where "classify any domain's problem under a Pillar and the answer becomes visible."

📝 Practice Questions

Click a choice to reveal the answer and explanation.

Question 1

A company reported the following state: deployment cycle of once per month, average commit-to-prod time of 14 days, incident rate of 28% per deployment, and average incident recovery time of 36 hours. How would you classify this under DORA ratings?

Question 2

Which of the following is furthest from being a consequence of Werner Vogels's "You Build It, You Run It" principle reflected in AWS tool design?

Question 3

A team's Deployment Frequency is at Elite level with 5 deployments per day, but its Change Failure Rate is very high at 45%. What is the highest-priority improvement?

Question 4

A company wants to adopt a GitOps model in an EKS environment. Between push-based (CodePipeline deploys to the cluster) and pull-based (ArgoCD polls Git), which model is more advantageous in a multi-account, multi-cluster environment, and why?

Question 5

Among CALMS's five axes, which other axes are directly affected when "Measurement" is not fulfilled?

Question 6

What is an appropriate pattern for implementing SRE's Error Budget concept on AWS?

Question 7

As an organization accelerates from "one deployment per quarter → one deployment per day," which AWS tool combination should be laid down first?

Question 8

Which is the most accurate description of the essential effect of a Blameless Postmortem?

Next Well-Architected Framework: Rereading It Through the Six Lenses of DevOpsWeek 1 · Day 2

On this page

  • The Birth of DevOps: The Moment the Silos Had to Break
  • CALMS: DevOps Maturity Decomposed Into Five Axes
  • DORA 4 Metrics: The Decisive Blow of Measurable DevOps
  • Mapping DORA → AWS Tools — What the Exam Really Asks
  • Werner Vogels's "You Build It, You Run It" — Redefining the Boundary of Responsibility
  • SRE and DevOps — Different Facets of the Same Thing
  • Pipeline-as-Code and GitOps — The Second Evolution of Automation
  • The Four Axes of the AWS DevOps Tool Map
  • The Starting Point of DORA Measurement, Seen Through the CLI
  • Wrapping Up — The Same Picture for the Exam and Real Work
  • Practice Questions