DORA metrics — deployment frequency, lead time for changes, change failure rate, and mean time to recovery — have become the standard vocabulary for measuring software delivery performance. The four numbers were derived from years of DORA's research about what separates high-performing engineering organizations from low-performing ones.
The metrics are useful. They are also frequently misimplemented and even more frequently misinterpreted. This post is how to measure them honestly and what the numbers actually tell you.
The Four Metrics
Deployment frequency. How often you deploy to production. Elite teams deploy multiple times per day; low performers deploy once per month or less.
Lead time for changes. The time from commit to production. Elite teams have lead times under an hour; low performers measure in weeks or months.
Change failure rate. The percentage of deploys that result in degraded service requiring remediation. Elite teams stay under 5%; low performers exceed 60%.
Mean time to recovery (MTTR). How long it takes to restore service after a production incident. Elite teams recover in under an hour; low performers take more than a week.
Together they describe both speed (deploy frequency, lead time) and stability (failure rate, MTTR). The original DORA research found that high-performing teams excel at all four simultaneously — speed and stability are not actually a tradeoff.
How to Calculate Each One Honestly
The mistakes start in the calculations.
Deployment Frequency
Definition: how often deploys to production happen.
deploy_count_per_day = count of production deploys / number of days
Sources:
- CI/CD system's deployment records
- Production deployment audit log
- Argo CD or Flux sync events
Common mistakes:
- Counting deploys to staging, not production
- Counting CI runs, not actual deploys
- Excluding rollbacks (rollbacks are deploys too)
For most teams, the right number comes from "production deploy" events emitted by the deployment pipeline. Make sure exactly one event per deploy, regardless of how many services are deployed.
Lead Time for Changes
Definition: time from commit to production.
lead_time = production_deploy_time - commit_time
For each commit that reaches production, measure the time. Aggregate as median or 75th percentile (mean is misleading because of long tails).
Sources:
- Git commit timestamps
- Production deploy events
- Need to join these via commit SHA
Common mistakes:
- Measuring "time from merge to production" instead of "commit to production." Commits sit in branches; that time counts.
- Including only the latest commit in a deploy. A deploy contains many commits; each has its own lead time.
- Using means instead of medians. Long-tail commits skew the mean.
Change Failure Rate
Definition: percentage of deploys that caused a problem requiring remediation.
change_failure_rate = (deploys causing incidents + rollbacks + hotfixes) / total deploys
Sources:
- Production deploy events
- Incident tracker (PagerDuty, Opsgenie, ServiceNow)
- Rollback events from the deployment pipeline
Common mistakes:
- Counting only declared incidents. Many incidents are quietly hotfixed without being reported.
- Counting only severity-1 incidents. Any deploy that required a fix is a failure for this metric.
- Excluding deploys that revealed problems hours later.
The most pragmatic approach: any deploy that required a hotfix or rollback within 24 hours counts as a failure. The 24-hour window is somewhat arbitrary but consistent.
Mean Time to Recovery
Definition: time from the start of a production incident to its resolution.
mttr = sum(resolution_time - detection_time) / count_of_incidents
Sources:
- Incident tracker timestamps
- Monitoring detection times
- Resolution timestamps
Common mistakes:
- Using the time the incident was created, not when symptoms started. Real impact began before someone created a ticket.
- Excluding incidents that resolved themselves. Auto-recovered incidents are still failure events.
- Cherry-picking which incidents count.
The honest version measures detection-to-resolution. If you cannot measure detection accurately, measure ticket-create-to-resolution and acknowledge the limitation.
Where to Pull the Data
For a non-vendor implementation, the typical data flow:
- Deployment events from the CI/CD platform. GitHub Actions, GitLab CI, ArgoCD, your platform emits structured events.
- Commit metadata from git. Available in the deployment event or fetchable from the API.
- Incident records from the incident tracker.
- Rollback records from the deployment pipeline.
Collect these into a database. A small DBT model can produce the four metrics on a schedule.
-- Simplified deployment_frequency view
SELECT
DATE_TRUNC('day', deployed_at) AS day,
COUNT(*) AS deploys
FROM deployments
WHERE environment = 'production'
GROUP BY 1;
The implementation is not the hard part; the discipline of consistent event emission is.
What the Numbers Tell You
The four metrics in combination paint a picture.
High frequency, low failure rate, short MTTR: elite. The system is well-engineered, tests catch issues, deploys are routine, recovery is fast.
High frequency, high failure rate: moving fast and breaking things. Speed without stability. Adopt better testing, smaller changes, or canary deploys before continuing.
Low frequency, high failure rate, long MTTR: legacy. Big infrequent deploys, more risk per deploy, slower recovery. The DORA improvement path is to deploy smaller changes more often.
Low frequency, low failure rate: cautious. Maybe too cautious — long lead time means slow response to opportunities and bugs. Adopt CI/CD and canaries to deploy more often safely.
Low MTTR, high failure rate: flaky systems with good operators. The team has good recovery practices but the system breaks too easily. Invest in stability before scaling.
What the Metrics Do Not Measure
DORA metrics measure delivery performance. They do not measure:
- Whether the right things are being built
- User happiness
- Code quality
- Engineering culture
- Architectural health
A team with elite DORA metrics can still ship the wrong features. A team with low DORA metrics can ship the right features (slowly). Treat them as one input among several, not the whole picture.
How Often to Look
DORA metrics are most useful on a weekly or monthly cadence. Daily readings are noisy; trend lines over a quarter are meaningful.
The four numbers should be one tab in the engineering dashboard, not a constant focus. If the team is staring at the metrics every day, you have introduced the wrong incentives.
Goal-Setting
Setting DORA metrics as targets ("we will deploy 10 times per day next quarter") creates Goodhart's Law problems. The metrics measure outcomes; targeting them directly produces gaming.
The right way to use them: as indicators that improvement work is paying off. "We invested in CI speed; lead time should drop." "We adopted canary deploys; failure rate should drop." The improvement is the goal; the metrics are how you check progress.
Tooling
Three options:
- Build your own. Pull data into a warehouse, query with SQL, visualize in Grafana. Most control, most setup.
- Open source. Four Keys (Google's reference implementation), Apache DevLake, OpenTelemetry's metrics standards. Lower setup cost.
- Commercial. Cortex, LinearB, Sleuth, Faros. Polished UIs, integrations out of the box, real money.
For a team just starting: build your own with a small SQL model. The metric definitions matter more than the visualization. Once you have a working baseline, evaluate whether a commercial tool adds enough value.
A Practical Implementation
For a Laravel + Kubernetes setup:
- Emit a deploy event from the deploy pipeline with
{commit, environment, started_at, finished_at}. - Collect incidents from PagerDuty with their start and resolution timestamps.
- Track rollbacks as a special deploy type.
- Write four SQL queries — one per metric — that aggregate weekly.
- Render in Grafana or a dashboard tool.
The total work is a week or two. The maintenance is minimal once the events flow reliably.
What Matters More Than the Numbers
The DORA metrics framework is most useful as a vocabulary. "Our lead time is two weeks; how do we cut it?" is a more productive conversation than "our development velocity feels slow."
The metrics give the team a shared language for talking about delivery. The discipline of measuring is more important than the specific numbers, and the conversations about improvement are more important than the dashboard itself.
Building visibility into your delivery performance and not sure where to start? We help teams set up honest DORA measurement and the improvement work that actually moves the numbers. scopeforged.com