Skip to content

Measuring SRE Outcomes

FDAI measures SRE improvement through outcomes and guard metrics, not through the percentage of decisions made automatically. Every comparison uses the same scenario set, a stated measurement window, and paired baseline and treatment evidence.

MetricWhat it answers
MTTR distributionHow long resolution takes at the mean, median, and p90
Auto-resolution rateWhich events reached the correct terminal outcome with no human touchpoint and no later rollback
Human touchpointsHow much operator work remains per incident
Change lead timeHow long a governed change takes from request to merge
Cost per resolved eventWhat attributable platform and inference spend each result consumed

Track change-failure rate, false positives, false negatives, rollback rate, policy-violation escapes, audit gaps, verifier failures, and mixed-model disagreement. An outcome improvement does not count if a guard metric regresses past its threshold.

  1. Freeze the scenario-set version and input distribution.
  2. Record the baseline model, rules, thresholds, adapters, and catalog versions.
  3. Run treatment against the same scenarios and observation window.
  4. Report sample size, missing data, confidence, and distribution, not only an average.
  5. Keep observation mode and enforce outcomes separate.
  6. Demote a capability when measured guard metrics regress.
  • Do not claim a multiplier without paired measurements.
  • Do not treat a missing projection as zero.
  • Do not merge mean and p90 into one latency statement.
  • Do not count a later rollback as successful auto-resolution.
  • Do not compare different scenario sets without labeling the change.
To learn aboutRead
Canonical formulas and windowsGoals and metrics
The named scenario sets and evidence levelsScenario validation inventory
How SLO burn measures workload impactSLOs and error budgets
How observation mode evidence controls promotionObserve, then enable changes
How audit evidence is reconstructedRead the audit log