datadog3 min read

Curated summary

How we built a real-world evaluation platform for autonomous SRE agents at scale

Read original(opens in new tab)

Bits AI SRE improved in isolated scenarios but lacked a way to detect regressions across the broader range of production incidents. The team found that tool-level tests and live replays could not capture failures caused by multi-step reasoning or changing telemetry. They built a replayable evaluation platform combining realistic investigation labels, scalable orchestration, and longitudinal performance tracking.

Subtle Regressions from Well-Intentioned Features

  • Adding a monitor’s service name to the agent’s initial context improved some internal investigations.
  • Across broader scenarios, it introduced irrelevant signals that confused the agent and degraded unrelated investigations.
  • Because there was no representative evaluation set, the team could not measure the change’s wider impact before internal misses exposed it.
  • The incident demonstrated the need to evaluate every change across diverse investigation types.

Limits of Tool Tests and Live Replay

  • Testing tools individually failed to capture errors caused by incorrect interactions between valid tool outputs.
  • Live investigation replay was difficult to scale because:
    • Results were not consistently aggregated.
    • Production environments changed.
    • Telemetry expired, making investigations unreplayable.
  • Standard evaluation frameworks assumed clean inputs and static datasets, unlike agents operating over production telemetry.
  • The team needed controlled, offline replay of realistic end-to-end investigations.

Evaluation Labels and World Snapshots

  • Each label represents one production-style investigation.
  • It contains:
    • Ground truth: the issue’s actual root cause.
    • World snapshot: the queries and signals available when the issue occurred.
  • The agent is never shown the root cause directly; it must reason from the preserved signals.
  • Labels must cover varied technologies and failure modes, including:
    • Kubernetes pod failures
    • Kafka lag
    • Bad-code deployments
    • Complex multi-service business failures
  • A narrow or overly clean dataset would make performance appear better than it really is.

Orchestrating Evaluations at Scale

  • The platform runs Bits against labels, scores the outcomes, and tracks quality over time.
  • It supports comparisons across:
    • Investigation categories
    • Model variants
    • Configuration versions
    • Evaluation runs
  • The architecture consists of a shared label set, an orchestration layer, and reporting infrastructure.
  • This allows teams to determine whether improvements in one domain, such as Kafka, regress another, such as Kubernetes.

Scaling Label Creation

  • The team initially created labels manually from Datadog alerts.
  • Manual labeling provided early coverage but consumed engineering time and remained far from representative.
  • They embedded label generation into Bits itself:
    • Customer feedback and investigation data are used to derive root causes.
    • Relevant queries are preserved as the world snapshot.
    • Each user interaction becomes a potential evaluation case.
  • This increased label creation rates by an order of magnitude and allowed coverage to grow with product usage.

Agent-Assisted Validation

  • Early labels required extensive human review, especially when feedback was ambiguous or reconstructed signals were uncertain.
  • As ingestion grew, manual review became a bottleneck.
  • Bits was then used to assist with validation by aggregating related signals, identifying relationships, and resolving ambiguous feedback before human review.

Practical Conclusion

Reliable agent improvement requires more than testing individual tools or replaying live incidents. A representative, production-derived label set combined with reproducible end-to-end evaluations makes regressions visible and enables safer iteration.

Continue with another curated summary.