New AnnouncementOpenSRE’s SRE Agent is now Open Source

The architecture of
reliability for CI/CD

When a pipeline run fails, OpenSRE gathers context, investigates, and opens a fix for you to review.

Layer 001

Failure intake

OpenSRE picks up failed workflow runs from GitHub Actions and normalises each one into a single investigation state.

Normalise failures into one state.

Layer 002

Context assembly

OpenSRE enriches the failure with:

  • The triggering commit and diff
  • Recent workflow and config changes
  • Run history and baselines for the workflow

Attach the diff, changes, baselines.

Layer 003

Execute actions

OpenSRE decides whether the failure is CI-caused, code-caused, or unresolved, and identifies plausible causes and investigation objectives.

Rank causes and hypotheses.

Layer 004

Agentic investigation loop

  • Plans next checks (logs, re-runs, diffs)
  • Executes against your repo and CI tooling
  • Synthesises evidence into evolving hypotheses

Check, re-run, update confidence.

Layer 005

Confidence gate

OpenSRE continuously evaluates:

  • Hypothesis confidence
  • Whether the fix is verified
  • Marginal value of further investigation

Stops and asks you when it isn't sure, or when more checks are unlikely to change the outcome.

Fix when sure. Ask when not.

Layer 006

Actionable fix report

OpenSRE produces:

  • Likely root cause(s)
  • Supporting evidence
  • Recommended next actions

Delivered to GitHub, Slack, or internal systems.

Evidence-backed fix, ready to review.