The architecture of
reliability for CI/CD
When a pipeline run fails, OpenSRE gathers context, investigates, and opens a fix for you to review.
Failure intake
OpenSRE picks up failed workflow runs from GitHub Actions and normalises each one into a single investigation state.
Normalise failures into one state.
Context assembly
OpenSRE enriches the failure with:
- The triggering commit and diff
- Recent workflow and config changes
- Run history and baselines for the workflow
Attach the diff, changes, baselines.
Execute actions
OpenSRE decides whether the failure is CI-caused, code-caused, or unresolved, and identifies plausible causes and investigation objectives.
Rank causes and hypotheses.
Agentic investigation loop
- Plans next checks (logs, re-runs, diffs)
- Executes against your repo and CI tooling
- Synthesises evidence into evolving hypotheses
Check, re-run, update confidence.
Confidence gate
OpenSRE continuously evaluates:
- Hypothesis confidence
- Whether the fix is verified
- Marginal value of further investigation
Stops and asks you when it isn't sure, or when more checks are unlikely to change the outcome.
Fix when sure. Ask when not.
Actionable fix report
OpenSRE produces:
- Likely root cause(s)
- Supporting evidence
- Recommended next actions
Delivered to GitHub, Slack, or internal systems.
Evidence-backed fix, ready to review.
