← All posts

What our experimental self-healing software loop shipped in two weeks

Thomas Daly·
Tech

We ran a pilot test between September 18 and October 1 of a self-healing loop pieced together with three separate services, all within our own repository. This loop opened 152 pull requests, none of which was written by a human. Our engineers merged 80 of them.

Three products, five handoffs, one human decision

Each of the loop's three layers does the one job it is good at. Fullstory is the production monitoring layer: it watches real sessions and flags when something is wrong. Replay QA does two separate jobs at opposite ends of the loop. Obvious writes the fix. Then a person reviews the PR and the QA report and decides whether to merge. That is the only human step, by design.

The loop: Fullstory (production signal) feeds Replay QA (root cause), which feeds Obvious (ships the fix), then round again.
The loop: Fullstory (production signal) feeds Replay QA (root cause), which feeds Obvious (ships the fix), then round again.
  1. Detect (Fullstory): surfaces the signal from real production sessions.
  2. Diagnose (Replay QA): reproduces it, finds the root cause from the recording, and files a GitHub issue with severity, evidence and a bug ID.
  3. Fix (Obvious): opens a PR on fix/qa-issue-N with the why, what changed and the tradeoffs, linked by Fixes #N.
  4. Verify (Replay QA): explores the PR preview for regressions, then reproduces the original bug against it: replaying the original session for frontend changes, reproducing it independently for backend changes.
  5. Merge (Engineer): the one human decision. The issue closes itself.

Reproducing a bug from a recording and proving a fix killed that bug use one mechanism twice. The second use is where this experiment got interesting.

152 opened, 80 merged, 80 verified

Fourteen days, one repository, no hand-written code. Here is where every one of the 152 pull requests ended up.

Where all 152 PRs went: 46 merged and verified by session replay, 34 merged and verified by independent reproduction, 45 still open, 27 closed without merging.
Where all 152 PRs went: 46 merged and verified by session replay, 34 merged and verified by independent reproduction, 45 still open, 27 closed without merging.

Every one of the 80 merged PRs landed with the original bug reproduced against the PR and shown to be gone. Forty-six were frontend changes, verified by replaying the session that found the bug. The other 34 touched the backend, where Replay QA reproduced the problem independently.

Underneath the PRs are 124 distinct bugs. Seventy-six are now fixed on main, and every one of those fixes was verified. That is a little over five bugs fixed a day with nobody writing the code. Eighteen of the 124 needed more than one attempt, because when a PR was closed or failed QA, the loop opened another one. One Segment PII leak took four tries.

Pace: 5.4 bugs fixed per day, 3-day median from PR to merge, 99 PRs opened in the first three days, 18 bugs needing a retry.
Pace: 5.4 bugs fixed per day, 3-day median from PR to merge, 99 PRs opened in the first three days, 18 bugs needing a retry.

Most of the 27 closed-without-merge PRs are those superseded attempts, not outright rejections. After the first three days cleared the backlog, merges kept pace with new PRs. One note on scope: seven of the 152 PRs have no linked QA bug. They are follow-ups, mostly the loop repairing CI on one of its own earlier PRs.

What a loop like this is good at

The 124 bugs, grouped by type. Two things in this chart changed how we think about this.

Bugs filed and fixed, by type. Every fix that landed was verified.
Bugs filed and fixed, by type. Every fix that landed was verified.

Performance is 43% of everything the loop fixed. That covers oversized payloads, API responses served no-cache and request waterfalls. One response was shipping about 975KB of journey-authoring trees on every project view. Bugs like these sit in a backlog for a year because no single person owns them and none of them is ever the most urgent thing. A loop does not care that they are boring.

Security bugs had the crispest reproductions. A security bug tends to fail in an unambiguous way, so proving the fix is unambiguous too. The fixes included GET /api/projects returning the private project list with a 200 and no authentication; unauthenticated project reads returning the full record, including owner_id; and /api/auth-me, which carries email and Auth0 id, being served stale from a CDN cache about 132 seconds later despite no-store.

Share of bugs fixed by severity: high 9 of 22, medium 39 of 62, low 28 of 40. High-severity bugs more often need a judgment call about intent, which is where we kept the human.
Share of bugs fixed by severity: high 9 of 22, medium 39 of 62, low 28 of 40. High-severity bugs more often need a judgment call about intent, which is where we kept the human.

Two ways to prove a fix

Verification is diagnosis pointed at the fix: reproduce the original bug against the PR and show it is gone. How Replay QA reproduces it depends on what the PR changed.

How the 80 merged fixes were verified: 46 by session replay (frontend changes), 34 by independent reproduction (backend changes).
How the 80 merged fixes were verified: 46 by session replay (frontend changes), 34 by independent reproduction (backend changes).

Either way, the engineer reviewing the PR was not asked to trust a model's account of its own fix. Each merge came with the original failure reproduced against the change and shown to be gone.

Generation was never the hard part. The loop produced 152 PRs in two weeks. What made 80 of them mergeable was proof, and that is the difference between a loop that generates pull requests and one engineers will actually merge.

Somebody already named the missing piece

While writing this up, we looked for anyone who had described the stage that made this work. The closest we found is a sentence from Chris Kelly, Head of Developer Experience at Augment, in his guide to building a software factory for LeadDev on 10 September:

A verifier deploys to an isolated instance, exercises the affected behavior, and posts inspectable proof — logs, screenshots, a replayable trace.
— Chris Kelly, Augment · LeadDev, 10 Sep 2026

We would like to make the case that the third item in that list is doing far more work than its position suggests.

A factory already has logs and screenshots. It does not have a replayable trace, and that is the artefact that separates a verifier that reports an outcome from one that hands over something a human, or a second agent, can interrogate. We read the published descriptions of agentic software factories, including Cloudflare's, Vercel's, Builder.io's, Augment's, OpenAI's and Uber's. Most have no reproduce-or-diagnose stage at all. Of the handful that do, Kelly's is the only one we have found that says what the stage should produce.

As far as we can tell, he is also describing something nobody ships today. Vercel's documentation is admirably direct about the same wall from the other side: investigation can reproduce a failure before implementation starts, but for bugs that need a browser, every browser session runs with a person present, and an unattended run falls back on recorded findings.

The only hard numbers we know of come from Google's automated-program-repair team, who measured both halves of this (arXiv:2502.01821).

Google APR: +30% more bugs fixed with a reproduction test; 28% success reconstructing one from the bug report. This loop: 80 of 80 merged fixes landed with the original bug reproduced and shown to be gone.
Google APR: +30% more bugs fixed with a reproduction test; 28% success reconstructing one from the bug report. This loop: 80 of 80 merged fixes landed with the original bug reproduced and shown to be gone.

Reproduction was the highest-leverage input they measured, and reconstructing one was the step that failed. In follow-up work, they note that their own engineers, using the deployed system, asked for the reproduction test to be included in the AI-generated patch, because that is what made them willing to trust it.

Our two weeks are a small, concrete version of the same finding. For the 46 frontend fixes, nothing had to be reconstructed: the session that found the bug was recorded as it happened. For the 34 backend fixes, the loop did the harder thing Google measured and reproduced the failure independently. The lesson is narrower and more useful than "agents can fix bugs now." A self-healing loop is only as good as its ability to reproduce the failure a second time.

What we are doing next

We are working on a self-healing loop that will be frictionless to set up, that can run in your own software factory's loop, so you can see the same results.

If you'd like to participate in the beta, tell us about your software factory and we'll follow up with you.

📧 Get on the beta list