Using production outcomes to improve coding agents
Would you drive on the highway blindfolded? If you’re like most people, then no. Yet, we subject our coding agents to the same fate. They can write PRs but have no ability to see how it behaves after merging. Did it cause a CI failure? How about a spike in prod incidents? What if the user engagement dropped? These are not just important signals that are invisible to the coding agent but also the key quality metrics for any piece of software. Our primary job should be to define these metrics, keep an eye on them, and let agents improve themselves.
At Autoheal, we've built a platform for exactly that: self-improving coding agents. You bring your coding agent while we provide all the self-improving scaffolding. You work on your PRs and define the quality metrics while we run a three-step improvement loop. First, trace and analyze PRs all the way to production and gather feedback on where they can be improved. Second, assemble the feedback into concrete coding agent modifications. Third, prove that the modifications actually work by turning your PRs into private, repeatable benchmarks. We don’t want to rely on hopes and prayers that our proposals work; we want to measure and show that they do.
Summary of results
None of the public benchmarks had enterprise-grade end-to-end data for the full lifecycle of a PR. Thus, we built our own small but high quality benchmark from a month of our engineering activity. We used human engagement as the signal to evaluate the PR quality. We also analyzed all CI/CD and production events and used Autoheal's AI coding optimizer agent to synthesize the learnings into six high quality skills to improve observability and testing. We then evaluated the effectiveness of the proposed skills by running coding agents on our private benchmark. Results:
The agent got more consistent. Before, 27% of the tasks passed consistently when attempted multiple times (passes in all attempts). After the skill improvement, 54% did.
Prior context files were unhelpful. As noted above, running the coding agent with existing context (AGENTS.md and skills) returned a consistency score of 27%. Removing the context increased the consistency score to 36%. Stale context is worse than having no context at all.
Metrics find the problem, benchmarks prove the fix

Quality metrics. As a PR moves from review to CI and then to production, it leaves a trail of evidence. Review comments, CI and staging failures, failed tests, alerts, incidents, and changes in user behavior all tell you something about the quality of the code. No single metric is enough. But taken together, they give you a useful picture of whether the change actually worked; well before the business impact shows up.
Private benchmarks. Metrics tell you something is wrong but not how to improve it. They might point you towards a fix but how can we know for sure that it works? Now, what if you could go back in time, make your hypothesized change, and re-run the whole process? This would verify the fix's effectiveness, and is it turns out is not that hard to do.
We rewind the repo to a commit before the PR, drop a coding agent into it with a prompt describing the task, and grade its output by running tests that encode the expected behavior. It's the closest thing to time travel we have. (Note: real time travel won’t be possible until January 4, 2031.)
We looked at public benchmarks first and passed. At best, models get hill-climbed on them. At worst, the tasks leak into training data. Either way, they say little about your own codebase. Public benchmarks also rarely cover anything beyond bug fixes and feature work, because those are the easiest tasks to turn into benchmarks.
How we built the experiment
We turned above ideas of quality metrics and private benchmarks into reality on our own engineering activity.
The benchmark. We combed through our repo's PRs and agent sessions, and picked 11 of the hardest tasks. These were complex, multi-part changes in Python, Go, and TypeScript that originally took many rounds of back-and-forth to land. We converted each into a single-shot prompt. Instead of the many turns a human needed to steer the original work, the agent gets one prompt and one shot.
Grading is pass/fail against deterministic tests: the agent must pass every test. The tests are loose enough to accept any valid approach and tight enough to reject a wrong one. Getting that balance right took most of our time. We used adversarial agents that tried to break the tests, plus plain manual review.
The skills. In parallel, agents on our platform watched code reviews, build pipelines, and production alerts on the same repo for 30 days. They turned each miss into concrete feedback: which log line would have sped up an incident investigation, or which test would have caught a regression before it shipped. From that feedback, Autoheal's optimizer agent wrote six skills, one on logging and five on testing (databases, Temporal workflows, deployments, alerts and metrics, and inter-service contracts). Together they add about 10k tokens of context. Excerpts from proposed skills:

The numbers
We ran GPT-5.6-Luna (max) in Codex under three conditions:
Existing Context (baseline): uses existing context files like CLAUDE.md in our repo.
No Context: those files removed, to see whether they were helping at all.
Prod Derived Skills: the baseline plus the six new skills, with an instruction to read them before editing.

Each task got three attempts. We measured consistency pass rate (all attempts pass), cost per task, and time per task. Cost and time barely moved, but the consistency score doubled from baseline’s 27% to prod derived skills’ 54%. The four hardest tasks failed in every condition. The agent didn't become smarter but became
The no-context row was a surprise. Getting rid of our context leads to a slightly more consistent agent, and we're a ~10-engineer AI startup that thinks about this for a living. Our best guess is that the files went stale as models improved over the three months since we last updated them. If you haven't tested your own AGENTS.md lately, it may be doing less than you think.
What makes the skills work? With skills, the agent added 33% more test code than with existing context. This is not surprising since the new skills contain testing guidelines, but increased consistency score points at higher quality tests. Unlike review comments which are based on simulations of failure scenarios, the skills are based on actual failures after the PR was authored by a coding agent. Your coding agent cannot typically see these failures or even if it does it cannot link them to the authoring agent session.
Could the skills be cheating? The feedback and the benchmark tasks came from the same 30 days, so it's fair to ask whether a skill accidentally became a cheat sheet for a benchmark PR. We checked. Four of the eleven benchmark PRs had some connection to the runs that produced the skills, but none of it explains the gains. The two tasks with a direct or indirect link failed every attempt in every condition. The two incidental mentions produced guidance unrelated to those PRs' code. This demonstrates that learning from failures generalizes to novel coding work in the repo.
So, does it work?
Reading the production outcome-aware skills made a measurable difference: the agent became far more consistent. This is still a small sample (11 tasks, three attempts each), and we can't yet separate the most important parts of the skills from spurious ones. Scaling up the benchmark and separating those two parts are next on our list.
You can try a version of this yourself tomorrow:
Run your coding agent on a few past PRs with and without your context files, and see whether they're earning their tokens.
Turn your last few postmortems into specific testing instructions ("execute X with Y absent; assert Z") instead of general advice.
This experiment is also why our platform doesn't stop at monitoring. The same loop, with quality metrics as input and benchmarks and skills as output, now runs in the AI coding evaluator and optimizer agents built on Autoheal. Any team can turn its own reviews, incidents, and postmortems into the same kind of evidence without doing it by hand. Book a demo →



