From tokenmaxxing to tokenoptimizing
We are all tokenmaxxers until the weekly quota is about to run out. Then, we rush to become tokenoptimizers. To curb this wayward behavior, we set out on a journey to find ways of optimizing a given agent harness to use cost effective models and agent architectures. Specifically, our goal was to reduce cost without sacrificing accuracy on two common engineering tasks that dominate our customers’ token spend: coding and incident response.
Routing each prompt to the cheapest capable model sounds great, but you rarely know how hard a task is until you're halfway through it. Switching models mid-task naively is also costly: each model keeps its own prompt cache, so every switch means re-writing the ever growing context into a new cache. We needed something smarter. Devin’s Fusion agent architecture is a great starting point, to which we add our own spin: per-turn dollar-budget reminders. The idea is for a smart but expensive model called "main” to drive a cheap but less smart model called “sidekick” while staying under the assigned budget.
Summary of results
We show that we can successfully re-configure two very different harnesses: Codex for coding tasks and our proprietary harness for post-coding platform tasks like incident response. Running Codex on a hard subset of DeepSwe 1.1, where the cheap model usually fails, fusion cut coding cost by 25% with nearly the same accuracy but ran 3x slower than the expensive solo model. On a privately curated benchmark of real production incidents, the Autoheal harness with fusion of a smarter model & cheaper model and dollar budgets returned 3x cheaper, 2x faster and 1.1x more accurate than with frontier model alone.
Agent Architecture

Fusion. The fusion architecture is a restricted form of the more general subagent architectures typically found in coding agents. The idea is to have two persistent agents that share all the tools, but differ in which sub-tasks they pick up. Powered by a smarter model, the main agent can decide what to delegate to the sidekick agent at any point in the session. Thus, there is no upfront routing and you ride the tailwind of frontier models getting better at delegating.
Budget reminders. The central challenge in the fusion design is defining how much work gets split between the two agents. More splitting increases communication overhead, while less splitting increases costs (the main agent does most of the work). Thus, what is the maximum amount of work that can be delegated to the sidekick without degrading overall quality?
Since frontier models are explicitly trained to be good at goal optimization, we experiment with assigning a dollar budget to each run and prompt the main agent to optimize cost. The budget reminders are only sent to the main agent and include a main vs. sidekick cost breakdown. They are sent every turn and are only used as an advisory signal; the agent is not terminated for going above the budget.
Coding Agent Cost
We first experimented with the baseline Devin Fusion architecture, no dollar budget, on a public dataset called DeepSwe 1.1. It's a popular benchmark used to evaluate coding agents, with a programming language mix roughly similar to ours. Here our goal is twofold: keep the benchmark close to our own coding agent setup so that results transfer well, and try to reproduce the Devin Fusion architecture.
Benchmark and the models. We started with GPT-6 Astra (High) and GPT-5.6 Luna (Max), but on DeepSwe1.1 the gap between their pass rate is pretty small: Astra 74% vs Luna 67% (using mini-swe harness). Running the whole benchmark is wasteful if the goal is to retain Astra-level accuracy while reducing cost. Thus, we focus on an 11-task subset of DeepSwe1.1 where Astra does well but Luna doesn't, 77% vs 16% pass rates. A side note: when we switch the harness from mini-swe to Codex the gap between the two models drops. While Astra’s accuracy stays the same, Luna’s increases from 16% to 36%.
Agent harness. We chose codex instead of the popular academic choice of mini-swe for the harness. Codex is what we actually use and love. We then implemented fusion in Codex which turned out to be surprisingly simple. Codex has tools that allow full flexibility in spawning and interacting with subagents. It is simply a matter of constraining the model to invoke and interact with subagents using following tool calls:
Use
spawn_agentwithfork_turns="none"to prevent copying context from main to sidekick when instantiating it.Use
followup_taskto communicate with the sidekick, which can wake up an idle sidekick.Use
wait_agentwith a large timeout to prevent constant checks on the sidekick.
These instructions go into the main agent prompt.
Results. Fusion was about 1.34x (25%) cheaper than the solo model with little loss in accuracy, though it took 3x longer to run. Definitely promising results but nothing spectacular. It would be a great fit for background tasks like modernization or migration or vulnerability remediation but not for live coding. Without this insight you might adopt this agent architecture, hear developers complain and reject this architecture when you could have reduced your coding costs substantially!

Next, let’s evaluate it on a different task and add the budget reminder trick to the implementation.
Incident Response Agent Cost
Benchmark. It is incredibly hard to find a high quality real-world public benchmark on incident response because that data is trapped within enterprise’s boundaries. Lucky for us we are deployed in some extremely sophisticated enterprises. We create a benchmark by snapshotting log data for 8 real incidents from our production environment, whose root causes were manually verified. Each task starts with a real ticket or alert and the agent must debug it using log data. Its proposed root cause must match the one accepted by the expert engineer. All results are averaged over 3 attempts per task.
Models. We expand to a wider variety of models to contrast how all frontier and open weight models behave on their own vs in the fusion pattern. We were also keen to see if you can fuse models from different AI labs to get the best of both worlds.
Agent harness. We have purpose-built our own agent harness for repetitive engineering work like incident response in extremely high volume enterprise environments. So we swap Codex for our own harness in these experiments. To emulate constraints of large enterprises, we set a hard budget of $2.50 per task, after which the task is terminated. The soft reminder budget is set at $1.50 per task and is only used in the fusion experiments. All open weight models are set to high reasoning effort and closed weight models to medium, except Luna, which is set to high.
Results. From the solo model results, we can see that Luna, MiniMax-M3, and Sonnet-5 offer the best balance between cost and accuracy. Opus keeps overshooting the hard $2.5 budget causing many of its runs to be incomplete.

Since our goal is to push the Pareto frontier, we run fusion experiments with the 3 models already at the frontier plus Sol and Kimi-2.7. The fusion results show that it not only reduces the cost but also increases the accuracy for almost all models!

The best fusion run, Sonnet 5 + Luna, is 3x cheaper and 2x faster than the most accurate solo run of Sonnet 5. Way stronger results than the coding agent benchmark. Why?

Debugging incidents using logs requires an agent to read a lot of big JSON payloads containing mostly irrelevant information which leads to a noisy context. Only a few log lines contain the evidence pointing to the actual root cause, but locating them is very hard. Thus, the agent must stack multiple needle-in-the-haystack searches to arrive at the real evidence.
When we let the sidekick do the mechanical work of assembling log lines and extracting information from them, we save the main agent's context and let it focus on the bigger picture. Furthermore, the budget reminders encourage the main agent to delegate work that the sidekick can perform. The budget also introduces a more practical axis for scaling agents, one tied directly to product outcomes rather than model internals like reasoning effort.
Conclusion
Fusion can be reproduced outside of the Devin CLI, in Codex, with minimal effort: 25% cost reduction with little loss of accuracy on a hard subset of DeepSwe1.1. When combined with budget reminders, it shows incredible promise for incident response agents, with a 3x cost and 2x time reduction.
If you want such improvements in your engineering agents, check out Autoheal’s self-improving software factory.
Book a demo today!



