← Articles

// FIELD NOTE

GPT-6 Sol Did Astra's Work for 40% of the Cost

Cory LaNou

Cory LaNou

GPT-6 Sol Did Astra's Work for 40% of the Cost

GPT-6 Sol Did Astra's Work for 40% of the Cost

Overview

We run AI agents on real client code every day, and every one of those runs costs credits. Our default has been GPT-6 Astra at low effort. When GPT-6 Sol came out at a fifth of the token price, the obvious question was whether we could just switch. The launch benchmarks looked good, but benchmarks aren't our clients' code, and a cheaper model that misses requirements just moves the cost into review and rework. So the night Sol launched, we replayed eighteen issues we'd already merged, across three client codebases, and graded every run. The video walks through what we found, including the one project where switching looked worse. This post is the longer version, with the per-project numbers, the receipt bug in detail, and the exact recipe so you can run the same test on your own repo.

The short version

Sol at high effort did work our blind reviewer rated just as good, for about forty percent of Astra's credits. The reviewer accepted eighteen of Sol's twenty-four runs, and sixteen of Astra's. Turning Sol up to extra-high cost more and didn't get us a single extra accepted change.

Now, there are caveats. This is eighteen issues, not hundreds. Sol took about three times as long per run. And on the one project where the pricing math has to be exact, Sol actually did worse. I'll go through the numbers first, and then how you can run the same test to figure out what your own projects should be using.

Why test at all

At launch, OpenAI showed Sol at extra-high effort ahead of Astra low on its AutomationBench workflow test. So we compared Astra low with two Sol settings: extra-high, written as xhigh in the configuration, and high. The question was whether either could give us comparable work for no more than half the credits.

We set that bar before looking at the results. Comparable meant the merge-ready rate had to land within five percentage points of Astra. If you set the bar after you see the numbers, you'll set it wherever the numbers landed, and the test tells you nothing.

The replay setup

The project names here are fictional. The code, the issues and the results are real.

  • Ferrydock is a database replication library.
  • Lanternhall is a multi-tenant application with billing and third-party integrations.
  • Kilnworks handles retail pricing and receipts.

We chose six previously merged issues from each, ranging from small fixes to larger features. For every run, we went back to the commit before the real pull request and gave the agent its own worktree. It got the issue text and the same instruction to implement the change with tests and follow the repository's rules. Then we removed the GitHub credentials, blocked pushes, and made sure a push actually failed. These were replays, so none of the candidates should be changing the live project.

Three ways to grade

Then we graded the results three ways, because any one of them on its own would have fooled us.

First, we ran the tests from the real merged change against each candidate. Second, a separate model reviewed the issue, the reference diff and three unlabeled candidates, without knowing which model wrote them. Third, we recorded credits from each run's usage report. That gave us a test result, a blind merge-readiness judgment and a cost, rather than treating any one of them as the whole answer.

We planned more repeats, but stopped after one full round and reran only the tasks where the settings disagreed. That left seventy-two runs, twenty-four per setting. The reruns aren't another random sample, so keep that in mind when we look at the totals.

The headline numbers

Sol high used about forty percent of Astra's credits. The blind reviewer marked eighteen of its twenty-four runs merge-ready, compared with sixteen for Astra. That's enough to pass the bar we set, but two additional judgments in a small test don't prove Sol is the better model. What we have is a promising cost result with comparable judged quality on this work.

Here are the totals across all seventy-two runs:

  • Astra low: 3,088 credits, 16 of 24 merge-ready, median run 5.6 minutes.
  • Sol extra-high: 1,986 credits (64% of Astra), 18 of 24 merge-ready, median run 16.6 minutes.
  • Sol high: 1,231 credits (40% of Astra), 18 of 24 merge-ready, median run 15.9 minutes.

If we divide all the credits, including failed runs, by the merge-ready results, Astra came to about a hundred and ninety-three credits per accepted candidate. Sol high came to sixty-eight. Those are candidates accepted by the blind reviewer, not pull requests shipped through our full production process. The first round alone pointed the same way: fifteen of eighteen for Sol high against fourteen for Astra, at thirty-six percent of the credits.

More effort, same result

Extra-high didn't improve on that. It also got eighteen of twenty-four merge-ready, but used sixty-four percent of Astra's credits, so it missed our cost bar. One run used three hundred and four credits where Sol high used thirty-nine in the same round. Another reached the ninety-minute cap. More effort didn't buy more accepted changes in this test.

The time tradeoff

There was also a time tradeoff. The median run was about sixteen minutes for Sol high and five and a half for Astra. If you're waiting on one fix before you can continue, that difference matters. If the work runs while you're doing something else, you may be willing to wait. The replay doesn't tell us how that tradeoff changes once an entire queue is running on Sol, and that's part of what the follow-up week is for.

The costliest miss

The expensive default had its own costly failure. On the hardest task, a ledger deadlock in Lanternhall, Astra low ran for eighty-seven minutes and used eight hundred and forty-three credits, about twenty-seven percent of its entire test bill. Sol high used a hundred and forty-eight credits on that task. Neither produced a merge-ready fix. No setting solved it in this replay, and the production fleet had needed nine attempts on the real issue. Spending more on a single pass didn't remove the need for another attempt.

Split by project

The project breakdown is where I'd be careful about changing defaults.

  • Ferrydock: Sol high merge-ready on 7 of 7 runs, Astra on 5 of 7.
  • Lanternhall: Sol high 6 of 7, Astra 4 of 7.
  • Kilnworks: Sol high 5 of 10, Astra 7 of 10.

On the two projects that are mostly code, Sol high won. On the pricing project, it lost. An overall average would hide the one project where switching looked worse.

The receipt that counted twice

The receipt-savings task explains why that exception matters. The issue explicitly said not to count the same discounts twice, and Sol high did exactly that in both runs.

Here's a simplified example to make the mistake visible. These aren't amounts from the client. Suppose two item discounts are three dollars and two dollars, and a five-dollar summary discount already includes both. The customer's savings are five dollars. Adding that summary to the two items would report ten.

A worked example like that gives the implementation and the review a concrete answer to check. It doesn't establish that a better spec would have rescued these runs; we'd have to rerun them to know. Extra-high made a different receipt mistake: when any item was returned, it applied return-only math to the whole receipt. Astra handled that task correctly both times. That's a reason to keep an exception for this project while we test further.

What the tests missed

You might expect the hidden tests to catch all of this for us. Across Ferrydock and Lanternhall, every setting passed or failed the real pull request's tests identically. Some failures depended on matching the reference implementation's names or error wording. Meanwhile, the blind review found differences in behavior and scope that those tests didn't separate. So the reviewer carried much of the quality verdict, and that verdict is still a model's judgment, not independent human approval of every candidate.

The limits

Eighteen tasks is a small sample, and we only repeated the disagreements. A machine reboot forced us to discard four runs and start them again, and three tasks later ran under heavy load. Excluding those tasks still left Sol high at about thirty-four percent of Astra's credits, but I'd be cautious about treating these timings as clean latency benchmarks. Each run was also one implementation pass, without the normal cycle of review and rework that a real pull request goes through.

What I'd change

So I wouldn't switch every project on the strength of this table. I'd use it to choose a live trial: Sol high for the projects where it held up, with the pricing project staying on Astra until we have better evidence.

Put together, that's the rule I'd start with. Sol high for routine work that can run in the background. Astra when you're waiting on the fix or the math has to be exact. Extra-high nowhere until something shows it earns its cost. And I'd measure the credits per merged pull request through the whole loop, because a cheap first pass can become expensive if it needs repeated repairs. Those are proposed changes, not a fleet migration we've already completed.

Managing it at every stage

That choice also has to survive day-to-day work. Without orchestration, you're managing model and effort across planning, implementation, review and audit yourself. It's easy to leave the expensive setting everywhere, or cut too far and spend the savings fixing mistakes. Across a queue, even a small team has a reason to care about that repetition.

Detent lets us set defaults per project, select model and effort from configured complexity rules, and give the review validator its own model. This is the relevant piece of one of our project files:

# detent.yaml (per project) — which model and effort each task level gets
agents:
    model_selection:
        normal_model: gpt-6-sol
        complex_model: gpt-6-sol
        levels:
            normal:
                model: normal
                effort: medium
            complex:
                model: complex
                effort: medium

Those rules use task metadata; they aren't a promise that the system understands how hard every issue is. The review side is configured separately, so the model that grades a change doesn't have to be the model that wrote it:

gate:
    kind: command
    run: env GOMAXPROCS=8 GOFLAGS=-p=4 make check
    validator:
        enabled: false
        # Leave empty to inherit the route default, or pin the review model
        # separately from the model that wrote the change.
        model: ""
        min_score: 0.8

That's how we keep the pricing exception and ask for stronger review where it's useful, without making every stage use the same expensive setting. We still have to measure whether the policy works.

Run it on your own repo

If you're considering a switch, try this on your own last eighteen merged issues.

  1. Pick six merged issues from each of three projects, from small fixes to larger features.
  2. Check out the commit before each real pull request and give the agent its own worktree.
  3. Give every model the same issue text and the same instruction.
  4. Remove credentials, block pushes, and confirm a push fails.
  5. Decide your pass bar before you look at a single result.
  6. Grade three ways: the real tests, a blind review of unlabeled candidates, and the credits from the usage report.
  7. Look at the missed requirements alongside the credits, then split the results by project before choosing a default.

The follow-up is a live week on Sol, measuring the cost per merged pull request, including the review and rework it takes to get there.

Want more AI development insights?

Subscribe to the newsletter for weekly tips on using AI in professional development.

Subscribe to Newsletter

// KEEP READING

More articles