🔍 Read the full analysis: What My September 2026 AI Stack Looks Like In Practice on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A Sept. 29, 2026 article describes a practical AI workflow built around Opus 5.5 for development and newly released GPT-6.1 Sol for detailed work and review. The author bases the choices on Artificial Analysis Intelligence Index v4.3.x scores and estimated task costs, while noting that benchmark results do not establish performance on every workload.
Thorsten Meyer published a report on Sept. 29 describing an AI workflow that uses Claude Opus 5.5 for building and newly released GPT-6.1 Sol for detailed analysis and review. The account compares models using Artificial Analysis Intelligence Index v4.3.x scores and estimated costs per task, arguing that cost can shape practical model choices when benchmark scores are close.
Meyer says Opus 5.5 is his main model for features, APIs, multi-file work and refactoring. He uses it at high effort, which the cited index lists at 54 points and $1.82 per task, or at xhigh effort for more demanding architecture, migration and trust-boundary work, listed at 56 points and $3.46 per task. At max effort, Opus scores 58 but costs $5.98 per task in the same comparison.
GPT-6.1 Sol launched on Sept. 29 at the same stated token prices as its predecessor, GPT-6 Sol: $2 per million input tokens and $10 per million output tokens. The index lists Sol at medium effort with a score of 48 and estimated cost of $0.21 per task; high scores 50 at $0.32, and xhigh scores 51 at $0.39. Meyer assigns it specific-file investigations and independent reviews of changes made with Opus.
The account gives other models narrower roles. Meyer lists GPT-6 Astra for agents and computer use, Sonnet 5.5 for scoped work such as documents and slides, and Luna for classification, extraction and routing. He describes Fable 5.1 and Astra as alternatives for particular tasks, rather than default choices. The source also reports that Jev, a decision model that cannot write sentences, handles high-volume yes-or-no and routing decisions.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
How Review Fits the Cost Curve
The workflow illustrates how a low estimated cost per task can make a second-model review practical as a routine step. Meyer says Sol at high or xhigh costs $0.32 to $0.39 per task in the index comparison, and he uses it to check meaningful changes made with Opus. He argues that using a different model family can provide a more useful second opinion than asking the original model to review its own output.
That is a reported practice, not evidence that the review catches every defect. Meyer cautions that reviewers can inherit problems from the same flawed requirements, that higher effort cannot supply missing requirements, and that passing tests alone does not establish that work is ready to ship. The account makes the human review time part of the cost calculation too: an illustrative example says a minute of extra human review can outweigh savings from cheaper model tokens.
Benchmark Scores Meet Task Costs
The comparison uses Artificial Analysis Intelligence Index v4.3.x, which Meyer describes as a map of general capability rather than a verdict on a reader’s own workload. In the article’s table, Opus 5.5 scores 58 at its top setting, while Sonnet 5.5 scores 56, Fable 5.1 and GPT-6 Astra each score 53, and GPT-6 Luna scores 37. The reported per-task costs range from $0.07 for Luna to $7.60 for Sonnet 5.5 at max effort.
Meyer highlights effort level as a cost driver within a model family. For Opus 5.5, the index table puts medium at 51 points and $1.34 per task, high at 54 and $1.82, xhigh at 56 and $3.46, and max at 58 and $5.98. He chooses high or xhigh for development, saying the extra effort is for harder problems. He recommends medium for documents and routine work, while identifying Sonnet 5.5 at high effort as its best value in his comparison.
The article reports that Sol’s high and xhigh settings took 57 and 69 seconds to produce a first token, respectively, according to the index. It also says Sol’s high setting used 25 million output tokens on the index, against a reported median of 82 million for comparable models. These are benchmark observations; they do not by themselves establish how quickly or economically the models will perform in a particular team’s workflow.
““shadow-test before you switch anything.””
— Thorsten Meyer
Limits of the Benchmark Comparison
The source does not provide independent evidence that the reported cost estimates or index scores predict results across different workloads. Meyer says readers should shadow-test before switching models. The account also states that Artificial Analysis had not yet published low or max settings for GPT-6.1 Sol as of the article’s publication, and that a one-point score difference falls within the noise.
It remains unclear how the described workflow performs over a larger set of projects, how often Sol’s reviews identify problems that would otherwise be missed, and how total costs change when human review time is included. The source calls its example about human review illustrative rather than measured. It also does not specify the underlying task mix or provide enough detail to independently reproduce every per-task cost estimate.
Test the Stack on Real Work
Meyer recommends shadow-testing models against the tasks they would handle before making a switch. For teams considering the same division of work, that means comparing output quality, review findings, latency and total task costs on their own examples, while keeping benchmark scores in their stated context. The source does not name a future evaluation date or report a planned follow-up; its next step for readers is to test the choices against their own workload.
Key Questions
What is the main model in Meyer’s workflow?
Meyer says he uses Claude Opus 5.5 as his main model for building features, APIs, multi-file changes and refactors, usually at high or xhigh effort.
What does GPT-6.1 Sol do in the workflow?
The account assigns GPT-6.1 Sol detailed investigations and review of work produced with Opus. Meyer lists its high and xhigh estimated costs at $0.32 and $0.39 per task in the cited index.
What benchmark supports the comparisons?
The article cites the Artificial Analysis Intelligence Index v4.3.x. Meyer describes it as a general capability measure and says readers should test models on their own workloads before switching.
Does the article establish that Sol is better value for every team?
No. The reported scores and per-task costs are benchmark comparisons, and the source does not establish how they translate to every workload. Meyer recommends shadow-testing, and says a one-point score difference is within the noise.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
