🔍 Read the full analysis: Where Opus, Sol, And Jev Fit In My AI Workflow on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Thorsten Meyer’s Sept. 29 workflow assigns Opus 5.5 to building, GPT-6.1 Sol to detailed review and Jev to high-volume yes/no decisions. The approach reflects a reported gap in task costs among models with relatively close benchmark scores, though the figures come from one index and may not predict performance on individual workloads.
Meyer’s reported default for development is Opus 5.5 at high effort, which the index scores at 54 and prices at $1.82 per task. He uses xhigh for harder work such as architecture, migrations and trust boundaries; that setting scores 56 at $3.46 per task. He says he rarely uses max effort, which scores 58 but costs $5.98 per task.
For a second review, Meyer uses GPT-6.1 Sol at high or xhigh. The index lists those settings at 50 points and $0.32 per task, and 51 points and $0.39 per task, respectively. Meyer says he assigns Sol focused reviews of files or code changes, rather than asking it to build. Its high and xhigh settings have reported times to first token of 57 and 69 seconds, making them less suited to interactive use.
The workflow assigns other models narrower roles: Sonnet 5.5 for scoped subtasks and documents, Astra or Fable for a second opinion when models disagree, and Luna for classification, extraction and routing. Meyer describes Jev as a decision model that cannot write sentences, used for routine yes-or-no and routing judgments. The source does not give Jev a score, price or release date.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
A Review Pass at Lower Cost
The workflow treats model choice as a question of quality at a given task cost, rather than picking one model for every job. If Meyer’s reported prices and scores hold for his workload, assigning routine review to a lower-cost model could make it practical to check more changes. He says a different model family reviewing Opus’s output provides a useful second perspective, while acknowledging that two models can still share a flawed specification.
The figures also suggest that the effort setting can materially affect cost. On the index, Opus 5.5 rises from 51 points and $1.34 per task at medium to 58 points and $5.98 at max. Those are benchmark results, not a guarantee that the most expensive setting will produce proportionately better outcomes for a particular team. Meyer’s advice to shadow-test before switching systems recognizes that distinction.
The Index Behind the Workflow
The source attributes its model scores and task costs to the Artificial Analysis Intelligence Index v4.3.x, describing it as a general capability measure rather than a verdict on any one workload. Its Sept. 29 snapshot lists Opus 5.5 at 58 points and $5.98 per task at its top setting; GPT-6.1 Sol at 51 and $0.39 at xhigh; and Luna at 37 and $0.07. The reported figures use the index’s task-cost estimates, not simply token prices.
Meyer also reports that GPT-6.1 Sol launched Sept. 29 at the same stated token prices as GPT-6 Sol: $2 per million input tokens and $10 per million output tokens. He says Sol’s medium setting scores 48 at $0.21 per task. These are figures cited in the source article; the material provided does not include independent measurements of Meyer’s own workloads.
“The practical reading: Sol is not the model I ask to build. It is the model I can afford to run on everything.”
— Thorsten Meyer
Workload Results Remain Unverified
The source does not provide controlled comparisons of these models on the same software tasks, nor does it show how often Sol catches defects in Opus’s work. Index scores and estimated task costs may differ from results in a reader’s own workflow. Meyer recommends shadow-testing before a switch, but the article does not report the results of such a test.
Some benchmark details are also incomplete. According to the source, Artificial Analysis had not published GPT-6.1 Sol’s low or max settings, and a one-point difference falls within measurement noise. Jev’s pricing, capabilities beyond the stated decision role, and evaluation results are not supplied. The source material also ends partway through a cost example, so it does not establish a complete comparison of model costs with human review time.
Test the Split on Real Tasks
Meyer’s stated next step for teams considering a similar setup is to shadow-test models on their own work before changing defaults. That means checking whether the models meet a team’s quality requirements and whether review savings persist once human oversight is included. The source does not announce a formal rollout, a follow-up benchmark date or further details about Jev.
For now, the workflow is Meyer’s account of how he uses the models as of Sept. 29, 2026. Readers should treat the reported prices, settings and scores as a dated snapshot; the material does not say when those figures will next be updated.
Key Questions
What does Meyer use Opus 5.5 for?
He uses Opus 5.5 at high effort for feature work, APIs, multi-file changes and refactors. He reserves xhigh for harder tasks such as architecture and migrations.
What role does GPT-6.1 Sol play?
Meyer uses GPT-6.1 Sol for detailed inspection and review, including focused checks of files and code changes. He says its lower reported task cost makes routine review more practical for his workflow.
What is Jev, according to the source?
The source describes Jev as a decision model that cannot write sentences. Meyer uses it for high-volume yes-or-no judgments and routing; the source gives no score or price for it.
Are the reported scores proof these models will work best for every team?
No. The scores come from the Artificial Analysis Intelligence Index v4.3.x, which the source characterizes as a general capability measure. Meyer advises testing models on a team’s own workload before switching.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
