Match the thinking to the task.

Field study · model & effort

How to choose a setting

  1. Start with the lowest effort that reliably clears the work.
  2. Raise it for an observed reason: a failed attempt, contradictory findings, a missing dependency, an unresolved tradeoff.
  3. When a run fails, name the failed requirement first. Check the evidence, the instructions and the tool results before raising effort or changing models. Higher effort cannot supply missing evidence.
  4. To compare two settings, run the same task in fresh sessions against the same checks. Keep the lowest setting that passes.

A worked example: my own routing, September 2026

From my log

My starting point in September 2026, not a rule. Check it on your work.

Explore another settingLess effort → more effort

Rows show documented API levels, checked September 12, 2026; an app’s menu may show fewer. Each row has its own scale. Matching labels do not mean equal capability or compute.

All twelve starting points, in plain text

Travis’s September 2026 field study. Alternatives are equal starting points. An absent trigger means no condition was recorded, not that changing effort is forbidden.

All twelve task cases
Task / environmentStarting pointReason / noteStep up or changeStep down
Dictated notes, capture, summaries (ChatGPT in the browser and on my phone)GPT-5.6 Sol · Low or GPT-5.6 Sol · MediumThe priority is faithful capture, not unsolicited analysis. Astra Light is fine if it is already selected.No condition recorded.No condition recorded.
Conversation and thinking aloud (ChatGPT in the browser and on my phone)GPT-5.6 Sol · HighSol carries ordinary conversation well and keeps the conversational and editorial character I like.GPT-6 Astra · Medium: when another analytical approach or stronger constraint tracking would help.No condition recorded.
Comparing tools, untangling questions (ChatGPT in the browser and on my phone)GPT-6 Astra · MediumMedium is Astra’s broad analytical band: quality and judgment without paying for High on every question.GPT-6 Astra · High: when important tradeoffs or contradictory evidence are being flattened.No condition recorded.
Serious research and analysis (ChatGPT in the browser and on my phone)GPT-6 Astra · HighSource quality still matters more than the label.GPT-6 Astra · Extra high: for unusually difficult synthesis, calculations, or verification.No condition recorded.
Art direction and reference analysis (ChatGPT in the browser and on my phone)GPT-5.6 Sol · High or GPT-6 Astra · MediumImage concepts, Suno direction, reference analysis. Higher effort does not select a better renderer and does not guarantee taste.one rung up on the model in use when many creative constraints interact, or repeated attempts drift.No condition recorded.
Small diffs and explanations (ChatGPT desktop app, reviewing my coding agent)GPT-6 Astra · MediumLocating code, explaining a function, a basic summary: a small inspection does not need the review default.GPT-6 Astra · Extra high: when the request becomes a substantive correctness review.No condition recorded.
Reviewing substantial changes (ChatGPT desktop app, reviewing my coding agent)GPT-6 Astra · Extra highCross-file behavior, architecture, requirements. Review is what I use the desktop app for, so Extra High is its standing default. Start with the change set; expand into dependencies only when warranted.GPT-6 Astra · Max: when one specific hard issue is still unresolved after a checked attempt. Include the failed attempt, the evidence, and the exact issue.GPT-6 Astra · Medium: for a small inspection, or bounded follow-through after the review.
Visual review of an interface (ChatGPT desktop app, reviewing my coding agent)GPT-6 Astra · Extra highCode inspection alone cannot tell whether a design works visually. With screenshots or the running interface in front of the reviewer.No condition recorded.No condition recorded.
Routine implementation (Claude Code in VS Code)Claude Opus 5 · MediumFamiliar patterns and scoped fixes rarely need more.Claude Opus 5 · High: when dependencies are unclear, reasoning is shallow, or fixes fail.No condition recorded.
Difficult implementation (Claude Code in VS Code)Claude Opus 5 · HighHigh is the vendor’s starting point on Opus 5. Extra High is justified by the work, never a standing default.Claude Opus 5 · Extra high: only if the specific problem remains unresolved.No condition recorded.
Substantial features, nuanced design (Claude Code in VS Code)Claude Fable 5.1 · HighMany interacting requirements need the planning High buys.No condition recorded.Claude Fable 5.1 · Medium: for mechanical follow-through once the decisions are settled.
Architecture, migrations, hard debugging (Claude Code in VS Code)Claude Fable 5.1 · Extra highBroad dependency changes and hard bugs are where I spend Extra High: the vendor’s setting for capability-sensitive, long-horizon work.Claude Fable 5.1 · Max: for one clearly defined frontier problem after a checked attempt.Claude Fable 5.1 · High: once the problem is understood and bounded.
Evidence & how the settings differ

Use the lowest effort that reliably clears the work, and escalate for an observed reason: a failed attempt, contradictory findings, a missing dependency, an unresolved tradeoff. Higher effort can improve planning and checking. It can also take longer, cost more, and expand past what you asked for. When a run fails, identify the failed requirement. Check the evidence, the instructions and the tool result before deciding whether to raise effort or change models.

Two layers, kept apart. The presets are my own routing, logged in September 2026 across the three places I work: where I start, and the conditions I actually change on. They are one practitioner's record, not measured task rankings; no matched trials sit behind them. The ladders reflect documented API support, checked September 12, 2026. Short cues combine vendor summaries with this guide’s practical advice; they are not measured task rankings. App menus can differ: availability depends on rollout, sign-in method and client, some menus show only a subset of these levels, and ChatGPT can label Low as Light.

Model choice changes the underlying capabilities. Effort changes how that model approaches the work. Higher effort cannot supply missing evidence.

None is a separate setting available here only for Sol; Astra’s API effort range starts at Low. Extra High and Max are different levels and stay separate. Anthropic recommends starting Fable 5.1 and Opus 5 at High; on Opus 5, thinking can be switched off at High or below.

The short version of my routing
  1. Capturing or chatting? Start with Sol.
  2. Untangling or researching? Astra Medium, or High.
  3. Reviewing my coding agent's consequential work? Astra Extra High.
  4. Implementing routine code? Opus 5 Medium.
  5. Implementing difficult code? Opus 5 High.
  6. Building substantial systems or making nuanced design decisions? Fable 5.1 High.
  7. Architecture, migrations, or hard bugs? Fable 5.1 Extra High.
  8. Still stuck on one clearly defined hard problem? Try Max, with the failed attempt and the exact issue in the prompt.
  9. Can the task be divided into independent workstreams? Consider explicit delegation, Ultra, or Ultracode.

Highest useful productivity, not highest available setting.

Orchestration modes: Ultra and Ultracode

Two product settings sit outside these effort ladders because they change how the work is organized, not only how hard one model thinks.

Ultra, in ChatGPT Work on eligible accounts and supported models, uses maximum reasoning and lets ChatGPT delegate suitable work to subagents on its own when parallel agents would materially improve speed or quality. At other levels you ask for subagents explicitly. Current local Codex releases spawn agents only after a direct request or an applicable project or skill instruction, so the same setting name does not promise the same behavior in every client.

Ultracode, in Claude Code, is a session setting rather than a model effort value. It sends Extra High to the model and switches on dynamic workflows for substantive tasks.

Use either only when independent workstreams make parallel work worth its overhead. Neither is a higher rung on a single model's ladder.

Benchmark snapshot

Artificial Analysis Intelligence Index v4.3, read from the provider pages on September 8, 2026. Each cell is the index score followed by the estimated cost in US dollars per index task. These are weighted estimates for one composite benchmark; they are not prices for your work, not a subscription allowance, and not a ranking of which model suits a given task. Fable 5.1 rows are the "default fallback" configuration on that site. None-effort rows are omitted. These historical values were preserved from the September 10 build and were not independently rechecked during the September 12 recovery.

Index score · USD per index task