Voitta Labs · The Effort Study the turning box  ·  the subject study  ·  the model index

Does thinking harder change the painting?

The main index varies the model. This one holds the model still and turns a single dial: how hard it is told to think. Same prompt, same one shot, every panel at identical size, so the only thing that differs across a row is the effort setting. Five providers, and no two dials alike — named rungs on Claude Code, Codex, xAI and DeepSeek, a token budget on Gemini — so each gets its own grid rather than pretending the levels line up.

Render Leonardo da Vinci's "The Last Supper" as a detailed SVG. Save it as last_supper.svg.

Anthropic · Claude Code complete

Five levels, set with --effort. Network, shell and file-reading tools were removed, so these models could only write. Tools actually called, read back from the session transcripts: Edit, Write.

lowmediumhighxhighmaxmore thinking →
Fable 5.1
Fable 5.1, low effort
low2 tries4m42s · 17 KB
Fable 5.1, medium effort
medium3 tries14m37s · 30 KB
Fable 5.1, high effort
high3 tries25m22s · 41 KB
Fable 5.1, xhigh effort
xhigh18m29s · 43 KB
Fable 5.1, max effort
max32m21s · 29 KB
Opus 5.5
Opus 5.5, low effort
low1m26s · 15 KB
Opus 5.5, medium effort
medium6m59s · 35 KB
Opus 5.5, high effort
high11m00s · 44 KB
Opus 5.5, xhigh effort
xhigh33m09s · 40 KB
Opus 5.5, max effort
max47m44s · 65 KB
Sonnet 5.5
Sonnet 5.5, low effort
low1m14s · 13 KB
Sonnet 5.5, medium effort
medium2m31s · 23 KB
Sonnet 5.5, high effort
high8m39s · 31 KB
Sonnet 5.5, xhigh effort
xhigh25m06s · 31 KB
Sonnet 5.5, max effort
max2 tries70m54s · 61 KB
Sonnet 5
Sonnet 5, low effort
low1m30s · 11 KB
Sonnet 5, medium effort
medium2m02s · 21 KB
Sonnet 5, high effort
high6m12s · 25 KB
Sonnet 5, xhigh effort
xhigh12m50s · 36 KB
Sonnet 5, max effort
max23m13s · 26 KB
Haiku 4.5
Haiku 4.5, low effort
low40s · 8 KB
Haiku 4.5, medium effort
medium59s · 10 KB
Haiku 4.5, high effort
high50s · 10 KB
Haiku 4.5, xhigh effort
xhigh47s · 6 KB
Haiku 4.5, max effort
max47s · 8 KB

OpenAI · Codex CLI complete · 5 refused by the provider

Seven levels, set with model_reasoning_effort. Latest generation only, one model per size class. The shell, network access, image viewing and image generation are all switched off, so these models compose the markup themselves. Codex's web tool is served by the model's own harness and no client flag reaches it, so it is closed by instruction instead and every cell is audited afterwards to show whether that held. Tools actually called, read back from the session transcripts: exec.

noneminimallowmediumhighxhighmaxmore thinking →
GPT-6 Astra
not available
this model cannot run at this setting
not available
this model cannot run at this setting
GPT-6 Astra, low effort
low6m56s · 26 KB
GPT-6 Astra, medium effort
medium8m00s · 31 KB
GPT-6 Astra, high effort
high19m44s · 46 KB
GPT-6 Astra, xhigh effort
xhigh30m15s · 76 KB
GPT-6 Astra, max effort
max33m50s · 73 KB
GPT-6 Sol
GPT-6 Sol, none effort
none2m02s · 13 KB
not available
this model cannot run at this setting
GPT-6 Sol, low effort
low7m34s · 18 KB
GPT-6 Sol, medium effort
medium4m44s · 29 KB
GPT-6 Sol, high effort
high5m35s · 29 KB
GPT-6 Sol, xhigh effort
xhigh5m49s · 29 KB
GPT-6 Sol, max effort
max9m17s · 40 KB
GPT-6 Luna
GPT-6 Luna, none effort
none2m06s · 13 KB
not available
this model cannot run at this setting
GPT-6 Luna, low effort
low1m37s · 10 KB
GPT-6 Luna, medium effort
medium1m27s · 9 KB
GPT-6 Luna, high effort
high2m17s · 13 KB
GPT-6 Luna, xhigh effort
xhigh4m19s · 22 KB
GPT-6 Luna, max effort
max4m19s · 21 KB
GPT-Reserve
GPT-Reserve, none effort
none1m44s · 11 KB
not available
this model cannot run at this setting
GPT-Reserve, low effort
low1m30s · 10 KB
GPT-Reserve, medium effort
medium1m46s · 11 KB
GPT-Reserve, high effort
high2m56s · 20 KB
GPT-Reserve, xhigh effort
xhigh3m37s · 22 KB
GPT-Reserve, max effort
max5m18s · 33 KB

Google · Gemini API complete · 2 refused by the provider

A token budget rather than named rungs: thinkingConfig.thinkingBudget, from 0 to the 32,768 ceiling, plus dynamic where the model chooses. Nothing had to be taken away here — a bare API call declares no tools at all, so there is no shell to script with and no way to fetch a reference. This is also the only leg that reports how many thinking tokens were actually spent, rather than leaving it to be inferred from the clock.

0102440961638432768dynamicmore thinking →
Gemini 3.1 Pro
not available
this model cannot run at this setting
Gemini 3.1 Pro, 1024 effort
10243,824 think1m07s · 12 KB
Gemini 3.1 Pro, 4096 effort
40966,815 think1m17s · 12 KB
Gemini 3.1 Pro, 16384 effort
1638411,289 think2m22s · 20 KB
Gemini 3.1 Pro, 32768 effort
3276812,329 think2m51s · 18 KB
Gemini 3.1 Pro, dynamic effort
dynamic13,065 think3m04s · 21 KB
Gemini 3.8 Flash
Gemini 3.8 Flash, 0 effort
01m01s · 36 KB
Gemini 3.8 Flash, 1024 effort
102447s · 28 KB
Gemini 3.8 Flash, 4096 effort
40962,786 think1m18s · 39 KB
Gemini 3.8 Flash, 16384 effort
163844,822 think1m45s · 50 KB
Gemini 3.8 Flash, 32768 effort
327684,061 think1m43s · 50 KB
Gemini 3.8 Flash, dynamic effort
dynamic9,767 think1m56s · 42 KB
Gemini 3.5 Flash-Lite
not available
this model cannot run at this setting
Gemini 3.5 Flash-Lite, 1024 effort
1024802 think16s · 9 KB
Gemini 3.5 Flash-Lite, 4096 effort
4096938 think19s · 11 KB
Gemini 3.5 Flash-Lite, 16384 effort
163843,513 think34s · 17 KB
Gemini 3.5 Flash-Lite, 32768 effort
327683,483 think31s · 14 KB
Gemini 3.5 Flash-Lite, dynamic effort
dynamic1,332 think30s · 18 KB

DeepSeek · DeepSeek API complete

The longest ladder here: reasoning_effort takes all seven of none, minimal, low, medium, high, xhigh, max, and none is a true off switch — both models return exactly zero thinking tokens for it. No tools are declared, so the markup is composed from recall. One caveat: this dial is real but noisy. Four runs at low spent 2,246 to 7,543 thinking tokens and four at max spent 7,344 to 14,399, so the medians differ threefold but a single cell can land out of order. Read the row as a trend, not cell by cell.

noneminimallowmediumhighxhighmaxmore thinking →
DeepSeek V4 Pro
DeepSeek V4 Pro, none effort
none1m04s · 17 KB
DeepSeek V4 Pro, minimal effort
minimal6,658 think3m15s · 30 KB
DeepSeek V4 Pro, low effort
low4,260 think2m19s · 22 KB
DeepSeek V4 Pro, medium effort
medium13,485 think3m59s · 25 KB
DeepSeek V4 Pro, high effort
high17,239 think4m45s · 25 KB
DeepSeek V4 Pro, xhigh effort
xhigh11,919 think3m27s · 24 KB
DeepSeek V4 Pro, max effort
max10,888 think3m52s · 28 KB
DeepSeek Flash
DeepSeek Flash, none effort
none18s · 14 KB
DeepSeek Flash, minimal effort
minimal5,141 think45s · 20 KB
DeepSeek Flash, low effort
low13,075 think1m25s · 23 KB
DeepSeek Flash, medium effort
medium17,669 think1m22s · 23 KB
DeepSeek Flash, high effort
high3,890 think38s · 19 KB
DeepSeek Flash, xhigh effort
xhigh14,165 think1m41s · 32 KB
DeepSeek Flash, max effort
max5,960 think55s · 25 KB

xAI · xAI API complete

Five rungs via reasoning_effort; the API refuses both none and max. Rows here are versions, not size classes — xAI ships no Pro/Flash equivalent, so “latest per size” is just Grok 4.7, and the two previous versions are shown for contrast. One tool is declared, write_file, and nothing else — no shell, no network — so the model still composes the markup itself from memory. Grok 4.7 needs it: without a file tool it reads “save it as last_supper.svg” as something it has already done and answers with one sentence and no drawing, fifteen times running. Reasoning tokens are reported directly.

minimallowmediumhighxhighmore thinking →
Grok 4.7
Grok 4.7, minimal effort
minimal140 think44s · 9 KB
Grok 4.7, low effort
low127 think1m02s · 12 KB
Grok 4.7, medium effort
medium278 think1m13s · 13 KB
Grok 4.7, high effort
high2 tries404 think31m42s · 19 KB
Grok 4.7, xhigh effort
xhigh2 tries400 think31m33s · 17 KB
Grok 4.6
Grok 4.6, minimal effort
minimal112 think39s · 10 KB
Grok 4.6, low effort
low154 think58s · 14 KB
Grok 4.6, medium effort
medium571 think1m56s · 28 KB
Grok 4.6, high effort
high3,220 think4m38s · 38 KB
Grok 4.6, xhigh effort
xhigh12,284 think7m30s · 43 KB
Grok 4.5
Grok 4.5, minimal effort
minimal208 think2m19s · 29 KB
Grok 4.5, low effort
low2 tries90 think31m10s · 13 KB
Grok 4.5, medium effort
medium116 think1m45s · 22 KB
Grok 4.5, high effort
high183 think1m48s · 22 KB
Grok 4.5, xhigh effort
xhigh459 think2m28s · 30 KB

Reading a cell. Each panel shows the file a model produced, at identical size across the whole study. Attempts above one mean an earlier try left no usable file; the prompt was re-sent unchanged. Not available means the provider refuses that setting outright — a fact about the model, not a failure to draw. Think is the reasoning tokens actually spent, where the provider reports them.

Composed or generated. In operation 04 the badge gives the verdict — computed, mixed or transcribed — from how much of the finished file was already sitting in the program as literal markup. A transcription hand-writes the SVG and wraps two lines of Python round it; that runs, but it is not a generator. Nothing in the prompt says which to do, so the choice is a result.

No reference, either operation. Every leg ran with the local MCP servers removed and no way to see the painting: the network, shell and file-reading tools are denied on Claude Code, the shell and image tools are switched off on Codex, and the three API legs declare no tools at all. Codex's web tool is served by the model's own harness, so it is closed by instruction and every cell is audited afterwards against its transcript. In operation 04 each program additionally runs in an empty directory with sockets and subprocess disabled, so it can neither fetch a reference nor read the drawings this repository already holds.

Does the dial do what the label says?

Does the dial do what the label says?

Three of the five legs report the thinking tokens actually spent, so for those the question is answerable rather than arguable: rank the settings as the provider orders them, rank the thinking spent, correlate. +1.00 means the dial tracks perfectly; 0 means the label predicts nothing. One sample per cell and a different number of rungs per model, so this is indicative — the raw counts are shown so it can be judged, not taken on trust.

LabModelrhothinking tokens by settingn
GoogleGemini 3.1 Pro+1.00tracks1024 3,824 · 4096 6,815 · 16384 11,289 · 32768 12,329 · dynamic 13,0655
GoogleGemini 3.8 Flash+0.93tracks0 0 · 1024 0 · 4096 2,786 · 16384 4,822 · 32768 4,061 · dynamic 9,7676
GoogleGemini 3.5 Flash-Lite+0.60loose1024 802 · 4096 938 · 16384 3,513 · 32768 3,483 · dynamic 1,3325
xAIGrok 4.7+0.80tracksminimal 140 · low 127 · medium 278 · high 404 · xhigh 4005
xAIGrok 4.6+1.00tracksminimal 112 · low 154 · medium 571 · high 3,220 · xhigh 12,2845
xAIGrok 4.5+0.40looseminimal 208 · low 90 · medium 116 · high 183 · xhigh 4595
DeepSeekDeepSeek V4 Pro+0.64loosenone 0 · minimal 6,658 · low 4,260 · medium 13,485 · high 17,239 · xhigh 11,919 · max 10,8887
DeepSeekDeepSeek Flash+0.43loosenone 0 · minimal 5,141 · low 13,075 · medium 17,669 · high 3,890 · xhigh 14,165 · max 5,9607
←→ effort   ↑↓ model   tab: draw / program   esc: close