The main index varies the model. This one holds the model still and turns a single dial: how hard it is told to think. Same prompt, same one shot, every panel at identical size, so the only thing that differs across a row is the effort setting. Five providers, and no two dials alike — named rungs on Claude Code, Codex, xAI and DeepSeek, a token budget on Gemini — so each gets its own grid rather than pretending the levels line up.
Five levels, set with --effort. Network, shell and file-reading tools were removed, so these models could only write. Tools actually called, read back from the session transcripts: Edit, Write.
Seven levels, set with model_reasoning_effort. Latest generation only, one model per size class. The shell, network access, image viewing and image generation are all switched off, so these models compose the markup themselves. Codex's web tool is served by the model's own harness and no client flag reaches it, so it is closed by instruction instead and every cell is audited afterwards to show whether that held. Tools actually called, read back from the session transcripts: exec.
A token budget rather than named rungs: thinkingConfig.thinkingBudget, from 0 to the 32,768 ceiling, plus dynamic where the model chooses. Nothing had to be taken away here — a bare API call declares no tools at all, so there is no shell to script with and no way to fetch a reference. This is also the only leg that reports how many thinking tokens were actually spent, rather than leaving it to be inferred from the clock.
The longest ladder here: reasoning_effort takes all seven of none, minimal, low, medium, high, xhigh, max, and none is a true off switch — both models return exactly zero thinking tokens for it. No tools are declared, so the markup is composed from recall. One caveat: this dial is real but noisy. Four runs at low spent 2,246 to 7,543 thinking tokens and four at max spent 7,344 to 14,399, so the medians differ threefold but a single cell can land out of order. Read the row as a trend, not cell by cell.
Five rungs via reasoning_effort; the API refuses both none and max. Rows here are versions, not size classes — xAI ships no Pro/Flash equivalent, so “latest per size” is just Grok 4.7, and the two previous versions are shown for contrast. One tool is declared, write_file, and nothing else — no shell, no network — so the model still composes the markup itself from memory. Grok 4.7 needs it: without a file tool it reads “save it as last_supper.svg” as something it has already done and answers with one sentence and no drawing, fifteen times running. Reasoning tokens are reported directly.
Reading a cell. Each panel shows the file a model produced, at identical size across the whole study. Attempts above one mean an earlier try left no usable file; the prompt was re-sent unchanged. Not available means the provider refuses that setting outright — a fact about the model, not a failure to draw. Think is the reasoning tokens actually spent, where the provider reports them.
Composed or generated. In operation 04 the badge gives the verdict — computed, mixed or transcribed — from how much of the finished file was already sitting in the program as literal markup. A transcription hand-writes the SVG and wraps two lines of Python round it; that runs, but it is not a generator. Nothing in the prompt says which to do, so the choice is a result.
No reference, either operation. Every leg ran with the local MCP servers removed and no way to see the painting: the network, shell and file-reading tools are denied on Claude Code, the shell and image tools are switched off on Codex, and the three API legs declare no tools at all. Codex's web tool is served by the model's own harness, so it is closed by instruction and every cell is audited afterwards against its transcript. In operation 04 each program additionally runs in an empty directory with sockets and subprocess disabled, so it can neither fetch a reference nor read the drawings this repository already holds.
Five levels, set with --effort. Network, shell and file-reading tools were removed, so these models could only write.
Seven levels, set with model_reasoning_effort. Latest generation only, one model per size class. The shell, network access, image viewing and image generation are all switched off, so these models compose the markup themselves. Codex's web tool is served by the model's own harness and no client flag reaches it, so it is closed by instruction instead and every cell is audited afterwards to show whether that held.
A token budget rather than named rungs: thinkingConfig.thinkingBudget, from 0 to the 32,768 ceiling, plus dynamic where the model chooses. Nothing had to be taken away here — a bare API call declares no tools at all, so there is no shell to script with and no way to fetch a reference. This is also the only leg that reports how many thinking tokens were actually spent, rather than leaving it to be inferred from the clock.
The longest ladder here: reasoning_effort takes all seven of none, minimal, low, medium, high, xhigh, max, and none is a true off switch — both models return exactly zero thinking tokens for it. No tools are declared, so the markup is composed from recall. One caveat: this dial is real but noisy. Four runs at low spent 2,246 to 7,543 thinking tokens and four at max spent 7,344 to 14,399, so the medians differ threefold but a single cell can land out of order. Read the row as a trend, not cell by cell.
Five rungs via reasoning_effort; the API refuses both none and max. Rows here are versions, not size classes — xAI ships no Pro/Flash equivalent, so “latest per size” is just Grok 4.7, and the two previous versions are shown for contrast. One tool is declared, write_file, and nothing else — no shell, no network — so the model still composes the markup itself from memory. Grok 4.7 needs it: without a file tool it reads “save it as last_supper.svg” as something it has already done and answers with one sentence and no drawing, fifteen times running. Reasoning tokens are reported directly.
Reading a cell. Each panel shows the file a model produced, at identical size across the whole study. Attempts above one mean an earlier try left no usable file; the prompt was re-sent unchanged. Not available means the provider refuses that setting outright — a fact about the model, not a failure to draw. Think is the reasoning tokens actually spent, where the provider reports them.
Composed or generated. In operation 04 the badge gives the verdict — computed, mixed or transcribed — from how much of the finished file was already sitting in the program as literal markup. A transcription hand-writes the SVG and wraps two lines of Python round it; that runs, but it is not a generator. Nothing in the prompt says which to do, so the choice is a result.
No reference, either operation. Every leg ran with the local MCP servers removed and no way to see the painting: the network, shell and file-reading tools are denied on Claude Code, the shell and image tools are switched off on Codex, and the three API legs declare no tools at all. Codex's web tool is served by the model's own harness, so it is closed by instruction and every cell is audited afterwards against its transcript. In operation 04 each program additionally runs in an empty directory with sockets and subprocess disabled, so it can neither fetch a reference nor read the drawings this repository already holds.