I asked it to judge itself.
Same question. Same four models. Three execution modes. Here is what happened.
The first Council, running Standard mode, returned a concise verdict at 54% confidence using live tools without a document. The second, running Extended mode on sandbox infrastructure, surfaced the single sharpest insight in the entire experiment while spending its 6x budget on PDF ingestion and research tool calls. The third, running Graph mode, decomposed the question into three independent claims, cross-examined each in structural isolation with native tools rather than explicit tool invocations, and reassembled them into the highest-confidence verdict at 73%.
The insight is not that Extended was cheaper than Graph. It was not. They both consumed 6x. The insight is that they spent those resources on fundamentally different things. Extended invested in external evidence. Graph invested in structural reasoning. Same budget. Different allocation. Different results.
Key Takeaways
- The Council's confidence rose with execution depth: 54% (Standard) to 69% (Extended) to 73% (Graph).
- Extended and Graph share the same 6x consumption ceiling. Extended grounds itself with explicit tool calls and a PDF; Graph decomposes into isolated claims with native tools.
- One gap remains: no cost-matched single-model baseline. The natural sequel.
The Decision HUD
The table below maps what you get, what each mode costs, and when to use it. The confidence scores are the Council's own assessment of how grounded each verdict is in the evidence. The token counts represent the total depth of deliberation across all rounds. The consumption column reflects actual debate execution, including tools and document surcharges.

The two deep modes are not tiered. They are routed. Extended and Graph share the same 6x consumption ceiling but spend it differently, and one spent it more effectively. Extended invested that compute into live research tools and PDF document ingestion to ground the debate in external evidence, reaching 69% confidence. Graph invested the same compute into structural decomposition: three independent claims, each deliberated in isolation, with tools baked into the model calls rather than issued as explicit invocations. It reached 73%. The gap was not tools versus no tools, both searched. The gap was linear deliberation versus structural isolation.
The Question
The prompt was self-referential: does a council of AI models produce more reliable decisions than a single frontier model, or does cross-model negotiation introduce failure modes that cancel out the diversity benefit?
The prompt was identical across all three modes: 184 characters asking whether a council of AI models produces more reliable decisions than a single frontier model. Extended alone additionally carried a 4,035-character research brief as a PDF, covering the diversity trumps ability theorem (Page, 2007), self-consistency methods (Wang et al., 2023), multi-agent debate (Du et al., 2023), LLM-as-judge biases (Zheng et al., 2023), and the Janis groupthink analogy. The four models (GPT-5.6 Terra, Claude Sonnet 4.6, Qwen 3 235B, and Gemini 3.6 Flash) ran in all three modes. The synthesiser, Claude Opus 5, produced every verdict independently.
The execution mode was the primary variable, though Extended also carried the research brief document.
Standard: Two Rounds, Honest About Its Limits
Standard mode runs on time-governed infrastructure with a 300-second ceiling. No document. Live tools enabled. Two rounds of deliberation at 3x consumption and 13,024 tokens.
The opening statements were distinct, not convergent. Claude Sonnet 4.6 foregrounded the overruled-minority problem, citing a "Council Mode" taxonomy that attributed 41 percent of incorrect responses to a correct minority being overruled, with a further 26 percent synthesis-override rate and a 19 percent correlated-hallucination rate. Gemini 3.6 Flash foregrounded feedback loops and cascade dynamics. Qwen 3 235B foregrounded coordination costs and conflicting priors. GPT-5.6 Terra framed it as an evidence problem: councils help only when independent, evidence-grounded checks outweigh persuasive error drift.
Round 2 was a single cross-examination pass. Each model critiqued a peer rather than the question. Claude Sonnet 4.6 surfaced the correlated-error statistic in this round, citing an ICML 2025 study finding that when two models were both wrong they gave the same wrong answer about 60 percent of the time. No model conceded ground, but none resolved the open questions either.
The verdict landed at 54 percent confidence, the lowest of the three. The Council was honest about what it could not measure: inter-model error correlation, a direct benchmark of heterogeneous councils against a single frontier model, and the magnitude of synthesis bias. None were resolved. The Council flagged its own blind spots rather than fabricate certainty.
Standard delivers a usable verdict at 3x consumption with tools enabled. Two rounds, 13,024 tokens, 54 percent. Use it for quick analysis where speed matters more than depth.
View the full Standard deliberation →
Extended: Three Rounds, Real Concessions
Extended mode runs on sandbox infrastructure with a one-hour execution window. Three rounds of deliberation with live tools and a 4,035-character research brief attached as a PDF. 6x consumption, 39,816 tokens. The extra cost came from document parsing and tool calls, not from additional rounds.
The opening round split rather than converged. Gemini 3.6 Flash and Claude Sonnet 4.6 took the pessimistic side: Gemini on feedback loops and correlated pretraining, Sonnet anchoring the 60 percent convergence statistic. GPT-5.6 Terra and Qwen 3 235B were conditional: councils help only when errors are meaningfully diverse and aggregation preserves evidence.
Round 2 was the stress-test round. GPT-5.6 Terra identified that Sonnet had miscited Du et al. as supporting correlation harms, when that paper actually reports debate gains. Gemini proposed the experiment's most concrete counterexample: a multi-agent software engineering pipeline with a deterministic compiler and test suite, where programmatic verification eliminates judge bias and debate drift.
Round 3 produced the insight that defines Extended mode's value. GPT-5.6 Terra refined its position to a precise operational statement: the reliability gain from councils localises in candidate generation, not selection. Diverse models produce a wider set of candidate answers. LLM-mediated consensus is a weak selector, inheriting position and verbosity bias. Where an external verifier exists, the diversity benefit is realised in full; where it does not, cross-model agreement is weak evidence.
The concessions this time were real and documented. Gemini 3.6 Flash restricted its degradation claim to unstructured debate with an LLM judge. Claude Sonnet 4.6 conceded the Du et al. citation error outright. Qwen 3 235B retracted its overstatement of the correlation gap. The verdict synthesised these refinements at 69 percent confidence.
Extended consumes 6x, the same as Graph. The difference is where the compute went: document ingestion and explicit tool calls grounded the debate in external evidence. If your question comes with a research brief or requires citation-chasing, Extended is the mode.
View the full Extended deliberation →
Graph mode does not run a single deliberation. It decomposes the prompt into independent claims, assigns a dedicated deliberation node to each, and reassembles the results through a final synthesis pass. The decomposer, also Claude Opus 5, extracts structurally distinct claims and routes each through its own three-round deliberation. The reasoning happens in structural silos, then recombines at the verdict. This experiment used no document. Tools are enabled by default. 6x consumption spent on structural reasoning.
The decomposer extracted three independent claims from the prompt. Each claim got its own deliberation node with the same four-model Council. The nodes ran in parallel on sandbox infrastructure, each with the same time ceiling.
The structure forced an analytical discipline that the other modes do not. A model could not pivot mid-round to a different claim. It could not bundle independent points into a single rhetorical position. Every claim survived or failed on its own reasoning.
The verdict surfaced a refinement absent from both Standard and Extended: the pathologies attributed to multi-agent systems have single-model analogues. Anchoring maps to prompt sensitivity. Groupthink maps to mode collapse. Confidence cascades map to autoregressive overconfidence. The failure modes exist in both architectures. The Council just makes them visible by structuring disagreement rather than burying it in a single decode.
The surviving position: councils are mechanistically distinct from single-model reasoning but not categorically novel. They amplify what structured reasoning already achieves, error visibility, not something qualitatively different.
Confidence reached 73 percent, the highest of the three. Not because Graph mode eliminates uncertainty, but because the decomposer forces the Council to surface it. Three independent claims, each verified or refuted on its own evidence, produce a verdict where no single model's error can cascade into false consensus.
Graph consumes 6x, the same as Extended. It spends that budget on structural decomposition. Tools are enabled by default; this experiment did not use a PDF. Three independent claim nodes, each cross-examined in isolation. Same consumption. Different allocation. If your question decomposes into independently-testable claims, Graph is the mode.
View the full Graph deliberation →
The three topologies produce structurally different reasoning from the same input. Standard is a linear pipe. Extended deepens it with document ingestion and live tool use. Graph decomposes into independent claims and reassembles. Extended and Graph share the same consumption ceiling. They reach it by different paths.

How the Same Model Changed Across Three Architectures
Claude Sonnet 4.6 held the strongest pessimistic position across all three modes. In Standard, it opened with the overruled-minority problem and surfaced the 60 percent correlated-error statistic during cross-examination. In Extended, it conceded the Du et al. citation error in the rebuttal round. In Graph, it contributed the mechanistically-distinct-but-not-categorically-novel refinement.
The arc is instructive. The same model, the same prompt, the same priors. What changed was how deeply the argument was tested.
Standard tested Sonnet's position once, in a single cross-examination pass. Extended tested it across two adversarial rounds, enough to force a genuine concession on the Du et al. citation. Graph isolated it per claim, forcing Sonnet to defend the independence assumption separately for every sub-claim rather than bundling them into a single rhetorical position.
The model did not change. The architecture did.
What This Means
The Council is not a black box that produces "better" or "worse" verdicts. It is an instrument. The mode you pick controls the depth of reasoning, the resource allocation, and the confidence of the output.
But the framing shifted during this analysis. Extended and Graph were initially assumed to sit on a pricing ladder: Standard cheap, Extended mid, Graph expensive. The execution traces told a different story. Extended and Graph run at the same 6x ceiling. They just spend it differently: Extended on explicit tool calls and a PDF, Graph on structural decomposition with native tools. And structure scored higher. Four points more confidence at the same cost. The difference was not whether the Council searched, both did. It was whether it reasoned about one question or three isolated claims.
Three modes. One question. The answer did not change. The Council is reliable under specific, verifiable conditions. What changed was how thoroughly the Council could prove it, and the finding that structural isolation scored higher than a single linear deliberation at identical cost, even though both modes ran web search.
One gap remains unclosed across all three modes. The experiment compared the Council against a single model, but not against a single model with matched computational budget. A cost-matched best-of-N single model (one frontier model running three to six independent samples with strict aggregation) was identified as the missing baseline in the Standard deliberation and never addressed across any mode. The Council catches its own errors, but it does not yet prove it catches more errors per inference hour than a well-tuned single-model alternative with equivalent resources. That experiment is the natural sequel.
Provenance
- Debates referenced: Standard, Extended, Graph
- Specifications referenced: Model Council
- Council roster, cost multipliers, and confidence scores verified against types.ts:
src/lib/types.ts - Consumption and tool audit verified against Supabase debate_events table:
tool_start, tool_result, and has_document columns - Articles referenced: Log #5: Three Rounds, One Verdict