The Rubber Stamp: Multi-Model Review Is Not a Second Opinion
Claude writes, Codex reviews. The industry is racing toward cross-vendor review. Five papers converge on one warning: a second model is not a second opinion.
Read →6 entries filed under Operations.
Claude writes, Codex reviews. The industry is racing toward cross-vendor review. Five papers converge on one warning: a second model is not a second opinion.
Read →Three Log articles named the same gap and left it hanging. We ran the test. Both found all three failures. The council won on what the scorecard missed.
Read →The Model Council shows confidence scores to one decimal place. 72.3%. 88.7%. 94.1%. The precision is compelling. The precision is also meaningless. The models are self-assessing. The orchestrator's aggregate confidence is generated by the same model that produced the synthesis. A percentage score generated by a language model is not a measurement. It is a number with a decimal point and no calibration.
Read →Four models agreed on every claim. The synthesis caught it. They had all pulled from the same training distribution and the same search snippet. Unanimity was not adversarial victory. It was statistical correlation wearing four badges.
Read →The split-based fenced code block approach was added as a Windows tool corruption workaround. It became the rendering bug that broke the Log for two sessions. The fix wasn't changing the code. It was changing the order.
Read →Promise.race timeouts, FK cascade failures, and the Supabase .catch() that silently dropped 57 of 63 deliberation records over six weeks.
Read →