The Instrument: Why Your Agent Will Delete Your Tests
The handshake told the critic what to read. The instrument tells the critic what to check. Code cannot be talked out of a math check.
Read →Architecture decisions, failure modes, and operational lessons from building NiraNexus-OS. Updated as we ship.
What shipped, what broke, what we learned. The operational record of building NiraNexus-OS, written before hindsight made it safe.
The handshake told the critic what to read. The instrument tells the critic what to check. Code cannot be talked out of a math check.
Read →A critic cannot audit what it cannot read. The acceptance-criteria handshake is not a prompt. It is the architecture that makes the Consensus Lie visible instead of invisible.
Read →Claude writes, Codex reviews. The industry is racing toward cross-vendor review. Five papers converge on one warning: a second model is not a second opinion.
Read →Three Log articles named the same gap and left it hanging. We ran the test. Both found all three failures. The council won on what the scorecard missed.
Read →The Verification Illusions trilogy proved that badges, consensus, and confidence bars manufacture authority. This article covers what was built to stop them. Not a governance framework document. Not a boardroom decision. A bash script and three functions. Twenty-one mechanical checks. The ones earned by a fire stay earned forever.
Read →The Model Council shows confidence scores to one decimal place. 72.3%. 88.7%. 94.1%. The precision is compelling. The precision is also meaningless. The models are self-assessing. The orchestrator's aggregate confidence is generated by the same model that produced the synthesis. A percentage score generated by a language model is not a measurement. It is a number with a decimal point and no calibration.
Read →Four models agreed on every claim. The synthesis caught it. They had all pulled from the same training distribution and the same search snippet. Unanimity was not adversarial victory. It was statistical correlation wearing four badges.
Read →The response card said GPT-5.6 Terra. The Execution Metrics said FALLBACK_TRIGGERED. GPT-4o wrote that analysis. The swap was silent, invisible, and correct by design. This is the fallback lie. The name on the badge is the model you asked for. The model that actually ran is buried in a collapsed data panel most users never expand.
Read →Microsoft ran 18 controlled experiments. Coding agents with their own test suite scored 221 out of 222 while the reusable library they were hired to build was completely dead. The Model Council separates making from checking. That is the difference between verification and confirmation.
Read →Every architecture article about the Model Council assumes the roster exists. This one explains where the models come from. Four cognitive personas, four providers, a two-tier fallback chain, and a daily cron audit that checks sixteen endpoints before a debate can fail.
Read →Same question. Same four models. Three execution modes. The verdict moved from 54% to 73% as the reasoning topology deepened. Extended and Graph consume the same compute ceiling but spend it on fundamentally different things.
Read →The Model Council does not vote. It cross-examines. Tracing a real deliberation through three rounds of structured conflict: code blocks, claim lifecycle, and the dissent protocol that preserves disagreement as evidence.
Read →A circuit breaker is not a budget control. It is a quality signal. Three yield points in the debate engine catch broken processes before they consume tokens and produce nothing. The code is the policy.
Read →The split-based fenced code block approach was added as a Windows tool corruption workaround. It became the rendering bug that broke the Log for two sessions. The fix wasn't changing the code. It was changing the order.
Read →Promise.race timeouts, FK cascade failures, and the Supabase .catch() that silently dropped 57 of 63 deliberation records over six weeks.
Read →