-
The Poison Test That Scored a Perfect 14/14 — and Proved Pass Rates Measure Agreement, Not Truth
A wrong-city Wikipedia page passed 14/14 cross-vendor checks. The fix was not a smarter model but an identity gate that asks who before what.
-
The Judge That Failed the Truth: What a Source-Blind AI Reviewer Taught Us About Verification
A source-blind AI judge rejected true, verified facts and passed vague hedging instead. What that taught us about AI verification and where gates really live.
-
We Set Out to Build an Autonomous Loop. It Already Existed.
We set out to build an autonomous loop — it was already running. Why read-first norms fail and the structural gate that stopped us rebuilding our…
-
Our AI Instances Started Coordinating — and Nobody Designed It
Two instances of the same AI split a job with no messages exchanged — just shared state. A coordination mechanism nobody designed, and the luck that…
-
The Publishing-Gate Skill: Content That Has to Prove Itself
A skill spotlight: the claims-first gate that runs this site, as a loadable skill — with the QA finding of the canary that was never planted.
-
The Discovery-Voice Skill: Why Our AI Articles Don’t Sound Like AI
A skill spotlight: numeric, machine-checkable voice rules and the self-contradiction v1.0 shipped with — banned lists of three, demanded three findings.
-
The Exact-Math Skill: Making an AI’s Algebra Provable
A skill spotlight: symbolic computation with a receipt — 200-bit cross-checks, class-specific verification, and the QA story of the flaw that would have failed correct answers.
-
The Project-Planner Skill: Making an AI Break Work Down Like a Project Manager
A skill spotlight: the MIT-licensed project-planner procedure that makes an agent size tasks at 2-8 hours, estimate three ways, and put a risk table on everything.
-
The Audit Trail Is the New Backlink
We built verification to keep our content honest. Then we noticed the proof layer was doing the job backlinks used to do. A working hypothesis, honestly…
-
One Prompt, One Website: What the Orchestration Actually Takes
The blueprint behind letting one prompt become a published page: six artifacts per run, gates that stop deploys, and the honest gap between design and running…