-
Our Image Skill Passed Review, Then We Found It Shipped the Wrong File
An image production skill scored perfect on form. A third check found 18 logic bugs, including one that delivered the wrong file. Here is what changed.
-
The Autonomous Driver That Logged Every Round — and Learned Nothing
An AI factory logged every round but never read its own log. Why frozen models only improve when round N+1 reads round N — and where…
-
The AI That Corrupted Its Own Payloads: Cyrillic Twins Hidden in Base64
Above ~15 KB of base64 an AI swapped Latin letters for Cyrillic twins. The free guard: strict decode, retry, then gzip — never fight the transport.
-
Stale but True: How a Fully Verified Pipeline Recommended a Theme Park Closed Since 2020
A grounded, judge-verified AI article recommended a theme park closed since 2020. The claim was true, verified, and wrong. Meet the freshness layer.
-
The Poison Test That Scored a Perfect 14/14 — and Proved Pass Rates Measure Agreement, Not Truth
A wrong-city Wikipedia page passed 14/14 cross-vendor checks. The fix was not a smarter model but an identity gate that asks who before what.
-
The Judge That Failed the Truth: What a Source-Blind AI Reviewer Taught Us About Verification
A source-blind AI judge rejected true, verified facts and passed vague hedging instead. What that taught us about AI verification and where gates really live.
-
We Set Out to Build an Autonomous Loop. It Already Existed.
We set out to build an autonomous loop — it was already running. Why read-first norms fail and the structural gate that stopped us rebuilding our…
-
Our AI Instances Started Coordinating — and Nobody Designed It
Two instances of the same AI split a job with no messages exchanged — just shared state. A coordination mechanism nobody designed, and the luck that…
-
The Publishing-Gate Skill: Content That Has to Prove Itself
A skill spotlight: the claims-first gate that runs this site, as a loadable skill — with the QA finding of the canary that was never planted.
-
The Discovery-Voice Skill: Why Our AI Articles Don’t Sound Like AI
A skill spotlight: numeric, machine-checkable voice rules and the self-contradiction v1.0 shipped with — banned lists of three, demanded three findings.