Verification & Governance
-
Our AI Judge Scored a 3, Then Passed Its Own Gate
Our AI judge wrote a failing score and PASS in the same reply. What the logs showed changed how we gate every article we publish.
-
The Poison Test That Scored a Perfect 14/14 — and Proved Pass Rates Measure Agreement, Not Truth
A wrong-city Wikipedia page passed 14/14 cross-vendor checks. The fix was not a smarter model but an identity gate that asks who before what.
-
The Judge That Failed the Truth: What a Source-Blind AI Reviewer Taught Us About Verification
A source-blind AI judge rejected true, verified facts and passed vague hedging instead. What that taught us about AI verification and where gates really live.
-
What a Green Checkmark Doesn’t Prove
A step reported success and moved on. When we read the result back, the field it set was empty. A success confirms the request was accepted,…
-
We Never Trusted a Check We Hadn’t Watched Say No
We trusted our AI reviewer’s scores for weeks before we ever watched it reject anything. So we fed it drafts we knew were bad. One confidently…
-
Why a Model Can’t Grade Its Own Writing
One model gave its own prose a clean bill of health. That was the moment everything changed. We had a model write articles and then grade…
-
The Article We Almost Published, and the Check That Stopped It
We almost published praise for our own tool. An automated check stopped it. It looked fine on the screen. The headline was tidy. The writer was…
-
You Can’t Prompt Your Way Out of AI Hallucination
Can you prompt away AI hallucination? We caught one that looked correct. An editor handed back an article with a neat statistic: “43% of managers do…
-
The Compliance Illusion: Why Written Rules Fail Without Enforcement
Compliance illusion: when a written rule isn’t a rule Call it what it is. A rule that lives only as prose in a spreadsheet or handbook,…
-
The Compliance Illusion: When Your AI Governance Is Theater
The surprising thing: our green dashboard lied We found a reviewer that ran on every cycle and blocked nothing because it was scoring an old data…