The SEO Audit That Approved a Page With No Headline, and How We Caught It

The SEO Audit That Approved a Page With No Headline, and How We Caught It

Posted by:

|

On:

|

The moment it broke

We built an automated auditor that must decide: block or pass, before a page goes live.

We thought the checklist and a single score would be enough.

It wasn’t.

In one run an approving reviewer marked the skill with top marks on form. The score was high. The verdict was pass.

Then the mandatory third checker ran and flagged twenty logic errors. Twenty. More than any other skill in that batch.

What we tried and what changed

Start with what the writer saw. They never saw the original sources. They saw a short brief distilled from three external search-advice texts and our house rules. That made word-for-word copying impossible by design.

We also required receipts: every check had to ship with evidence and a stated weight, and the system would publish a checksum so anyone could recompute the verdict from the article alone. That was in from the start.

Before version 1.1 the auditor re-used its own earlier measurements. It accepted the evidence that had been attached without replaying the measurement against the raw input. That mattered.

Here’s the concrete failure. An article reached an 80 percent score and passed. But it had no main heading and the target keyword did not appear in the opening. The checklist should have blocked that. It didn’t because there was no critical gate — the threshold alone allowed the pass.

That was the turning point. The third checker didn’t just annotate; they re-ran checks against the raw input and found twenty logic errors the approver missed. Every one could be traced to either a mis-measured field or a missing hard stop.

We made two concrete changes the same day. First: every verdict now re-measures its own evidence directly from the raw input. Every quoted field must match character for character. One mismatch kills the run; there’s no rescue by score. Second: we added a critical gate so that certain failures block publication regardless of the percentage score.

After those edits the same checklist could no longer be tricked into approving a page missing its main heading or lacking the keyword at the top. The audit now ships with a full receipts table plus the checksum described so a stranger can recompute every check from the article alone. And version 1.1 closed all twenty logic errors on the same day it found them.

Why that worked

We learned a simple fact: measured evidence that isn’t replayed is untrustworthy. If you accept supplied measurements, you inherit prior mistakes. Replaying measurements against the raw input exposed mismatches we otherwise missed.

Second fact: a single aggregate score hides critical failure modes. An article can hit a high percentage yet fail a vital requirement. A critical gate makes that failure visible and non-negotiable.

Third fact: transparency plus recomputability deters sloppy evidence. When every check must come with traceable evidence and an exact checksum anyone can rebuild, reviewers tighten up. The receipts are not theater; they are an accountability mechanism.

A counter-example we saw

Not everything broke before the fix. Some skills in the same batch had sensible checks and genuine metrics, and the third checker found few or no problems. The failure was not universal. It clustered where we depended on supplied measurements and had no hard stop for essential requirements.

How we know

We can point to the audit records. The writer’s input was the distilled brief, verified. The pre-1.1 runs show approvals with high scores but missing headings and missing opening keywords. The third checker’s report lists twenty discrete logic errors; every one maps to a specific check and raw-input mismatch. Version 1.1 edits close those items and add the critical gate. The receipts and checksum format are recorded so an independent party can recompute the result from the published article alone.

Questions we got

Why force character-for-character matching?

Because fuzzy acceptance hides mistakes. We tried loose matching and it let small but consequential divergences slip through; exact matching exposed them immediately.

Does a critical gate reject too many useful pages?

It can, if mis-specified. That’s why the gate is reserved for essential on-page requirements only — headings, required disclosures, core metadata. Non-essential items remain in the scored checklist.

Can the receipts be faked?

Not easily. The receipts plus a stated checksum construction let anyone recompute the checks from the article alone; that recomputability is the deterrent. If a receipt doesn’t rebuild, the checksum fails and the evidence collapses.

The runnable skill file lives in the members’ library.

Rule: always re-measure your own evidence and block on essential failures before you count percentages.

Sources: production audit records and the version 1.1 skill file updates from our content pipeline.


How we know

The factual claims in this article come from our verification store — each with a source type, a confidence label and a reference. The method is documented on How we know.

– Our production line wrote this audit skill from a curated brief distilled from three third-party search-advice sources; the writer only ever saw the distilled patterns, never the source texts, so word-for-word copying was impossible by design. | source: documented source | conf: verified | ref: EBZ skill-fabrikk kjøring 2026-07-24 + skillfil v1.1
– Before any verdict, the audit re-measures its own evidence directly from the raw input: every quoted field must match character for character, and a single mismatch blocks publication with no way for a high score to rescue the run. | source: documented source | conf: verified | ref: EBZ skill-fabrikk kjøring 2026-07-24 + skillfil v1.1
– It is assumed a scored checklist keeps bad pages out. Version 1.0 proved otherwise: with a plain 80 percent threshold and no critical gate, an article missing its main heading and with no keyword in the opening could still be approved for publishing. | source: documented source | conf: mythbuster | ref: EBZ skill-fabrikk kjøring 2026-07-24 + skillfil v1.1
– Each verdict ships with receipts a stranger can recompute: a table of every check with evidence and weights, plus a checksum whose exact construction rules are stated so the same hash can be rebuilt independently from the article alone. | source: documented source | conf: verified | ref: EBZ skill-fabrikk kjøring 2026-07-24 + skillfil v1.1
– A mandatory third checker found 20 logic errors in this audit skill that the approving reviewer had missed, the most of any skill in the batch of three, and every one of them was closed in version 1.1 on the same day. | source: documented source | conf: verified | ref: EBZ skill-fabrikk kjøring 2026-07-24 + skillfil v1.1

Leave a Reply

Your email address will not be published. Required fields are marked *