AI Content ROI: 270 Guides, 19,000 Credits, 4 Clicks

Posted by:

|

On:

|

Ninety days after the last of 270 guides went live, we opened the search data expecting a curve and got a single number instead: four clicks. Roughly 4,400 impressions, average position near 40, four clicks. The run had taken twenty days and about 19,000 credits at roughly 70 credits per guide, and that was the whole measured AI content ROI of it.

Nothing was flashing red. Every row in the queue had reached PUBLISHED. Every judge gate had returned PASS. Every scorecard was green. The board we had built to tell us whether the factory was working said, correctly, that the factory was working.

What we checked first, and why none of it found anything

We assumed something had broken quietly, because that is usually the answer: thin output, near-duplicate pages, an indexing problem.

The duplication theory died first. We compared 354 posts pairwise on eight-gram overlap; every pair came back below 0.10. The pages were not variations of each other. Then we did the arithmetic we should have done at the start. At average position 40, a click-through rate of 0.1 to 0.2 percent on 4,400 impressions predicts four to nine clicks. We got four.

That is the uncomfortable part. Four clicks was not a symptom of failure. It was the pages earning precisely what their position entitled them to. There was no bug to bisect and no gate to tighten. The machine had done exactly what we asked of it, 270 times.

The defect sat upstream of everything the gates could see

Province capitals and towns of under a thousand people had entered the queue at identical priority. A regional hub with real search demand and a village nobody looks up were the same kind of row to the pipeline, which had no notion of kinds of rows. It had a notion of valid rows.

That is what every gate we had built measured. Is this claim true, is this structure well-formed, is this text original, did this publish. All correctness gates — and correctness gates answer a different question from value gates. Any number of the first provides no coverage of the second. A green board proves the absence of detected error, not the presence of worth.

Nobody had estimated, before building, what a single row could be worth if the build went perfectly. Not the expected value — the best case. That number exists to identify rows whose best case is still not worth having, and those rows are invisible to a system that only asks whether output is correct.

AI content ROI is decided before generation, not measured after it

The comparison that would have saved the run could only have happened at one moment: before any of the 270 existed, with the full candidate set in front of us. A pipeline that discovers work one item at a time cannot compare items, and comparison is the entire mechanism. Once a row is built its cost is spent, and the only question left is whether it came out correct — which it did, 270 times.

It also explains why better writing would not have rescued the AI content ROI here. The guides were good; quality was never the variable. Sharper ones would have moved position 40 to perhaps 35, on rows whose ceiling was low either way.

The gate fired on its own author the first time we ran it

When we turned this into an actual pre-build gate and pointed it at our own queue, the first version passed 191 of 197 candidates. We had scored each candidate on the number of consolidated sources behind it, and the spec caps an original at eight sources. Everything with eight or more tied at the top. The score stopped discriminating exactly where the interesting candidates were.

A floor that nothing fails is not a floor. It was a budget wearing a threshold’s clothes. The repair was not a higher number but a ceiling that measures something which actually varies across candidates instead of saturating. The rerun rejected 141 of 186.

What this means if you are about to run a batch

Three things carry into any repeated build — a bulk publish, a migration, a generation run, anything where one procedure repeats across rows.

  • Verify negligible unit cost against the total, not the unit. Seventy credits is nothing; 270 rows of it is 19,000.
  • Record the estimate before the build even when you distrust it, because an item with no prior estimate can never afterwards be judged as anything except complete.
  • Publish what you declined next to what you built. A silent cut reads later as though the whole candidate set was covered, which is the same error one level up.

Keep the honesty about your own counterfactual too. In our grounding run the roughly 200 suppressed rows were calculated, never built, and never confirmed to have failed, so every claim about the AI content ROI we saved is one unrun experiment away from being an assumption.

The runnable procedure — how to enumerate, what unit to estimate in, how to derive a floor that holds — lives in the member library.

How we know

Grounded in: our own content factory run — 270 guides published across twenty days, audited 14–15 August 2026, and the resulting gate run live on our skill-originals queue on 17 August 2026. Verified: ninety days of search data read back at roughly 4,400 impressions, four clicks and average position near 40; 354 posts compared pairwise on eight-gram overlap with every pair below 0.10; and 191 of 197 candidates passing the first floor before it was rebuilt and rejected 141 of 186. The runnable procedure lives in the member library.

Leave a Reply

Your email address will not be published. Required fields are marked *