That one stung
We write travel guides about Spanish coastal towns. Short sentences first. Then the long, embarrassing one.
On 2026-07-23 our Benalmádena guide recommended Tivoli World as an attraction worth visiting. The park has been closed since 2020. The line came from a respected travel wiki and had passed our two reviewers before publication. Both reviewers verified the stored fact: the park exists. They did their job. They did not check whether it was open.
It hit us when two independent, outside audits ran the numbers on the same pages and scored Nerja 82 and Benalmádena 78 on 2026-07-23/24. Those scores forced us to look again. The auditors flagged the Tivoli line. That was the turning point. The outside scores, not anything internal, forced the rebuild.
The experiment that made us change
We tried repairing the problem in one go. One big AI call: assemble facts, structure ten sections, build tables, write Q&As, add links, and check freshness. It failed. Hard.
So we split the work into passes. First pass: assemble stored facts and mark every sentence that suggests someone should go somewhere. Second pass: run a live lookup for each marked sentence against an official or municipal source. Third pass: rewrite or retire the line if the lookup disagreed or had no confirmation. Fourth pass: human review focused only on the differences. The multi-pass version produced the same 2,700-word guides with tables, Q&As and a freshness box—and it passed review. The single-call approach could not keep track of forty-plus facts plus live verification at once; splitting made each step checkable.
Moment of clarity: the stored fact was true, and that was our blind spot. Checking verifies truth, not current status. “Stale but true” is its own failure mode.
The rule that came out of it
Any sentence that implies “go visit X” or says “X is worth it” must first pass a live lookup against an official or municipal source before it may be used as a recommendation. A monthly audit now retires outdated statements, writes corrected ones and rebuilds the affected pages. We codified that on 2026-07-23/24. The tested, runnable version lives in the members’ library.
One small representative detail we kept for illustration: a sentence like “Visit Tivoli World for a day of fun” is flagged automatically for live verification; if the official site or municipal page shows the venue closed or has no current information, the line is retired or rewritten. That’s the taste—only a single line, not the full recipe.
Why this works
Encyclopedia-class sources are strong on timeless facts and weak on current status. We learned that with Tivoli World: the park exists (timeless). It stopped operating in 2020 (status). A stored fact and two reviewers confirmed existence and so passed, but no live check showed it was closed. Adding a live lookup closes that gap. The monthly sweep handles things that change slowly—closures, reopening, price changes—so we don’t carry stale recommendations for months.
Splitting into passes also made errors visible. When each pass has a clear, testable job, reviewers and auditors can point at the exact step that went wrong. That made fixes surgical instead of sprawling.
A counter-example we kept
We did not throw out every third-party source. For museum histories, architectural dates and long-standing facts, encyclopedia-style pages still win. Our change targets sentences that steer readers to act now: go, visit, book, try. For timeless context we still rely on respected sources; for present-tense recommendations we require live confirmation.
How we know
The evidence is in our logs. The Tivoli error is recorded and dated. The two outside audits and their scores are in the audit folder for 2026-07-23/24. The multi-pass rebuild produced passing reviews; the one-shot attempt did not. Those are verifiable events, not feelings.
Questions we got
Why did two reviewers miss that Tivoli was closed?
Because both reviewers verified that the stored fact was true—the park exists near Benalmádena. Our checks confirmed existence and provenance, not whether the site was open today. The reviewers did what we asked; the gap was in what we asked them to check.
Won’t live lookups slow production a lot?
They add time, yes. Splitting into passes avoids a single, fragile step that tries to do everything. The multi-pass approach we tested returned guides at scale and kept review time predictable. There is an operational cost; we accepted it because it prevented publishing bad recommendations.
Could this create false negatives—places wrongly retired?
Occasionally. Official pages can be poorly maintained. That’s why the live lookup is one pass among others, and why human review focuses on mismatches. We also keep a monthly sweep that rechecks retiring decisions; if an official page reappears or a venue reopens, we rebuild the page. It’s not perfect, but it’s better than telling readers to visit closed places.
Truth can be true and useless. A true sentence can still be a bad recommendation.
Sources: internal logs and audit records dated 2026-07-23/24 (Benalmádena/Tivoli World, outside audits, and the freshness-rule entry).
How we know
The factual claims in this article come from our verification store — each with a source type, a confidence label and a reference. The method is documented on How we know.
– The fix also broke a one-shot habit, proven on 2026-07-24: a single AI call could not use forty-plus facts, structure ten sections, build tables, write fifteen Q&As, add links and check freshness all at once. Split into separate verifiable passes, the same production line now delivers roughly 2,700-word guides with tables, Q&As and a freshness box — and passes review. | source: documented source | conf: verified | ref: our living-lab log + seksjon ‘our freshness law’ Bevis 2 (fler-pass-loven), 2026-07-23/24
– Encyclopedia-class sources turned out strong on timeless facts and weak on current status: neither Wikipedia nor the travel wiki carried any reliable signal about whether Tivoli World was open, what it cost, or its opening hours. The incident is written up in our project log as the reference case behind the freshness rule. | source: documented source | conf: observed | ref: our living-lab log + seksjon ‘our freshness law’ Bevis 1, 2026-07-23/24
– What actually triggered the overhaul wasn’t our own gate. Two outside AI-driven SEO audits, run as independent second opinions on 2026-07-23/24, scored our Nerja guide 82 out of 100 and Benalmádena 78 — and it was those outside scores, not anything internal, that forced the deep-plus-fresh rebuild of the article factory. | source: documented source | conf: sourced | ref: our living-lab log + seksjon ‘our freshness law’ parter/oppgave, 2026-07-23/24
– On 2026-07-23 our travel guide for Benalmádena, a beach town in southern Spain, recommended Tivoli World as an attraction worth visiting. The theme park has been closed since 2020. The fact came from a respected travel wiki, and it had passed checking by two independent AI reviewers before publication. | source: documented source | conf: verified | ref: our living-lab log + seksjon ‘our freshness law’ Bevis 1, 2026-07-23/24
– The assumption that a sourced, double-checked fact is safe to publish fell on 2026-07-23: the Tivoli World line was technically true — the park does exist near Benalmádena — so both reviewers passed it. Checking verifies truth, not current status. ‘Stale but true’ is a failure class a verification gate cannot see. | source: documented source | conf: mythbuster | ref: our living-lab log + seksjon ‘our freshness law’ Bevis 1, 2026-07-23/24

Leave a Reply