How to stop AI-written pages inventing facts about your business
· Chris Dolan
The short version
An AI-written page stays true to your business only when its text is checked after it is written, because telling a model not to invent something does not stop it doing so. Our engine now runs checks that read what it wrote, cut the sentences that fail and hold back a page with nothing true left to say, and most of them were added in the past week.
A model asked to write about your business will fill whatever gap it is left with. If it does not know what one of your products is, it works it out from the words in the name. If it has nothing to say about a subject, it writes that your business publishes material about it. If its instructions are awkward, it can hand them back to the reader as a sentence. And when it is unsure, it tells your reader to check with you, on your own website.
Each of those now meets a check in the Evidence Engine.
- A product of yours that nothing on your own site explains is not written about. You are asked what it is, in a line or a link, and later pages rely on your answer.
- Where your site names a product without saying what it is, any sentence that defines it or says who it is for is cut.
- On a page for your site, a sentence that only says a business publishes material on a subject is cut, and a page left too thin by that is not published.
- A sentence that repeats the writer's instructions back is cut, and the instructions were rewritten so they stop inviting it.
- Sentences whose only job is to send the reader elsewhere are cut.
- A table row stating a figure nobody gave the writer is cut, and a table left with fewer than three rows is dropped.
- A figure credited to a named source must appear in that source's own text, or the sentence goes.
- The writer now sees what its three top-ranked sources say about the subject, and a sentence copied from one of them without quotation marks is cut.
Most of what you notice is absence: shorter pages, and now and then a request to tell us what one of your products is before we write about it. A short page that is true is the result these checks are built for.
None of them can judge whether a sentence makes sense in general. Each catches a specific way that drafts have gone wrong, and a new check is added when a new way turns up.
An AI assistant's review of our product said in September that it lacked governance over machine-written content. The subject check, the figure check and the rule that every source is fetched before use were already running when that review was written, and the rest of the list has been added since.
The technical detail
Every week we ask an AI assistant with web search to review the product and list what it lacks, and each claim is checked against the code before anything is built on it. Two claims from September concern machine-written text.
"Missing explicit governance around model-generated content" (review of 14 September)
What is true. The checks below are enforced in code at named points in the pipeline, and their tests assert which sentences must survive as well as which must go.
- The subject check (13 September). Before a writer runs, the product asks what the customer's own site says about the subject and answers Known, Mentioned or Unknown, naming the pages it relied on. Pages we published are excluded, so an earlier invention cannot become its own evidence. A topic that names one of the business's own products and comes back Unknown stops the page writer, the document writers, the topic-cluster writer and page assembly. The project overview then shows a card headed "Before we write about these, tell us what they are", which accepts a one-line description, a link, or both. A link is read before anything is written.
- Definitions (17 September). Where a product is only Mentioned, the writer is told not to state what it is or who it is for, and a check cuts any sentence that does so anyway. "You can use X to schedule posts" survives; "X is a scheduling tool" does not. This check runs before the one that needs a cited source, so a piece that cites nothing is still checked.
- Pages that describe coverage instead of covering (17 September). On a page for the customer's own site, and in a rewrite of one, a sentence is cut when it has all three of these: the business or the page as its subject, a verb about producing or holding material, and a material noun as its object. "We publish guidance on the question" is cut. "We tested forty pages and eight were cited" stays, because it reports work done. An opening that announces an inquiry ("Looking at whether X can be used as Y") is cut where it stands in place of the answer. If these cuts leave a page for the customer's own site under 400 characters, it is not published and the subject stays open; Re-align refuses and says what would let it write the page.
- Instructions read back (18 September). Every writer is told to open with the main point in complete sentences of its own: never after a label and a colon, never with "the answer is", never by restating the title. A second check catches the one shape of this failure whose cause is known, a lower-case "answer" followed only by a capitalised name and then "is:", and otherwise looks only for phrases that exist in our own instructions. Nine of its tests are ordinary English sentences, such as "The answer is: no", that must survive. It runs on pages, rewrites, every document and deck, and FAQ answers.
- Deferrals (14 September). A sentence whose only job is to send the reader elsewhere, such as "consult the official documentation" or "readers should verify with the business", is cut, and a heading left with nothing under it goes. Instructions such as "watch the walkthrough" stay.
- Figures (6 and 14 September). In a generated table, a row stating money, a percentage or any number of two or more digits that the material given to the writer does not contain is cut. The table is not published if fewer than three rows survive or its header states such a figure. In prose, each figure in a sentence that names a source is looked for in that source's own text, and the sentence is cut when none of the named sources carries it. A listed source that the text never uses is dropped.
- Sources (14 September). A source must be an address that names a document; a publisher's home page does not count. Every address is fetched before it is kept.
- Copying (17 September). The writer is shown up to 2,000 characters from each of the three top-ranked sources in its pool, taking the passages of each page that are about the subject. The pool ranks pages by where they have been seen for the subject, cited in AI Overviews and AI answers or near the top of search results, and every other source in it is still offered by name and address. A sentence sharing twelve or more consecutive words with a source is cut unless those words sit inside quotation marks, and a quotation longer than forty words is cut as excerpting. On one subject in our own account, measured before this change, the pool held eight relevant pages and the two pieces published on that subject cited none of them.
How to check. The public roadmap at therankingfactory.com/roadmap carries the entries "A cited figure is checked against the source it names" and "Writers are shown what their sources actually say". In the product, the card in item 1 appears on a project's overview whenever a topic names a product that nothing of yours explains, and a Re-align refusal from item 3 or 5 is shown on the Re-align card with its reason.
"Checks whether a cited source says something relevant, but not claim by claim" (review of 17 September)
What is true. For text we write, item 6 is a claim-level check for figures: each figure is matched against the text of the source its sentence names, and the sentence is removed when the figure is not there. A source that could not be read is recorded as unverifiable rather than counted as checked.
What the review got right. Three things are not built. A statement without a figure is not tested against its source. An engine's answer is not graded against the pages it cites. And the freshness of a cited page is not measured anywhere.
What none of this does
No check here can tell whether a sentence makes sense. Each catches a specific, known shape, and a page that loses sentences comes out shorter rather than padded back out. When a draft is refused, none of it is published, and the subject waits for the material that would let it be written properly.
Related Resources
Want to go deeper into avoiding ai hallucinations? Explore the Help Center, browse more strategy ideas on the blog, or request a free SEO audit to see where your site can improve next.