An auditable AI fact-check of an exhibition wall text

Published Sep 24, 2026Reviewed Oct 7, 2026

A pipeline of AI research agents that fact-checked a 1990–2000 chronology written for a contemporary-art exhibition, one verdict and source trail per paragraph, with the curator's original file left untouched.

At a glance

Problem
A contemporary-art curator needed every factual claim in an exhibition wall text, a year-by-year chronology of world events from 1990 to 2000, checked against sources before it went on the wall.
Constraint
The original Word file could not change, parts of the text were deliberately blacked out, and a claim without a source had to be reported as unverified, never as false.
What I built
Four AI research passes, one per period, wrote a status, note and source list for every paragraph into JSON coverage logs, and one build script generated every deliverable from those logs.
Result
All 441 non-empty paragraphs came back about two hours after the request arrived: 20 marked to fix, 8 priority corrections on page one, 278 comments in a copy of the original, 542 distinct sources, and a hash check proving the original was untouched.

Overview

The request arrived by email in the early afternoon: a Word file and one line asking whether it could be run through a strong AI model for fact-checking.

The file was the wall text for a contemporary-art exhibition. It held 880 paragraphs, 441 of them non-empty, each a short line about one or two events between 1990 and 2000: elections, wars, court rulings, record releases, product launches. It was written as art, so some lines are deliberate interpretation, and some words are covered by black blocks as part of the work.

Asking a model to "fact-check this document" returns confident prose about some of the claims, with no record of which paragraphs it actually read. The goal was provable coverage instead: one verdict per paragraph, with its sources, in files the curator could open without new software.

The reply went out about two hours after the request arrived. The agent run itself, from downloading the attachment to sending the reply, took a little over half an hour.

Coverage before conclusions

The unit of work is the paragraph, identified by its position in the Word file, empty paragraphs included. That position becomes the paragraph ID, so a number in the report points to the same place in Word.

A coordinating agent split the text into four periods: 1990–1992, 1993–1995, 1996–1998 and 1999–2000. Three sub-agents researched the first three periods in parallel, and the coordinator took the last one itself. Each pass wrote one JSON record per paragraph: status, note and source URLs. The passes produced 135, 118, 112 and 76 records.

The 1990–1992 pass seeded every paragraph as unverified, with a note saying so, and upgraded a paragraph only after writing a specific finding for it. A paragraph the agent never got to therefore stays visibly unverified instead of silently passing.

The three sub-agents also named the ID field three different ways: paragraph, paragraph_id and id. The build script accepted all three, then refused to generate anything unless coverage was exact:

assert len(records) == len(texts) == 441
assert len({r['id'] for r in records}) == 441
assert set(texts) == {r['id'] for r in records}
assert all(r['status'] in labels for r in records)

No paragraph skipped, none reviewed twice, no status outside the vocabulary.

In a single long pass, a skipped paragraph looks the same as a paragraph with no problems. With the coverage log, "did it check everything" is an assertion that either passes or stops the build.

Six statuses, and "unverified" means unverified

A true/false verdict would have misrepresented most of the text. Each paragraph got one of six statuses:

StatusParagraphsMeaning
Fix20A specific factual error or a misleading claim
Refine210Partly confirmed; a date, figure or wording needs tightening
Confirmed130The visible factual content is supported
Interpretation37The author's reading of events, not a checkable fact
Unverified11Not enough source evidence either way
Redacted33Too little visible text to assess

"Unverified" means no adequate source was found in the time available, and the report, register and read-me all state that this is no evidence the claim is false. Blacked-out words were never reconstructed or guessed at; confirmation applies only to the visible text.

Every "fix" and every "confirmed" verdict cites at least one source.

The corrections were the kind a careful visitor would repeat. One line merged separate statistics from a truth commission's report into a single claim the commission never made. Another compressed a 78-day standoff that ran from July to September into "78 days in July". A third said "not one officer is hurt", and the investigating commission's own report mentions an injured commander.

Suggested English replacement lines were offered as editorial options and never inserted, because the surrounding text belongs to the author.

The original never changes

The curator's file is the reference everything else points into, so the pipeline treats it as read-only evidence. The build script hashes it with SHA-256 before generating anything and checks the hash again after the last file is written.

Review comments were anchored in a copy of the original with python-docx, on every non-empty run of the paragraph they refer to. Paragraphs marked "confirmed" or "redacted" got no comment, which leaves 278 comments. Each comment is signed as an AI fact-check and carries the paragraph ID, status, date, note and linked sources. After saving, the script reopens the copy and proves it adds notes and changes nothing:

check = Document(OUT / 'WALL-TEXT-su-komentarais.docx')
assert [p.text for p in check.paragraphs] == [p.text for p in original.paragraphs]
assert len(check.comments) == comment_count
assert hashlib.sha256(ORIGINAL.read_bytes()).hexdigest() == original_hash

The commented copy keeps every paragraph's text, and the original's hash is unchanged.

One log, every deliverable

Every document comes from the same coverage records through one Python script, and LibreOffice converted the generated DOCX report to PDF.

Sources found after a period's review was written did not go in as hand edits to an output file. They went into the build as explicit, named overrides: two edits inline, three in a dictionary of late updates, and one from a small JSON file read at build time. Six paragraphs changed this way, and the whole package can be rebuilt from the logs.

The curator received:

  • A report as PDF (19 pages) and editable DOCX, opening with the 8 priority corrections.
  • A copy of the original Word file with comments at the exact spots.
  • A register of all 441 paragraphs as CSV and as a single HTML file with text search and a status filter, which runs offline and sends nothing anywhere.
  • A short read-me explaining the statuses and that the work was done by AI.

The report is in the curator's language, and the replacement lines are in English to match the wall text. A JSON register and a metadata file with the original's hash, the paragraph count, the comment count and the status totals stay with the project as the audit record.

The register cites 542 distinct source URLs from 342 sites, with at least one source on 372 of the 441 paragraphs. Before sending, the coordinator rendered the PDF's first page to an image and checked it visually.

Limitations

  • The research was done by AI and every deliverable says so; the value is speed plus a complete, checkable trail, with no claim of infallibility.
  • "Refine" is a wide bucket of 210 paragraphs: it holds near-confirmations and claims that need real rewording, so its notes need reading one by one.
  • Eleven paragraphs remain unverified; the search stopped there, which says nothing about whether those claims are true.
  • The record schema was not fixed before the agents started, which is why three ID field names reached the build; I would hand every agent the same schema up front.
  • The hash is taken when the build script starts, so it covers the build only; the research steps ran before it.
  • Replacement lines are drafted in isolation, so the author has to fit them to the full text, especially next to redacted passages.

Related case studies

Start with one process

A one-hour call costs €80. Afterwards you get a written plan, whether or not we work together.