Inside Audiala's Editorial Process: How a One-Person Team Publishes 170,000 Verified Pages

We're transparent about being AI-assisted. Here's exactly what that means: the research, the grounding contract, the gates, and what happened when we audited our own judges.

Inside Audiala's Editorial Process: How a One-Person Team Publishes 170,000 Verified Pages

Audiala publishes audio guides and written guides for more than 170,000 pages across 11 languages — researched, written, translated, and maintained by a very small team with a lot of automation. We've always said we're AI-assisted and transparent about it. This post is the long version of that sentence: what the pipeline actually does, where humans sit in it, and how we check the system itself — including the time we fact-checked our own fact-checker and learned something uncomfortable.

Step 1: research before writing — always

No guide starts with a model "knowing things." Every guide starts with a fact packet: structured research assembled from Wikidata and Wikipedia (dates, coordinates, classifications), official tourism and municipal sources (hours, access, seasonal details), and targeted web research for recent developments. Sources are recorded with the packet, so every claim in a finished guide can be traced back to where it came from.

This matters because of the single most important rule in our pipeline:

Step 2: the grounding contract

The writing model's job is to turn the fact packet into readable, warm, narratable prose — using only what's in the packet. We call this the grounding contract. The model is not asked to contribute its own knowledge, however good that knowledge might be, because a claim we can't trace is a claim we can't verify, and our product is read aloud into people's ears at the actual monument.

We enforce the contract with automated gates: schema validation (every section present and well-formed), fact-check gates that compare generated claims against the packet they were drawn from, confidence tagging for single-source information, and contradiction flags when sources disagree. Content with unresolved claims doesn't ship. Our narrative story formats go further still — every factual claim is classified as sourced or explicitly framed as dramatization.

Step 3: we pick writers with blinded trials, not vibes

Which AI model writes the guides is not a matter of taste. Candidate models face the incumbent in blinded head-to-head duels on identical fact packets, scored by two independent AI judges on factuality and style separately — because we've measured that they move in opposite directions: the prettiest prose tends to carry the most ungrounded content. A model is promoted only on a clean sweep under both judges, plus better cost, speed, and format reliability. Our current writer won 35 of 35 duels against the previous one while being five times cheaper — the full numbers are public in our small-models deep dive.

Step 4: we audit the auditors

Here's the uncomfortable part, and the reason we trust this system: this August we went back and fact-checked 406 claims that our own judges had flagged as fabrications, verifying each one against the web with sources. It turned out 81–94% of them were true — real knowledge the models added that simply wasn't in the packet. Our judges were measuring contract compliance, not truth, and we had been calling that "hallucination."

Two things followed. First, we corrected our own published claims — the write-up above carries the revised numbers and a gallery of real failures. Second, the audit strengthened the system: on verified-false claims per 1,000 words — actual errors, checked against the world — our chosen writer scored 3–6× better than the alternatives, and we found that fabrication risk tracks neither model price nor size. We keep the grounding contract because traceability is the product; we now also verify what the metric itself means. Every eval logs its packets, outputs, and verdicts precisely so it can be re-litigated later.

Step 5: humans, where humans matter

We don't claim a human reads every one of 170,000 pages line by line. We claim something more honest: every page passes the same gates, and automation routes anything doubtful to human eyes — low-confidence claims, source contradictions, flagged translations, and every reader report. Editors spot-check new guides, review the flags, and own the editorial guidelines the models follow. The founder personally reviews the evals that decide which models get hired and fired.

Step 6: staying current

Tourist information rots. Opening hours shift, monuments close for restoration, prices change. We run a tiered freshness strategy: safety-relevant changes are updated as soon as we're aware of them; popular destinations sit on recurring refresh cycles; we monitor upstream sources and flag guides whose underlying facts changed; and guides display their last-reviewed date so you can judge freshness yourself.

And the most valuable signal is still you. If something in a guide is wrong or stale, tell us via the support page — reader reports go straight into the review queue.

The commitment

Being AI-assisted at our scale isn't something to hide; it's something to engineer honestly. The pipeline above is what lets two hands publish a library this size without publishing fiction: research first, a strict grounding contract, blinded trials for every writer, audits of the auditors, humans on the judgment calls, and receipts at every step. When we improve the process — or catch it being wrong — we'll keep writing it up here.

editorialtransparencyai-systemsquality
Hear it in the city

Walk it with Audiala.

Audio guides that follow your feet. Your first city is free — no card asked.