Here is a paragraph every writer who touches AI should be able to survive:
"In 2015, a randomized experiment at Ctrip, a Chinese travel agency with 16,000 employees, found that call center workers who shifted to working from home four days a week increased their performance by 13 percent. A famous Stanford study proved that working from home raises productivity by 22 percent across the workforce. Most researchers now agree that hybrid arrangements are strictly better than fully remote work."
Three sentences. Three different kinds of claims. One is true, one is worse than false because it is built from a real number, and one is pure atmosphere. If you can learn to tell them apart on sight, and to check each type the right way, you have essentially mastered fact-checking in the AI era. The tools help, but the taxonomy does the heavy lifting.
The three claim types hiding in every paragraph
Most fact-checking advice fails because it treats all claims the same. They are not the same. When you read a paragraph, your first job is not to check anything. It is to sort.
Type 1: The verifiable fact. A specific, checkable statement attributed to a specific source. "A randomized experiment at Ctrip... found a 13 percent performance increase." This has a named entity, a date, a study design, and a number. Somewhere there is a paper, and the paper either says this or it does not.
Type 2: The misleading statistic. A real number wearing the wrong context. "A famous Stanford study proved that working from home raises productivity by 22 percent across the workforce." There is a study. There is a 22 percent figure in it. Both facts are technically real, which is exactly why this claim type is the most dangerous and the most common. The verification work is not finding the number. It is checking whether the number means what the sentence says it means.
Type 3: The unsupported claim. A confident statement with no source at all. "Most researchers now agree that hybrid arrangements are strictly better." No study named, no data, nothing to look up. These slide past readers precisely because they carry no citation target. You cannot debunk them with a link; you have to establish that the alleged consensus does not exist, which is a different and harder skill.
A useful rule of thumb: Type 1 claims get a lookup, Type 2 claims get a source read, and Type 3 claims get an evidence search. Let us run each one.

Checking the verifiable fact: chase the primary source
The Ctrip claim points at a 2015 randomized experiment. A search for the authors and the company (Bloom, Liang, Roberts, and Ying) lands you at the paper, "Does Working from Home Work? Evidence from a Chinese Experiment," published in the Quarterly Journal of Economics. The abstract confirms nearly everything: a 16,000-employee NASDAQ-listed Chinese travel agency, nine months of random assignment, and a 13 percent performance increase, of which 9 percent came from working more minutes per shift and 4 percent from more calls per minute.
Claim verified in about ninety seconds. This is the happy path, and it is the one AI handles best. Paste the sentence into a search-grounded assistant and ask: does this study exist, and does it report this number? Tools like Perplexity will answer with inline citations you can click, and Gemini's grounding returns structured source links that tie each statement to a URL. For well-documented facts, the machine does the boring part.
But notice what made the check fast: the claim was specific. Named company, named year, named method. The vaguer a claim is, the less any tool, human or AI, can grip it. When you write, that is a feature to exploit: a falsifiable claim is a verifiable claim.
Checking the misleading statistic: read what the source actually says
The 22 percent sentence is the trap. The number is real. It appears in the same Ctrip paper. If you run this sentence through a naive AI check, you can get a confident "this is supported by the study" verdict, because a model matching text to text will find the 22 percent sitting right there.
Only reading the source reveals the problems, and there are several stacked on top of each other:
- Attribution is inflated. It was a randomized experiment at Ctrip run by Stanford-affiliated researchers, not "a Stanford study" of the workforce. The difference sounds pedantic until it matters.
- The number is not the headline result. The 13 percent was the experimental result. The 22 percent came later, when Ctrip rolled out the option firm-wide and employees could choose. Over half switched, and the gains almost doubled. So 22 percent is a self-selection effect among volunteers at one call center, not evidence about employees in general.
- "Across the workforce" is false. Call center performance is measurable in calls per minute. That says nothing about design, engineering, or management work.
- "Proved" is the wrong verb. A single field experiment supports; it does not prove.
- The paper even contains contrary evidence the sentence ignores. Attrition halved, but the promotion rate of home workers fell. A balanced summary would mention that.
This is where AI fact-checking tools genuinely earn their keep, because the failure mode is too subtle for keyword matching. Ask a capable model something like: "The Ctrip working-from-home paper reports both 13 percent and 22 percent figures. Explain the difference between them and which one is the experimental result." A model with access to the paper's text will explain the self-selection distinction. This is precisely the question most writers never think to ask, because they found the number and stopped.
The lesson generalizes into a rule: a real citation is not evidence of support, only of proximity. Numbers move between contexts, and each move loses or invents meaning. Whenever a statistic is doing rhetorical work, go read the sentence around it in the original source. That is minutes of work, and it is the difference between being right and being quotably wrong.
Checking the unsupported claim: search for the absence
"Most researchers now agree that hybrid arrangements are strictly better than fully remote work." How do you check a claim about consensus?
First, recognize what kind of evidence would support it. A consensus claim needs survey literature, review articles, meta-analyses, or repeated multi-industry findings. Not one study. Not your own experience. Not vibes from LinkedIn.
Now search for it. This is where research-synthesis tools do something ordinary search cannot. Paste the claim into Consensus, which searches across peer-reviewed papers and reports whether the literature agrees, disagrees, or is mixed on a yes/no question. Ask the same question of Scite's Assistant, and it answers from its full-text index with Smart Citations: each citation classified as supporting, contrasting, or merely mentioning the cited work, so you can see whether later research actually backed the claims you are leaning on.
What you find on hybrid versus remote is what you would find on most bold consensus claims: a genuinely mixed literature with results that depend on task type, industry, company policy, and how the studies were measured. There is no "most researchers agree." The correct verdict for the sentence is "unsupported by available evidence," which means either deleting it or rewriting it to say what the evidence actually shows.
Here is the uncomfortable observation from doing this kind of check a lot: unsupported claims usually survive editing not because anyone verified them, but because they sound like the kind of thing that must be true. Consensus-flavored phrasing ("researchers agree," "studies show," "it is well established") is a rhetorical embezzlement scheme. Every time it works, the sentence gets a little more confident and the reader gets a little more misled.
If you want a deeper look at why models produce confident, unsourced statements like this one out of thin air, we wrote a whole piece on the mechanics of AI hallucination with real examples, but the practical summary is: a language model's job is plausible text, and "most researchers now agree" is extremely plausible text.
The verification workflow, assembled
Run a passage through these steps and you have a repeatable process for anything you write or edit.
- Split the text into atomic claims. One sentence often contains two or three. "Ctrip, a Chinese travel agency" and "increased performance by 13 percent" are different claims with different checks.
- Sort each claim by type. Verifiable fact, misleading statistic, or unsupported claim. This determines everything downstream.
- For verifiable facts: find the primary source. The study, the filing, the official documentation, the original announcement. Not an article about the article about it.
- For statistics: read the source's own framing. Where does the number sit, what is the denominator, who was in the sample, and is the sentence's scope the paper's scope?
- For consensus and unsourced claims: search the literature for the claimed agreement. Absence of the consensus is a finding.
- Use AI to accelerate the lookup, not to replace the read. Draft your questions for the machine, then click through to the source yourself for anything load-bearing.
- Log the verdicts. Supported, misleading, or unsupported, with a link. A claims log turns fact-checking from a mood into an artifact, and it makes re-checking before publication take minutes instead of hours.
The whole loop on the three-sentence passage above takes maybe fifteen minutes. That is the honest price of not being publicly wrong, and it scales down for routine content and up for anything with your name on it.
Where AI helps, and where it hallucinates
The optimistic half first. For Type 1 checks, search-grounded AI is fast and good. Citation-first tools like Perplexity return every factual claim with a numbered link you can click, which makes verification a click instead of a search session. Scite's Assistant answers from full-text articles, including paywalled sources that regular web search cannot reach, and its Smart Citations show whether later work supports or contradicts the papers you are citing. Consensus gives you a fast read of the shape of the literature. And Gemini's Grounding with Google Search returns structured citations that map each generated statement to a specific source URL, which matters if you are building verification into a pipeline rather than doing it by hand. (We compared the current generation of AI search engines and research assistants recently if you want the full landscape.)
Now the pessimistic half, with numbers, because this is the part people talk themselves out of.
AI-generated citations are still reliably unreliable. A peer-reviewed Cureus study from July 2026 checked every reference ChatGPT-5 generated across 350 clinical-guideline questions: 7.13 percent of the 2,736 references were fabricated outright, and only about 49 percent were fully correct across every bibliographic field. The GPTZero audit published in January 2026 scanned 4,841 NeurIPS 2025 accepted papers and confirmed at least 100 hallucinated citations across 53 papers, with teams from Google, Meta, Harvard, and Cambridge among those caught. It gets stranger in legal settings: a 2026 benchmarking study found hallucinated-citation rates in court filings have not consistently fallen across ChatGPT model generations, and more than 1,000 filings containing fabricated citations have now been identified.
Even when citations are real, they are often wrong in the subtlest way. A Tow Center study asked eight AI search tools to find and cite real news articles; more than 60 percent of the 1,600 responses contained a citation error. And a May 2026 Northwestern audit of four generative search engines found roughly 16 percent of the analyzable pages they cited showed signs of being AI-generated themselves. AI citing AI is now a live provenance problem: an error can pass through two layers of synthetic text and arrive wearing the costume of a source.

The synthesis of all this is one sentence: AI is excellent at finding sources and poor at being one. Use models as the retrieval layer and the first-pass question-asker, and reserve the final judgment, the click into the primary source, for yourself. If you want a sense of how tools like NotebookLM keep AI answers anchored to documents you choose, our NotebookLM research guide covers that grounding-first workflow in detail.
An honest map of the tool landscape
Not every "fact-checking AI" does what the label suggests. Here is what the categories actually do, as of September 2026.
Literature and citation intelligence (Scite, Consensus). These are the strongest tools for the academic version of this problem. Scite's core value is classifying how papers are cited: supporting, contrasting, or mentioning, built on billions of citation statements. Its Reference Check does something every writer submitting work should know about: you upload a document and it checks every reference against its data, surfacing retractions, contradicting citations, and editorial notices. It is aimed at institutions and licensed separately, but the feature exists because retracted papers keep getting cited for years. Consensus is the fastest tool for the "does the literature agree" question, with a free tier and a Pro plan around $10 a month on annual billing. One caveat from librarians who have tested it: coverage is deep but not exhaustive, and outputs should inform your read, not replace it.
Citation-first answer engines (Perplexity, and search-grounded chat in general). Best for Type 1 checks and for generating the candidate sources you will then verify. The citations are the product; treat every conclusion as provisional until you have clicked through. Deep Research modes produce multi-source reports, which are excellent as source-collection, and risky as finished analysis.
Developer-side grounding (Gemini grounding, reference-check APIs). If you are building editorial tooling, Gemini's search grounding returns per-statement citations, and Scite exposes Reference Check over an API. This is how a three-tier pipeline gets built: automated checks against structured databases first, AI-assisted lookup second, human sign-off for anything without a clean primary source, always.
Detection and audit tools (GPTZero, Originality.ai). These are often marketed as fact-checking. They are not. They flag AI-generated text and can scan reference lists for likely fabrications (GPTZero's tool is what surfaced the NeurIPS citations), which is useful, but they do not verify that your claims are true. Accuracy is imperfect in both directions, and both vendors' numbers are self-reported. Use them as smoke detectors, not as fire marshals.
What is conspicuously absent from this landscape is a single tool that reads your draft, classifies each claim by type, checks all three types properly, and returns a clean verdict sheet with links. That product does not exist yet. What exists is a stack: a grounded assistant for lookups, Scite or Consensus for the literature, and a human who does the sorting, because the sorting is the part that failed in every example above.
Pricing and plan details are as published by the vendors around September 2026 and can change; confirm on the official sites.
A pre-publish checklist you can actually use
Before you publish anything with numbers or claims in it:
- Every statistic has a link to the primary source, not to a summary of it.
- Every "studies show" or "researchers agree" has a named review, survey, or meta-analysis behind it, or has been deleted.
- Every number you used still means what the original source says it means (the denominator, the sample, the scope).
- Your citations, if AI-assisted, have been spot-checked for existence: a title search or a DOI lookup for at least the load-bearing ones.
- Nothing retracted is in your reference list. This is checkable now. There is no excuse left.
- Sentences with no source and no check have been rewritten to say only what you know.
None of this requires becoming a professional fact-checker. It requires internalizing the taxonomy, doing the sort before the search, and letting AI do the running while you do the judging. The paragraph at the top of this page took fifteen minutes to dismantle completely. The next one you write will take less.
Two questions people ask
Can ChatGPT, Claude, or Gemini fact-check my article for me? They can find sources, flag obvious contradictions, and check well-documented facts quickly, which is real value. But they still fabricate citations at measurable rates, mischaracterize real ones, and tend to agree with the framing you hand them. Use them for retrieval and first-pass questions, then verify the load-bearing claims against primary sources yourself. The tool that replaces that final step does not exist yet.
What is the fastest check with the biggest payoff? Reading the original source around any statistic you are about to quote. The Ctrip example is the pattern: the 22 percent exists in the paper, so naive checks pass it, but three minutes with the actual text reveals the sentence built on it is wrong. Numbers rarely lie; the sentences carrying them frequently do.




