GEO (Generative Engine Optimization) is the craft of making your pages easy for AI assistants to read, quote, and cite. This playbook explains how to get cited by AI, with every move labeled by honest evidence: measured, official, or snake oil.
A customer opens an AI assistant and types a question. The kind of question your business has answered for years. You wrote the guide for this. You tested it, you lived it, your page answers it better than anyone else's.
The answer appears. It's a good answer. And at the bottom, in a neat little list, sit the sources. Your competitor's name. Your competitor's link. Their page gets named. Yours gets skipped.
Nothing happened. Nobody clicked the wrong thing. The assistant simply picked the pages it could quote easily, and your page didn't make the cut. It stings because it's silent. With old search, at least you could see yourself at position seven and climb. Here you're just... not in the answer.
Here's the promise of this playbook: you can't buy a citation, and nobody honest can guarantee you one. But you can make your page the easiest one in the room to quote, and there's real evidence for which changes help, which are just official advice, and which are snake oil. That's what this post is: a ranked playbook with the receipts attached, written so a ten-year-old could follow it.
In short:
- AI assistants pick sources the way a librarian picks passages: readable, clearly labeled, and backed by receipts.
- You can't buy or guarantee a citation - but you can make your page the easiest one to quote.
- Some moves are measured to work, some are official advice, and some are snake oil. This post labels each one honestly.
- You'll know it's working by checking the scoreboard, not by wishing.
The new librarian in town
Imagine your town just got a second library. The old one, Google Search, is a giant card catalog. You hand it a slip of paper with words on it, it checks which cards match those words, and it points you to a shelf. It never reads the books. It just matches cards.
The new library works differently. It has a friendly librarian. You walk up and ask a real question, out loud, the way you'd ask a person. The librarian fetches about five books, reads them, and then blends the best parts into one answer she tells you out loud. Then she points at the pages she used: "that bit came from this book, that bit from that one."
In plain words: a generative engine is a search engine plus an AI model. It fetches a few pages, reads them, and retells them as one answer with citations.
That new librarian is what researchers call a generative engine. It's two machines glued together: a search engine (the part that fetches pages) and an AI language model (the part that reads and retells them with citations). You've met them: ChatGPT search, Perplexity, Google's AI Overviews and AI Mode.
This creates a new game with a new name: GEO, short for Generative Engine Optimization. A research team from Princeton University coined the term in a paper published at the KDD 2024 conference. Their one-line idea: old search matched keywords on cards, but the librarian actually understands and retells, so keyword-matching tricks don't apply anymore (paper §2.2). You may also hear GEO called answer engine optimization or AI search optimization. Same idea, same moves.
"Wait," you might ask, "GEO vs SEO - how are they different?" Here's the kid-friendly version: old search was a matching game (does the card contain the words?). Generative engines are an understanding game (can I retell this page in my own words?). Different game, different moves.
One more thing before we go on. In the old game, winning meant being result number one. In this new game, "winning" is fuzzier. The Princeton researchers measure visibility instead of rank: how many words about you appear in the answer, how early in the answer they show up (a front-row seat counts more than the back row), and how much the whole answer leans on your page. No single number - more like a newspaper report about how much the librarian talked about you.
You can read the full paper yourself: GEO: Generative Engine Optimization (Aggarwal et al., KDD 2024) on arXiv.
What AI assistants actually cite (the honest evidence)

Back to our librarian. Two books sit side by side on the same shelf. Both cover the topic. She quotes one and ignores the other. Why?
Because one page hands her a clean, checkable sentence with a receipt attached, and the other makes her dig for it.
The short answer: AI assistants cite pages that are easy to quote, and the same moves help you get cited in AI Overviews, ChatGPT, and Perplexity. Easy to quote means fluent sentences that stand alone, numbers you can check, and direct quotes from real, named sources.
That's not a guess. The Princeton team ran a real experiment called GEO-bench: ten thousand search queries, real web pages, and a simulated AI assistant that fetched the top five results and generated answers. Then they measured which changes to a page actually made it get quoted more. They also ran a smaller live test (200 queries) on the real Perplexity.ai.
Three things made pages quotable, test after test:
- Prose that's easy to lift out. Smooth, well-formed sentences that can stand on their own. Not walls of jargon. Not fragments.
- Numbers you can check. Concrete statistics instead of vague claims.
- Direct quotes from credible people and sources. Real words, really said, really findable.
In the researchers' own words: "Generative engines value not only content but also information presentation." What's on the page matters, but how it's laid out matters almost as much.
Think of it as receipts. When your page says "most users leave in ten seconds," the librarian has to take your word for it. When your page says "load times under two seconds kept 70% of visitors, per the 2024 Web Performance Report," you've handed her a receipt she can pass along to the visitor. Receipts make her comfortable quoting you.
The receipts showed up in the measurements, too. In the Princeton paper, adding statistics, adding direct quotations, and citing your sources each improved visibility by roughly 30 to 40 percent (paper-measured, Position-Adjusted Word Count, vs pages left unoptimized). Making the writing more fluent and easier to understand added another 15 to 30 percent. Stacked together, the best methods reached visibility gains of up to 40 percent overall, and up to 37 percent on the live Perplexity.ai test.
Now the honest part, because this playbook only works if we're honest. The scientists tested a librarian setup from 2023: the simulated engine fetched the top five Google results and used GPT-3.5 to write the answers. Today's engines are newer and smarter. So treat these numbers as strong hints, not laws of physics. The paper itself also warns that what works varies by topic - statistics win in some subjects, quotes in others.
And one more finding that should make small sites sit up. When the researchers split pages by their Google rank, the sites sitting at rank five - the ones losing the old game - gained up to +115.1 percent visibility from adding citations, while rank-one sites actually lost ground - about 30 percent down in the paper's table. The librarian isn't only picking the famous book. She's picking the page with the receipts, wherever it sits.
Even the engines admit they're not perfect judges. OpenAI's own help page for ChatGPT search says that "search results and citations can be incomplete, outdated, or incorrect" - which is why ChatGPT shows its links and asks people to open them. You can read their explanation here: Searching the web with ChatGPT (OpenAI Help Center).
The playbook: eight moves, honestly ranked

Picture a ladder of trust. The top rungs are moves scientists actually measured. The middle rungs are what the librarian's own manual recommends. The bottom rung is what to skip - because it was measured and it failed.
Start with the measured moves: add statistics, add direct quotes, cite your sources, and write short, clear sentences. Those four carried the biggest visibility gains in the research.
Every move below gets three things: what to change today (concrete, copy-able), why it works (the librarian-reason), and an honest evidence label. No move gets dressed up as more proven than it is.
Rung 1 - measured by the paper (do these first)
1. Add statistics. Replace vague claims with numbers. "Our customers save time" becomes "customers cut onboarding from 3 weeks to 4 days." Why: a number is a receipt the librarian can quote without softening it. Evidence: paper-measured - Statistics Addition landed in the +30-40% visibility band in the Princeton experiment.
2. Add direct quotations. Find people worth quoting - researchers, officials, named experts - and quote their exact words. "I think people overestimate what one tool can do," not a paraphrase of that feeling. Why: a direct quote is pre-wrapped. The librarian can lift it word-for-word without misquoting anybody. Evidence: paper-measured - Quotation Addition was the single best standalone method in the study's tables.
3. Cite your sources. Every factual claim gets a reference. Not a vague "experts say" - an actual named source the reader could check. Why: citations are proof the librarian can hand the visitor, and they made the biggest jump for smaller, lower-ranked sites (+115.1% for rank-5 pages in the study). Evidence: paper-measured, and it worked best of all when combined with other moves.
4. Make it fluent and easy to understand. Short sentences. One idea per sentence. Words a kid knows. Why: the librarian retells your page in her own words - a smooth, simple sentence is easy to retell; a tangled one is easy to skip. Evidence: paper-measured - Fluency Optimization and Easy-to-Understand Language together lifted visibility 15-30%. (You're reading a demonstration right now.)
Rung 2 - stated in the engines' official docs
5. Use clear headings and make the main content easy to spot. Someone skimming your page should find the actual answer without hunting past ads, widgets, or sidebars. Why: Google's own guidance for AI experiences says to "make it easy to distinguish main content from other content," and to put important information in plain text - the librarian reads text, she can't read your beautiful infographic. Evidence: official-doc-stated (Google Search Central, May 2025) - this is the librarian's manual, not a lab measurement.
6. Add structured data that matches the visible text. Structured data (schema.org markup) is a little label inside your page that tells machines what each thing is - this is the recipe, this is the price, this is the author. Why: Google's docs say structured data helps systems understand your page, but it "must match the visible text" - honest labels only. Evidence: official-doc-stated. This is not a finding from the Princeton paper - a common mixup you'll see online.
Rung 3 - sensible practice, labeled honestly
7. Answer first, in standalone sentences. Put the answer in the first sentence of a section, stated so completely that it makes sense even if lifted out alone. Why: the librarian quotes sentences, not introductions. If your answer only makes sense after three paragraphs of wind-up, there's nothing to lift. Evidence: inferred, not directly measured. It follows from the paper's fluency findings plus how the engines quote pages - but no study has isolated this one by name. Honest label: a strong bet, not a proven fact.
8. Keep it fresh. Update your numbers, dates, and claims when the world moves. Why: OpenAI itself warns that citations "can be outdated" - a librarian who spots a five-year-old statistic may quietly pick a newer book. Evidence: part official-doc (OpenAI's warning), part common sense. The Princeton paper never tested freshness on its own, so treat this as reasonable practice, not measured fact.
The skip rung - measured to fail or do nothing
- Keyword stuffing. Repeating your target phrase over and over. In the study it scored below doing nothing at all (17.7 vs 19.3 baseline) - worse than leaving the page alone. Imagine shouting the book's title at the librarian repeatedly. She ignores shouting. That world is gone.
- Bossy, "authoritative" tone. Sounding impressive and confident didn't help overall - "no significant improvement" in the paper's measurements. The librarian isn't moved by swagger. (One exception: in debate- and history-style topics it did top the list, so know your genre.)
One more finding worth your time: combining moves beats any single move. In the paper's experiments, pairing fluency work with statistics beat every lone strategy by a clear margin, and citing sources shone brightest when stacked with others (paper §5.3). Don't pick one rung. Climb.
If all this leaves you wondering how different engines judge answers differently - they genuinely do - the same skill of testing claims empirically applies when you pick the right AI model for every job. And notice what changed underneath: old search matched keywords on cards, while today's model-driven engines actually read and retell. That's the shift the whole playbook rides on.
For Google's own short list of what it wants from pages - Googlebot access, clean headings, honest structured data - their guidance is public: AI features and your website (Google Search Central).
What you can't control (and the snake oil to skip)
Now for the part nobody sells you: the librarian's own head.
You can hand her the best passages in the world. You cannot stand behind her chair and point at your book. You can't watch her think, and you can't pay for a better chair-spot. Anyone selling you a chair-spot is lying to you.
The honest answer: no one can guarantee you an AI citation. The engines are black boxes, and Google says a page just needs to be indexed and eligible for a normal snippet. Nothing more.
Here's what's genuinely outside your control, straight from the engines' own docs:
- No guarantees, ever. These systems are black boxes - the Princeton paper says creators have "little to no control over when and how their content is displayed." Google adds that AI Overviews are shown only when they add something to regular search, and they "often don't trigger" at all. Your page can be perfect and still not appear.
- There is no AI-Overview secret handshake. Google states it plainly: a page just needs to be indexed and eligible for a normal search snippet. "There are no additional technical requirements." No special code, no special file, no membership card.
- Each librarian has her own habits. ChatGPT decides on its own when a question needs a web search. Perplexity runs two different crawlers - PerplexityBot, which you can allow in robots.txt, and Perplexity-User, which fetches pages when a visitor asks and generally ignores robots.txt (their own crawlers page explains the difference). Google's AI features use something called query fan-out: they fire off several related searches across subtopics, which means the answer's source pool is wider than page one of Google - you don't have to rank first to be in the running (official-doc-stated, Google Search Central).
Because engines change under your feet - new models, new habits, new rules - the durable practices below beat clever tricks every time. The engines themselves keep getting retrained, as we covered in how AI models actually get smarter during training.
Now, the snake-oil shelf. Four bottles, each with the honest rebuttal:
- "Guaranteed AI rankings." Impossible on its face. The engines are black boxes, Google promises nothing, and nobody outside those companies knows the recipe. Anyone guaranteeing placement is selling smoke.
- "Our llms.txt file gets you cited." llms.txt is a small text file some sites add, hoping AI engines read it. Google has said officially, in its AI optimization guide, that Google Search - including its generative AI features - does not use llms.txt or any "special" markup or AI text files. This one is debunked at the source: Google's guide to optimizing for generative AI features.
- "Keyword stuffing is back." Measured, tested, and worse than doing nothing (see the skip rung above).
- "Just sound more authoritative." Measured, tested, no overall effect. Swagger isn't a receipt.
One more, for balance: sometimes you'll see headlines claiming AI Overviews destroy traffic, or vendors quoting big statistics about who gets cited. The honest position is that both directions get exaggerated - Google reports that clicks from AI Overviews tend to be higher quality (their claim, not independent data), and vendor citation statistics you see online are industry-reported numbers, not verified measurements. Take both with a grain of salt.
How to know if it's working
You wouldn't know if your team scored without a scoreboard. For years, AI citations had no scoreboard - the stadium was dark. In June 2026, someone finally switched on the screen.
To measure your LLM visibility, check three scoreboards: Search Console's AI performance reports, referrer visits from perplexity.ai or chatgpt.com, and a monthly log of which sources the engines cite.
The official scoreboard. On June 3, 2026, Google launched Generative AI performance reports inside Search Console (official-doc-stated). They show you impressions from Google's AI features - AI Overviews, AI Mode, and Discover - broken down by pages, countries, devices, and dates, and the data also flows into your normal performance report. It's rolling out gradually to a subset of sites first, so if you don't see it yet, sit tight.
The referee's notes. Watch your analytics for visits coming from perplexity.ai or chatgpt.com referrers. This is common practice among site owners, not an official spec the engines document - so treat it as a useful signal, not a complete count.
The DIY citation check. The most honest method costs nothing: once a month, ask the engines the same handful of questions your customers ask, and log which sources they cite. The Princeton paper's own visibility metrics - how many words about you, how early in the answer, how much the answer leans on you - make a fine scoring sheet for your log.
The honest gap. There is no official free dashboard that counts your citations across all engines - Google's new report covers Google's features only. Paid vendor tools exist, but their numbers are industry-reported, not verified measurements; if you use one, treat it as a rough weather report, not a scoreboard.
The long game
The books a librarian keeps pulling off the shelf aren't the loudest ones. They're the clearest, most trustworthy ones that keep getting borrowed.
That's the whole long game: extractable answers, receipts (statistics, quotes, citations), clean structure, and freshness. No magic. These outlive every engine update, because every future engine - whatever model, whatever interface - still needs one thing above all: passages it can trust and quote.
And that sting from the beginning - watching a competitor get named while your better page sat unquoted? That's not the end of the story. It's a to-do list. The move now is simple: be the easiest page for the next librarian to quote.
GEO FAQs
Can you optimize specifically for Google AI Overviews?
No - and Google says so itself: there are no additional requirements beyond normal Search eligibility (indexed, snippet-eligible). There is no special code, file, or handshake. Work the fundamentals in this playbook instead.
Does keyword stuffing help in AI assistants?
No. In the Princeton study it performed below the unoptimized baseline - worse than doing nothing. AI assistants understand and retell language; they don't match keywords, so shouting your phrase repeatedly just makes you the loud book on the shelf.
Does llms.txt make Google's AI features cite you?
No. Google stated officially that Search - including its generative AI features - does not use llms.txt or other "special" markup or AI text files. Anyone selling llms.txt as a Google citation booster is selling snake oil.
Can small sites actually get cited over big domains?
Yes, measurably. In the paper, rank-5 sites gained up to +115.1% visibility from citing sources, while rank-1 sites lost ground - about 30 percent down in the paper's table. The librarian picks the page with the receipts, not just the famous book.
How long until I see citations?
Honestly: no measured timeline exists anywhere in the research. Don't watch the calendar - watch the scoreboard: Search Console's AI performance reports, referrer checks, and a monthly log of which sources the engines cite for your questions.




