On November 16, 2013, at the Padova Astronomical Observatory, I gave a talk titled “Web Communication— from Search Engine Optimization to Semantic Engine Optimization. I wasn’t being prophetic. I was describing where the ground already seemed to be shifting, back when “semantic search” was Google’s polite way of saying it was starting to understand meaning instead of just matching strings. Thirteen years later, it reads less like an old slide deck and more like an early draft of a problem that hadn’t been named yet.
It wasn’t a one-off. I’d made a version of the same case in 2011 at the Genoa Science Festival, again in 2012 at the Journées Hubert Curien in France (later rebranded Science & You, where I spoke again in 2015), and again in 2017 at a UWE Bristol seminar I titled “Be Visible or Vanish.”
Two books and a chapter in a manual on scientific writing came out of that same stretch, all built around one recurring question: if nobody finds your work online, does it even count as published?
Generative Engine Optimization, or GEO, is the newest name for that question — and it won’t be the last. What keeps changing is only the audience being optimized for: first other humans scrolling a feed, then a ranking algorithm, then a semantic layer starting to guess at meaning, and now a chatbot deciding, in the half-second it takes to assemble an answer, which handful of sources are worth citing out of the millions it could have chosen.
Most researchers I talk to have no idea that mechanism is already running underneath every question their audience types.
From ranking to selection
Traditional SEO trained an entire generation of communicators, myself included, to think in terms of position: where does your page land on a results list that a human being will scroll through, top to bottom, hoping the answer is somewhere in the first five entries. That world hasn’t disappeared, but a second, stranger one has grown up alongside it.
When someone asks ChatGPT, Perplexity, or Gemini a scientific question, there is no list to scroll. There’s a synthesized answer, built through retrieval-augmented generation, or RAG: the system quietly searches the web for what it judges to be the most authoritative material, then rewrites it into something conversational and hands it to the user as if it had known the answer all along. If your press release, your abstract, or your paper isn’t legible to that retrieval step, it simply never enters the conversation. Not lower-ranked. Absent.
That distinction matters more than it sounds.
Ranking is a spectrum; you can be page two and still exist.
Selection is binary. An AI system either surfaces your work as a source or it doesn’t, and there’s no equivalent of scrolling further to find you if it doesn’t.
Writing for a reader that doesn’t skim, it parses
Large language models read differently than people do, and that difference has very concrete consequences for how a research communication piece should be built. Where a human reader tolerates — sometimes even enjoys — a slow build-up, a model rewards structure it can parse cleanly and penalizes ambiguity it has to resolve.
The most immediately useful habit is what I’d call FAQ-mapping: writing your abstract, or the opening lines of a press release, as if it already anticipated the exact question a curious reader would type into a chatbot.
“What did this study find, and why does it matter ” shouldn’t require three paragraphs of throat-clearing before the answer appears; state it early, in plain declarative sentences, and let the nuance follow.
Models tokenize and synthesize text, and dense subordinate clauses or double negatives genuinely make that job harder — not because the model is unintelligent, but because ambiguity resistant to a single clean parse is exactly the kind of thing a synthesis step tends to smooth over or drop.
Format matters almost as much as phrasing. Well-structured Markdown or clean HTML tables are, for lack of a better description, catnip to these systems: a clearly labeled table of results is far more likely to be lifted intact into a generated answer than the same numbers buried in a paragraph. The same goes for genuinely useful bullet points — not the decorative kind, but ones that isolate a discrete, extractable fact per line. And on the more technical side, structured metadata (Schema.org markup identifying author, institution, publication date, and subject) does for AI crawlers roughly what a well-organized filing cabinet does for a librarian: it removes the guesswork.
Earning the citation, not just requesting it
Structure gets you read; it doesn’t get you trusted. The second half of GEO is closer to what I’d call semantic authority, and it has more in common with old-fashioned scholarly rigor than most people expect from anything associated with an algorithm.
Generative engines show a measurable preference for text that itself cites credible sources and verifiable statistics — the logic being, more or less, that a well-documented piece is more likely to be a reliable node to cite in turn. A press release littered with vague superlatives and no anchoring data is not just bad writing; it’s now also a weaker candidate for algorithmic trust.
One small habit pays outsized dividends here: whenever a piece introduces a technical term a general audience won’t already know, give it an explicit, quotable definition — literally a sentence structured as “We define [term] as…” Models tend to lift that exact formulation when a user later asks for a definition of the concept, which means the phrasing you choose in that one sentence may end up being the phrasing an AI puts directly into someone else’s mouth.
Rebuilding the press release around a paywall that isn’t going away
None of this matters more than in the science press release, which has always occupied the strange position of being the one openly accessible document standing between a paywalled paper and a curious public.
GEO raises the stakes on that role rather than replacing it. A release built around what I think of as the direct-answer principle — cramming the core finding, its significance, and its novelty into the first hundred words — gives a generative engine everything it needs without having to infer anything.
Closing with a short, explicitly labeled “Key Takeaways” section, three or four lines summarizing the research in isolatable form, essentially hands the model a pre-packaged citation it doesn’t have to construct on its own; these sections tend to get lifted into chat responses almost as-is, which is either flattering or slightly unsettling depending on your mood that day. And when you do link out to the original paper, the anchor text carries real weight: “the MIT fusion study” tells a crawler something useful, while “click here” tells it nothing at all.
The upside of taking all this seriously is bigger than better metrics. A well-optimized, openly accessible press release lets an AI system know your research exists and represent it accurately even when the paper itself sits behind a wall the model can’t cross — which means the paywall stops being the full stop it used to be. And because, by most estimates, only a small fraction of research institutions are currently writing with any of this in mind, showing up correctly right now is less about chasing a trend and more about claiming a mostly empty room before it fills up.
A short, honest list of tools
Almost everything marketed as a “free GEO tool” is a free trial wearing a free tool’s clothes — one scan, then a paywall. Worth knowing that going in, so here’s the list split honestly in two.
Actually free, no trial clock running: HubSpot’s AI Search Grader is the one I’d point a research office to first — a real product from an established company, no card required, giving both a visibility score and a sentiment read (not just whether you’re mentioned, but how the AI characterizes you when you are).
Google Search Console‘s experimental AI-powered configuration lets you query, in plain language, how your site performs in AI Overviews specifically — coverage is still partial, but it comes straight from Google and costs nothing.
And the method every paid tool quietly recommends you also do: open ChatGPT, Perplexity, Claude, and Gemini yourself and ask the questions your audience would ask. It doesn’t scale past a handful of queries, but fifteen minutes of doing this by hand tells you things a score never will — how you’re framed, in what tone, alongside which competitors.
Free tier, limited, worth trying once before you commit to anything: Peec AI and Otterly both offer a one-off free snapshot report — useful for a baseline reading, not for ongoing monitoring. For genuinely continuous tracking across engines and competitors, that’s where the paid platforms (Profound, Writesonic, and others) start to earn their subscription — but that’s a second-stage investment, not a first step.
The part I’ve been putting off
Everything above is the part I can give away in an article: the vocabulary, the habits, a couple of tools worth trying this week. What I’ve been building, more slowly, is a training for research communicators that turns those seven or eight principles into an actual workflow — the kind of thing a research office could run every time a paper comes out, rather than remembering once and forgetting by the next press release. It’s still taking shape, but it’s coming.
If the 2013 talk was early, I’d rather this one be on time — and I’d rather you hear about it from me before you need it.
People also ask
What is Generative Engine Optimization (GEO)?
GEO is the practice of writing and structuring research content so that AI systems like ChatGPT, Perplexity and Gemini can find, understand and cite it. Unlike traditional SEO, which competes for a position on a results page, GEO is about being selected as a source inside an AI-generated answer — where your work is either included or absent, with no page two to scroll to.
How is GEO different from SEO?
SEO is about ranking: where your page lands on a list a human scrolls through, so page two still exists. GEO is about selection: an AI either surfaces your work as a source or it does not. Ranking is a spectrum you can survive on; selection is binary.
How can researchers make their work easier for AI to cite?
Lead with the finding in plain declarative sentences (FAQ-mapping), use clean structure such as Markdown tables, extractable bullet points and Schema.org metadata, and cite credible sources and data yourself. Explicit definitions phrased as “We define X as…” and a short “Key Takeaways” section help too, because models tend to lift those formulations almost verbatim.
