Here is the claim that starts arguments with SEO people. If your GEO audit opens with a site crawl, you have already audited the wrong object.

Crawlers measure a website. Generative engines answer questions. Those are different systems with different failure modes, and a report full of missing meta descriptions and orphaned URLs tells you almost nothing about why ChatGPT names your competitor three times in a row and never says your name once. I have read audits that were technically flawless and strategically worthless for exactly that reason.

The term has a specific origin worth knowing before you sell or buy one of these. GEO, short for generative engine optimization, came out of a 2023 research paper by a team from Princeton, Georgia Tech, the Allen Institute for AI, and IIT Delhi, later presented at KDD 2024. Their finding was blunt: source content carrying quotations, citations, and hard statistics got surfaced by generative engines more often than content that read like marketing copy. That is a content and corpus finding. It is not a crawl finding. It sets the shape of the work.

So a real geo audit measures two things. What do the engines say about you, and where did that answer come from. Everything else is downstream of those two questions. What follows is the seven-point sequence I run, in order, and the artifact it produces.

A GEO audit is not an SEO audit in new packaging

Magnifying glass resting on a printed volume chart next to a phone, the close inspection a GEO audit demands

Rank is a property of a page. Citation is a property of a corpus.

That single distinction explains most of the confusion. In classic search, one URL competes for one query, and the audit asks whether that URL is technically fit to compete. In generative search, the engine assembles an answer from many sources, some of which you own and most of which you do not. Your homepage might be perfect and still lose, because the three review sites, two Reddit threads, and one YouTube comparison that the engine actually pulled from do not mention you.

This has a practical consequence for scope. An SEO audit is bounded by your domain. A geo audit is bounded by your category. You are auditing a conversation you are only partly a participant in, which means the deliverable has to include work on properties you do not control. Any auditor who hands you a list of on-site fixes and calls it done has scoped the job as if it were 2019.

The second difference is measurement. SEO has a stable, queryable index. Generative engines are stochastic. Ask the same question twice and you may get two different answers with two different source sets. That is not a bug you can engineer around, it is the operating condition, and it changes how you collect evidence. You need repetition and dated snapshots, not a single point-in-time crawl. Anyone reporting a single ChatGPT answer as proof of anything has misunderstood the instrument.

What are you actually auditing?

Four surfaces, and they fail in different ways.

The prompt surface is the set of questions a buyer types when they are trying to solve the problem you solve. Not your keywords. Your keyword list is a list of things people search when they already know the vocabulary. Prompts are longer, messier, and full of constraints: budget, industry, timeline, integration, region. A category term like “PR agency” is a keyword. “Which PR agency should a Series A fintech in Chicago hire if the budget is under 8k a month” is a prompt. The second one is what you audit against.

The source surface is the set of domains the engines cite when answering those prompts. This is the layer most audits skip and it is the one that predicts results. If every answer in your category cites G2, Reddit, two trade publications, and a Substack newsletter, and you are absent from four of the five, you have your roadmap. No amount of homepage rewriting closes that gap.

The entity surface is what the models believe about you as a thing in the world. Your name, category, founding, location, leadership, product names, pricing, and how all of those connect to other entities. Entity failures are quiet and expensive. A model that thinks you are a different company with the same name, or that you were acquired in 2022, will describe you wrong in every answer forever until the underlying sources change.

The retrieval surface is the plumbing. Can the bots that matter fetch your pages, and does anything useful survive rendering. This is the only part of the work that resembles a traditional technical audit, and it is the smallest part of the job.

Build the Answer Ledger first

Two colleagues comparing printed graphs and on-screen data, the kind of side-by-side review the Answer Ledger enables

Every audit I run starts with an artifact I call the Answer Ledger. It is a spreadsheet, it takes about a day to fill, and it is the only part of the deliverable clients keep coming back to six months later.

One row per prompt, per engine, per run date. Nine columns. The prompt text. The engine. The date. Whether your brand appeared at all. Where it appeared, meaning first named, in the middle of a list, or in a footnote. Every domain cited in that answer, comma separated. The exact sentence in which your brand was named, copied verbatim. The competitors named. And a free-text note for anything strange.

The verbatim sentence column is the one people want to delete and the one that earns its keep. It is where you discover that the engine names you but describes you as a “press release distribution service” when you sell managed media relations, or that it hedges with “reportedly” because its only source is a five-year-old directory listing. Presence is a binary. Framing is the thing that closes or loses deals, and you cannot see framing in a presence score.

Run 40 to 60 prompts across four intents. Category definition prompts (“what does an answer engine optimization agency do”). Comparison prompts (“X vs Y for mid-market SaaS”). Recommendation prompts (“best options for a bootstrapped ecommerce brand”). Verification prompts (“is X legitimate”, “who founded X”, “what does X cost”). Run each across ChatGPT, Perplexity, Gemini, and Google AI Overviews on the same day, because comparing an answer from March against one from August tells you about model updates, not about you.

Then sort the cited-domains column and count. That frequency table, ranked, is your source surface. It usually fits on half a page and it is more actionable than any keyword report you have ever received.

Check the retrieval surface: can the bots read you?

Now the plumbing, and it is quick.

Open your robots.txt and read it as a stranger would. The user agents that matter in 2026 include GPTBot and OAI-SearchBot from OpenAI, PerplexityBot, ClaudeBot, Google-Extended, Applebot-Extended, Amazonbot, Bytespider, and CCBot for Common Crawl. Many sites blocked several of these during the 2023 and 2024 wave of scraping anxiety, then forgot. I still find blanket blocks on marketing sites whose owners are paying someone to improve their AI visibility. Those two facts sitting in the same building is a special kind of expensive.

Understand what each block costs you before you unblock. GPTBot governs training data collection. OAI-SearchBot governs whether ChatGPT’s search feature can retrieve your pages live, which is the one that affects answers this quarter. Google-Extended governs Gemini and grounding, and Google’s own documentation is clear that it does not affect Search ranking or AI Overviews, which run off the standard Googlebot crawl. So a site can block Google-Extended and still show up in AI Overviews. Auditors who report all of these as one undifferentiated failure are guessing.

Then check rendering. Fetch a key page with JavaScript disabled and see what remains. Several retrieval crawlers do not execute JavaScript the way Googlebot does, and a page that renders its entire value proposition client-side can arrive at the model as an empty shell with a nav bar. If your pricing, your service definitions, or your FAQ answers only exist after hydration, they may as well not exist.

Last, look at whether you have an llms.txt file, and be honest about it. The proposal came from Jeremy Howard at Answer.AI in September 2024 and it is a sensible idea: a plain-text map of your most important content for machine readers. Adoption by the major engines is not confirmed, so treat it as cheap insurance rather than a fix. It costs an hour. It is not a strategy.

Who is describing you when you are not in the room?

Take the top 20 cited domains from your Answer Ledger and mark each one present or absent. This is the highest-yield hour in the whole exercise.

Most brands score between two and five. That means fifteen to eighteen of the properties that shape every AI answer in their category have never heard of them. The fix list writes itself from there, and it is a mix of earned media, review platform presence, community participation, and third-party listicles.

Pay attention to what type of source dominates. Categories cluster. Software categories lean on review platforms and Reddit. Professional services lean on trade press, association directories, and local news. Consumer categories lean on YouTube, publisher listicles, and forums. Health and finance lean hard on institutional and government sources, which is a much taller wall to climb and worth knowing before you promise anyone a timeline.

There is a second cut worth making on the same data: owned versus earned versus community. Owned means your domain and your properties. Earned means journalism, analyst coverage, and third-party editorial. Community means forums, review platforms, and user-generated discussion. Tally what share of citations in your category falls into each bucket. A category where 70 percent of citations are community sources needs a completely different budget than one where the citations are 70 percent trade press. I have seen agencies pitch six months of blog production into categories where the engines were citing almost nothing from brand-owned domains, and the audit would have caught it in an afternoon.

Also check the age of what gets cited. Sort your cited URLs by publication date where you can find one. If the engines are citing 2021 roundups in your category, the opportunity is to be in the 2026 roundup that nobody has written yet, and pitching that piece to the publication is a better use of budget than another blog post on your own domain.

Where does your entity break?

Ask each engine four questions about your company in a fresh session with no context. What is it. Who founded it. What does it cost. Is it any good.

Log every factual error, and log the hedges too. Hedging language is diagnostic. When a model says “based on available information” or “I could not find detailed information,” it is telling you that its sources are thin, contradictory, or old. That is a different problem from being described incorrectly, and it needs a different fix. Thin sourcing is a PR problem. Contradictory sourcing is a data hygiene problem.

Then trace the errors backward. A wrong founding year almost always traces to a Crunchbase profile, an old directory, or a stale press release still sitting on a wire service page. A wrong category descriptor traces to how you described yourself three years ago in the one article that got syndicated everywhere. Fixing the model is not possible. Fixing the sources it reads is, and it is unglamorous work: updating your Crunchbase and LinkedIn entries, correcting your own about page, getting a Wikidata item right, making sure your Organization schema and your sameAs links agree with each other and with reality.

Name collisions deserve their own pass. Search your brand name alone and see what else carries it: a band, a defunct startup, a regional hardware chain, a character in a video game. If a collision exists, the audit needs to record how each engine handles it, because the fix is different depending on whether the model conflates you with the other entity or ignores you in favor of it. Conflation gets fixed with disambiguating context, meaning your category and location have to travel with your name everywhere it appears. Being ignored in favor of a better-known namesake is a volume problem, and it is slow to solve.

Consistency beats cleverness here. Write one description of your company in one sentence, and use that exact sentence everywhere: schema, about page, press release boilerplate, LinkedIn, directory profiles, speaker bios. Models resolve entities by triangulating repeated language across independent sources. Six clever variations of your positioning read as six weak signals instead of one strong one.

What does a finished GEO audit hand you?

Four things, and if you get fewer than four you did not get an audit.

The Answer Ledger itself, dated, with the raw sentences intact so you can rerun it in 90 days and compare like with like. A ranked source gap list naming the specific domains, subreddits, publications, and review platforms you are missing from, in order of citation frequency. An entity correction list with each error, its likely source, and who owns the fix. And a retrieval fix list, which is usually short and should be finished within a week.

What you should not get is a score out of 100. Composite scores feel authoritative and they hide the only thing that matters, which is which of the four surfaces is broken. A brand at 62 because its entity data is a mess needs a completely different quarter than a brand at 62 because it is absent from every review site. The number is the same. The work is not.

Here is your next hour. Open a blank spreadsheet, write the nine Answer Ledger columns across the top, and fill twenty rows: five prompts across four engines, run today, with the cited domains and the verbatim sentence copied in. Twenty rows is enough to see the shape of your problem, and it will tell you more than the last three reports anyone sent you. The other thirty prompts can wait until tomorrow.