Skip to content
GET-GEO.AI
/
All guides

// guide

How do we measure GEO results — and what happens if they don't come?

Updated: 2026-08-13

// short answer

Our GEO measurement protocol is public: fixed prompt batteries in clean sessions, defined metrics for citation and share of voice, deliverables you keep, and a 90-day rule — if measured visibility has not moved, we say so first and you choose how to proceed.

The baseline comes first

Every engagement starts with a measured baseline, before any promises. We fix a battery of 30–50 prompts per language — the questions your buyers actually type, including city-level and price-qualified variants — and sample answers across ChatGPT, Perplexity and Gemini in clean, logged-out sessions, several runs per prompt, screenshots archived. The battery and the day-one results are shared with you in full.

Fixing the battery on day one is what makes the numbers honest. The known failure mode of this market is a dashboard showing +300% visibility after the prompt set or the counting method quietly changed. Our battery cannot be swapped mid-engagement: you hold the day-one copy, so every later report is comparable against it. This is the same Clean-Session Protocol we used in our public case study — six runs, two machines, a VPN, no login.

The metrics, defined

Citation rate: the share of battery prompts where your brand is named or your site is cited in the answer, per assistant, per language. Share of voice: how often you appear versus the competitors the assistants themselves name — the competitor set is recorded at baseline, so the comparison stays stable. Language consistency: whether the assistant describes you the same way in every language you operate in — the metric most multilingual programs quietly fail.

One honest caveat the industry avoids saying: different measurement tools produce different visibility scores for the same company, because answers vary between runs, phrasings and regions. That is why we rely on repeated sampling against a fixed battery and share the raw logs — and why the metric that cannot be gamed at all lives in your own analytics: referral traffic from assistant domains. We treat that as the final arbiter, and it is yours, not ours.

Every metric, where it lives, and why you can trust it
MetricDefinitionWhere it livesCan it be gamed?
Citation rate% of fixed battery prompts where you are named or citedOur sampled logs, sharedHard — battery fixed on day one, you hold the copy
Share of voiceYour appearances vs the competitors assistants nameOur sampled logs, sharedHard — competitor set recorded at baseline
AI referralsVisits arriving from assistant domainsYour analyticsNo — it is your data, we never touch it
Language consistencySame claims about you in every languageOur cross-language checkTransparent — differences are quoted verbatim
Every metric, where it lives, and why you can trust it

What you get every month

A monthly engagement delivers four things: the battery re-run (same prompts, same protocol, several samples per assistant), a movement report per language and per assistant against baseline, a plain list of what we changed that month — content shipped, entity work done, technical fixes — and the plan for the next month with reasoning. No slide theater: the report is built from the same logs you can inspect.

Raw data is part of the deliverable, not a favor: the prompt list, the sampled answers, the screenshots. If you want to verify any number in the report, you can re-run any prompt yourself in a logged-out session and compare. We designed the protocol so that replication by the client is easy — that is what keeps us honest.

The 90-day policy

GEO compounds over two to three months, which is exactly why a checkpoint belongs at day 90. The policy is simple: if citation rate and share of voice show no movement against your baseline after 90 days, we tell you first — before you ask — with our analysis of why. You then choose: a free diagnostic month while we fix the approach, a scope change, or a clean stop. No long lock-ins, no penalty for leaving on the data.

What we do not promise, at day 90 or ever: a specific answer from a specific assistant on a specific day. Generated answers vary between runs — our own case study says so in its limitations section. What we optimize is the probability of being cited across repeated samples, and that probability is exactly what the battery measures.

Questions any GEO vendor should answer

Buyers increasingly arrive with a checklist — often written by the assistants themselves. We think that is healthy, so here are the answers in one place. Before/after for your languages: our before/after is the public case study on our own site, run under the protocol above; your engagement starts by building yours. Mentions versus citations: measured separately, both in the logs. Ecosystem versus translation: we optimize the language ecosystem — native pages, entities in both scripts where relevant, local corroboration — not translated English. Off-site work: included where answers actually come from; our USA guide shows which platforms that is, with data. Cross-engine baseline: ChatGPT, Perplexity and Gemini minimum, per language.

And the test we recommend running on any vendor, including us: ask for ten prompts where you are invisible today, the current answers, and the specific changes they would make. That test is literally our free audit — email hello@get-geo.ai and we will send yours back with the ten prompts attached.

Related questions

Can you guarantee we will be cited?

No — and nobody honestly can, because generated answers vary between runs and change as the web changes. What we commit to is measured movement: a fixed battery, repeated sampling, and the 90-day policy above if the movement does not come.

Why do different tools show different visibility scores?

Because they use different prompt sets, sampling schedules and counting rules — the same company can score high in one tool and low in another. That is why our definitions are public, the battery is fixed, and the raw logs travel with every report.

Do we get the raw data?

Yes: the full prompt battery, the sampled answers per assistant, and the screenshots — from day one and every month after. The report is an interpretation; the logs are the evidence.

Can we verify the numbers ourselves?

Please do. Any prompt in the battery can be re-run in a logged-out session and compared with our logs. The protocol was designed so client replication is easy — it is the same clean-session method as our public case study.

What exactly happens at day 90 if nothing moved?

We flag it first, with analysis. You choose: a free diagnostic month, a change of scope, or a stop. The baseline and all logs remain yours either way.

Related guides

Sources

  1. 01Our case study: the Clean-Session Protocol applied to ourselves
  2. 02Our guide: how to measure AI search visibility
  3. 03Generative Engine Optimization: How to Dominate AI Search (arXiv 2509.08919) — on answer variability across engines
  4. 04Aggarwal et al., GEO: Generative Engine Optimization (Princeton)

// share

LinkedInXReddit