Skip to content
GET-GEO.AI
/
All guides

// guide

What should you ask a GEO agency before you hire one?

Updated: 2026-08-24

// short answer

Ask for evidence rather than process: an anonymised before/after dataset from a real client, prompt counts per language, how mentions are told apart from recommendations, how share of voice is computed, what happens if visibility rises and revenue does not, and who actually does the work. Below are the ten questions, what a good answer contains, and our own answers.

Why the questions matter more than the pitch

GEO is young enough that a convincing methodology page can be written in an afternoon. Proof takes quarters. That asymmetry is the buyer's whole problem: the pitch you are reading and the result you would get are only loosely related, and nobody in this market has a track record long enough to fall back on instead.

Buyers have started showing up with checklists written by the assistants themselves — they ask ChatGPT what to ask us before they ever send an email. In August 2026 we ran that exercise on ourselves: we asked ChatGPT, with browsing enabled, to assess get-geo.ai as a potential vendor. It scored our understanding of GEO, our multilingual positioning, our measurement protocol and our transparency at 9 out of 10 each — and our proven client results at 5, our maturity as an agency at 4. That is one run of one model producing a heuristic, not a metric, so hold the numbers loosely. The gap it named is not loose at all, and it is the right gap: knowing how to make a GEO agency visible is not the same as having proved you can do it for somebody else's business.

So here are the ten questions that exercise produced, what a good answer to each one contains, and our own answers — including the first question, where our answer is currently weak.

The ten questions

Every one of these separates process from evidence. The pattern to listen for is specificity: a good answer names a number, a language, a source or a date. A weak answer names a capability.

Send them in writing and keep the replies. Half the value of this list is that it is comparable — three agencies answering the same ten questions produce a document you can actually read side by side, which no amount of discovery calls will give you.

Ten questions, and how to read the answers
The questionA weak answer sounds likeA good answer contains
Can we see an anonymised 90-day before/after dataset from a real client?Our client results are under NDA.Prompts, baseline, after, per language — brand names removed, method stated.
How many prompts have you measured in our language, for clients?We cover all major languages.A count per language, and the battery size used per engagement.
Show us a client whose visibility grew in one language independently of English.It all lifts together.One named language, the baseline, and the movement against it.
What is the split of the work — content, technical, digital PR, entity, off-site?We take a holistic approach.Rough percentages, and who executes each part.
Which sources do you treat as authoritative in our market and language?High-authority websites.Named platforms for that language, and evidence answers actually cite them.
How do you tell a mention apart from a recommendation?We track overall visibility.Two separate counters, each defined, with an example of both.
How exactly do you compute share of voice?Against your competitors.The competitor set, the date it was fixed, and the formula.
What happens if visibility rises and revenue does not?Visibility takes time to convert.A named checkpoint, a diagnostic step, and a way out.
Who does the work — the founder, a team, or subcontractors?Our team of experts.Roles by name, and what happens when that person is unavailable.
Will you run a fixed-scope 90-day pilot with an agreed measurement protocol?We recommend a twelve-month engagement.A pilot scope, the protocol in writing, and who owns the data afterwards.
Ten questions, and how to read the answers

Our answers, including the uncomfortable one

Question one is where we are weakest, so it goes first. We do not yet have a client before/after dataset to show you. What we have instead is our own: a public case study in which the subject is us, run under a clean-session protocol we published, with its limitations stated inside the case rather than in a footnote. That is a real measurement and a poor substitute for a client result, and we would rather say so than dress it up. What we have changed is the design of the work: every pilot is scoped so that it produces a publishable anonymised dataset — prompts, baseline, after, per language, brand removed — from the first engagement onward. The first client to sign is buying a discount and contributing the proof.

The rest of the answers are short, because short is the point.

  • Prompts measured in your language: for clients, zero so far — we will not pretend otherwise. Our own battery runs 30–50 prompts per language across ChatGPT, Perplexity and Gemini, and that is the size we scope for a pilot.
  • Language-independent growth: unproven for clients. Our own multilingual runs are public, and they include the languages where we did worse, which is the part that makes them worth reading.
  • Split of the work: roughly half technical and entity work, a third content, the rest off-site corroboration — adjusted after the baseline, because the baseline is what tells us which of the three is actually blocking you.
  • Authoritative sources: they differ per language and we name them per engagement, from the answers themselves — we read which sources the assistants leaned on for your category before we decide where to be present.
  • Mention versus recommendation: two counters, never merged. Being named in a list and being told to start here are different outcomes, and only the second one moves revenue.
  • Share of voice: computed against the competitor set the assistants themselves named at baseline, fixed on day one so the comparison stays honest as the engagement runs.
  • Visibility up, revenue flat: that is a diagnosis, not a debate. It usually means the prompts we won are not the prompts your buyers ask, and the fix is the battery, not more content.
  • Who does the work: today, the founder does the analysis and the measurement, with specialists brought in per task. That is a real dependency, one of the risks the model flagged, and you are entitled to price it in.
  • The pilot: 90 days, fixed scope, protocol agreed in writing before day one, and the data is yours whatever happens at the end.

Four numbers that should never be merged

Most AI-visibility reporting collapses four different things into a single score, which is how a dashboard can climb while nothing changes for the business. Ask any vendor to separate them: mention rate, whether you were named at all; recommendation rate, whether you were the answer rather than an item in a list; share of voice, your appearances against the competitors the assistants themselves name; and citation quality, which sources the model leaned on when it named you. Being named is one achievement. Being named because the model read an industry study, a review platform and your own documentation is a different and much sturdier one.

Definitions matter more than the numbers here, because a vendor who has not written the definitions down can adjust them later. Ours are published in full, along with the protocol and the 90-day policy, in our guide on how we measure — that page is the long form of this section.

The comparison test: one mini-audit, three agencies

The strongest thing a buyer can do is refuse the tender format and run a small identical task instead. Send the same brief to three or four agencies and ask each for the same artefact: ten prompts where you are invisible today, the current answers quoted verbatim, the competitors those answers name, the specific changes they would make in the first 30 days, and the price of a 90-day pilot.

Then compare on five axes rather than on impressions. Are the prompts the ones your buyers would actually type, or generic category terms? Do the competitors they found match the ones you know? Do they say which sources the answers cited? Are their KPIs defined or vibes? And is the pilot priced as a pilot, or as a retainer with a different name? An agency that cannot produce this in a week is telling you something about how it will run month four.

Ours is free and the artefact is the same one described above: email hello@get-geo.ai and we will send yours back with the ten prompts attached. Run it against two competitors of ours — the comparison is the point, and we would rather lose it on the data than win it on a call.

Red flags

None of these mean an agency is dishonest. Each of them means a specific thing you will not be able to verify later, and later is when it matters.

  • “Under NDA, but trust us.” Anonymised data exists for exactly this situation — numbers without names. Refusing the format, not the names, is the flag.
  • A guaranteed position, or a promise to “get you into ChatGPT”. Generated answers vary between runs; anyone guaranteeing a specific one is guaranteeing something they do not control.
  • A dashboard whose prompt set can change mid-engagement. If you do not hold a copy of the day-one battery, every later chart is unfalsifiable.
  • A single visibility score with no published definition. Ask what it counts. If the answer takes more than two sentences, it counts whatever is convenient.
  • A multilingual pitch with no per-language numbers. Multilingual GEO fails language by language, so a single aggregate hides exactly the thing you are buying.
  • A case study where the agency is its own client and does not say so in the opening paragraph. Ours is one — which is why we say so in ours, and again here.
  • Refusal to hand over raw logs. The report is an interpretation. The logs are the evidence, and they should travel with every report by default.

Related questions

Is it not strange for an agency to publish the questions that expose its own weak spot?

It is unusual, not strange. Buyers now get this list from an assistant anyway — we would rather meet the questions with written answers than improvise them on a call. And the one weak answer is temporary: it closes the day our first pilot produces a publishable dataset.

What if an agency refuses to share raw logs?

Then every number in every report is unverifiable, including the good ones. Anonymised logs — prompts, sampled answers, screenshots, brand names removed — cost nothing to share and are the only thing that lets you re-run a measurement yourself.

How many prompts make a defensible baseline?

For a single language, 30–50 prompts sampled repeatedly across at least three assistants is enough to see movement without drowning in noise. What matters more than the count is that the set is fixed on day one and that you hold a copy of it.

Should we start with a pilot or a retainer?

A pilot, always, and not only with us. Ninety days is long enough for GEO work to compound and short enough that a wrong choice costs you a quarter rather than a year. A vendor who will only sell twelve months is pricing their own uncertainty into your contract.

A model scored you 5 out of 10 on proven client results. Why hire you?

Because that score measures our history, not our method — and the method is public, testable and already produced a measured result on a live subject. If proven client history is what you need most, hire someone older. If a published protocol, raw data and a 90-day exit matter more, that is the trade we are offering, and we have stated it plainly rather than hoping you would not ask.

Related guides

Sources

  1. 01Our guide: how we measure GEO results, and the 90-day policy
  2. 02Our case study: we asked ChatGPT to recommend a GEO agency
  3. 03Our guide: how to measure AI visibility and citations
  4. 04Aggarwal et al., GEO: Generative Engine Optimization (Princeton) — on answer variability between runs

// share

LinkedInXReddit