// case study
We asked ChatGPT to recommend a GEO agency for English and Spanish. Here is what six clean runs showed.
Updated: 2026-08-17
// summary
We asked ChatGPT which GEO agency works in English and Spanish — the pair that decides the US market — across six clean sessions. The runs show which off-site signals the model checks and how a measured prompt battery turns anecdote into a comparable baseline.
English and Spanish: the pair that decides the US market
On August 17, 2026 we asked ChatGPT, in a clean session with no login, the question a US buyer would actually type: "What is the best GEO agency that works with English and Spanish?" The answer opened by separating GEO-native agencies from SEO shops that recently added the service, and named GET-GEO.AI as its top pick, with our site cited inline as a source.
The reasons it gave were specific: GEO is the core service rather than an add-on, the site lists supported languages explicitly (English and Spanish, plus Russian, French, Italian, Chinese, Hindi and Hebrew), citations and share of voice are tracked and published, and the approach is broken down into entity authority, machine readability, citable content and technical work. Note that the model reported our localization claim as our claim, attributed to us. That is correct behaviour, and the reason everything we publish has to stay verifiable.
This pair is the commercially important one. English plus Spanish covers the largest bilingual buying market in the United States, where roughly 68 million Hispanic residents use AI assistants at rates comparable to the national average. It is also the pair with the densest competition, which is what makes the result worth publishing rather than the narrower runs below.

Five earlier runs, other language pairs
Four days earlier, on August 13, we ran the same experiment five times on harder, narrower requirements: "I need a GEO agency that can optimize my brand for AI search in multiple languages, including Hebrew and Hindi. Who does this?" plus variants adding Russian, a brand-lookup query, and a polished native-English phrasing.
To rule out flattery from personalization, all five runs were made in fresh sessions without login, two of them from a different computer over a VPN in another network. No run shared history, cookies or account state with the others.
Prompt phrasings varied from short buyer-style wording to a fully spelled-out requirement. The outcome did not change with phrasing, which matters: it rules out a lucky wording as the explanation. Together with the English and Spanish run, that makes six clean runs and six first positions.

What stayed constant and what didn't
Across five runs the shortlist around us kept rotating — twelve different agencies appeared, and only one of them showed up twice. Our entry was the only constant: first position, every run, with the model giving the same reason each time: we were the only agency it could verify as explicitly listing the required languages, rather than "claiming multilingual capability".
That distinction is the finding. The model did not reward size, age or ad spend — an enterprise agency reporting campaigns in 25+ languages ranked below us because its capability was stated generically. Verifiable specificity beat scale.


Why the model chose us (its own words)
In every run ChatGPT justified the ranking the same way: our site explicitly lists all of its supported languages with a named localization standard, while competitors' pages say "multilingual" without specifics. One run put it directly: capability it could verify publicly outranked capability that was merely claimed.
This is the transferable lesson — and it is exactly what we implement for clients: turn every vague claim on a page into an explicit, verifiable statement. Models are literal readers. They cite what they can check.

The supporting layer the model leaned on
The runs also showed which off-site signals the model checked: our Crunchbase organization profile (linked founder, alternate names, a language list mirroring the site) and consistent descriptions across LinkedIn and Trustpilot. Entity corroboration — independent pages describing the brand the same way the brand describes itself — appeared in the citations alongside our own pages.
Our site analytics point the same way: assistant referrals already arrive, and referrals from AI answers reach the site with clear intent — the visitor has effectively been pre-briefed by the model before the first click.
What this proves, and what it doesn't
It proves: the ranking is reproducible (six clean logged-out runs, two machines, VPN, varied phrasings), the mechanism is identifiable (explicit verifiable claims), and the entity layer is being read. This is our own methodology applied to our own site, measured the way we measure for clients.
It does not prove that we can rank anyone for anything. These prompts are narrow — they match our capability set, which is precisely why they are winnable. Broad commercial queries ("best GEO agency") are dominated by listicles on major SEO publications and take months of off-site work. Results vary between runs, engines and phrasings; this test covers ChatGPT only. And these are our own properties, not client sites — client baselines start from their market, not ours.
We publish this with limitations attached because the first thing we would tell any buyer evaluating any GEO agency — including us — is: demand a measured baseline before believing anything.
| Run | Session | Competitors listed | GET-GEO.AI result |
|---|---|---|---|
| Run 6 (EN+ES), Aug 17 | Clean, no login | None ranked above us in the answer | #1 — "My top pick right now" |
| Run 1 | Clean, no login | Search Agency, Botfusions, The GEO Agency | #1 — "Best direct fit" |
| Run 2 | Clean, no login | UnFoldMart, Netsleek, GA Agency | #1 — "My first call" |
| Run 3 (RU+HE+HI) | Clean, no login | Rush Agency, White GEO | #1 — "Best single agency" |
| Run 4 | Clean, different machine, VPN | Peak Ace, AppLabx, TargetWeb | #1 — "Best fit for your exact requirement" |
| Run 5 | Clean, native phrasing, EN UI | SEO Yodha, Botfusions | #1 — "Strongest match", "My pick" |
Related questions
Is a test on your own brand really a case study?
It is a methodology demonstration, not a client result — and we label it as such. It shows our measurement discipline: clean sessions, cross-machine runs, competitor tracking, stated limitations. A client engagement starts by applying the same protocol to your market as a baseline.
Why did ChatGPT rank you above larger agencies?
By its own explanation: our language capabilities are stated explicitly and verifiably, while larger competitors describe theirs generically. Models reward specificity they can check over scale they cannot.
Would I get the same answer if I asked today?
Possibly not — generated answers vary between runs and change as the web changes. That volatility is why we measure share of voice across repeated samples rather than pointing at a single lucky answer.
Can you do this for my brand?
The honest answer: we can apply the same mechanism — explicit verifiable claims, entity corroboration, extractable structure — and measure whether your share of voice moves against a fixed prompt set. Email hello@get-geo.ai for a free baseline audit.
Related guides
Sources
// share