Blog
14 min read
AI visibility monitoring

API vs Interface: Why AI Monitoring via API Shows the Wrong Answers

Why AI answers collected through an API differ from what users see in the app. Surfer data on 13,779 answers, our own before/after comparison of Gemini moving from API to the live interface, and a checklist for vetting an AI visibility tool.

AI visibility monitoringmethodologyAPIdata collection
Vladislav Puchkov
Vladislav Puchkov
Founder of GEO Scout, GEO optimization expert

Two ways to collect an AI answer

Via API. The tool sends the prompt to a developer endpoint and gets the model output back. The tool itself picks the model, temperature, whether web search is on, and how it is configured. It is fast, stable and usually cheaper: response formats barely change, and scaling to thousands of prompts is easy.

Via the interface. The tool asks the question where a person asks it — chatgpt.com, gemini.google.com, Alice, the Yandex or Google results page — and records what appears on screen: text, source links, cards, product carousels, ads. It is slower and more expensive: interfaces change, collection needs constant maintenance, and each answer takes tens of seconds.

The price gap is the main reason many tools choose the API. When we surveyed per-request pricing across Russian GEO services, vendors that sell both modes charge several times less for "answer from model memory" than for "answer with web search". The question is what you give up for that discount.

Why the API answer is not the app answer

A different model or version

AI vendors do not disclose which exact model serves app users on a given day. The API sells the "closest equivalent", and the monitoring tool picks it. Surfer notes that, per Google's documentation, AI Overviews runs on a model that is not sold through the API at all, so they had to reuse the AI Mode method for it.

A different system prompt

Apps ship long instructions that set tone, format, how to treat brands and ads, and how to cite. The API has none of that, only what the tool sends. As Surfer data scientist Maciej Gruszczyński puts it, a single sentence in the system prompt can turn a response by 180°, and consumer apps run extensive instructions that vary by product.

Web search and source selection

In the app, the product decides whether to search, how many queries to run and which pages to cite. Via the API, search is a tool with whatever parameters the vendor configured. In Surfer's data, interfaces searched on 88% of prompts for ChatGPT and 100% for the other four products, while APIs searched on 83% for ChatGPT, 85% for Gemini, 94% for AI Mode, and 100% only for Perplexity. Different search behavior means different sources, and sources decide which brands make it into the answer.

Location and language

The app knows where the user is. "Where to get my apartment renovated" asked from Moscow and from Novosibirsk yields different answers, and that is part of what your buyer sees. An API call is tied to no city by default. When collecting from the interface, GEO Scout sets the browser location to the brand's city, as described in the methodology.

Cards, carousels and ads

This is the most visible difference: the API has none of these blocks. The model returns text; the app draws product carousels, business cards from maps, images and ads around it. In our database over the last 30 days (5,938 non-empty Alice AI answers across 132 brands):

  • 59% of Alice AI answers had a Yandex Direct ad labeled "Promo" next to them — we track these in AI Advertising;
  • 39% included business cards from maps;
  • 15% included a product carousel with prices and stores.

ChatGPT over the same 30 days showed business cards in 11% of answers and product carousels in about 1% — on Russian-language prompts they are still rare. The shares depend on niche and query type, but the point is different: an API-based tool will report zero for all of these. If your product ranks first in an Alice AI carousel, an API-based report will not know it. More on carousels in the E-com guide.

Sub-queries and modes

Before answering, an AI engine runs several searches — the query fan-out. The API does not always expose them. Grok is a clear example: when we moved Grok to interface collection on September 18, 2026, answers started arriving with sub-queries that the API had never returned at all, plus the full list of sources found, marked as cited or only retrieved. More on sub-queries as a metric in the Query Fan-Out guide.

Search surfaces are a separate case. Yandex Search with Alice is not a chat; it is a block on the results page. That block is what users see, so GEO Scout captures the "Alice AI Quick Answer" from the live results page.

What Surfer's study found

Surfer ran 1,000 prompts through five AI products — ChatGPT, Perplexity, Gemini, Google AI Mode and Google AI Overviews — twice: via the interface and via the API. 13,779 answers in total, collected on August 4, 2026, published on September 25.

Which brands. Brand list overlap between API and interface (Jaccard) was 15.5–23.8%, and 21.3–31.6% after merging name variants of the same brand. Gemini had the highest overlap — roughly 3 brands out of 10.

How many brands. The API named more brands in all five products. ChatGPT is the extreme: 13.8 brands per answer via API versus 7.9 in the interface.

Length. API answers were mostly longer: 2.05× for ChatGPT, 1.27× for Gemini, 1.06× for AI Mode. Perplexity was the exception, with a slightly shorter API answer.

Sources. The direction varies by product, but the gap is everywhere.

ProductSources per answer: interfaceSources per answer: APIDomain overlap
ChatGPT12.13.14.8% (lowest in the study)
Google AI Mode22.214.6—
Perplexity10.2719.526.7% (highest)
Gemini3.46.5—
Google AI Overviews10.314.6 (AI Mode method)—

At the page level, overlap was 19.7% for Perplexity and 4.8–6.2% for ChatGPT. For Gemini, AI Mode and AI Overviews page-level overlap could not be measured: their APIs return per-call redirect tokens instead of destination URLs.

Limitations Surfer acknowledges. Each prompt was asked once, and AI models are non-deterministic, so part of the disagreement is randomness rather than the collection method. Model parity in the API is approximate. AI Overviews reused the AI Mode API method. One AI Mode method covered 813 of 1,000 prompts before hitting a Google search quota.

The first limitation matters most: without it, you cannot tell how much of the gap is the model's own noise. That is what we tried to account for.

Our data: Gemini before and after moving to the interface

Until August 31, 2026, GEO Scout collected Gemini answers via the API, using Gemini 3.5 Flash-Lite with native Google Search grounding. Since August 31, Gemini has been captured from gemini.google.com — see how to track brand visibility in Gemini. The rollout started on August 26, so the two periods overlap by a few days. Every stored answer records its collection method, so we did not have to split the data by date.

Sample. We took prompts that have answers from both methods: 47 prompts from 13 brands, mostly in Russian. API: 63 answers from July 27 to August 31. Interface: 64 answers from August 27 to September 24. Empty answers were excluded, as were answers replaced by our follow-up request for sources. We averaged per prompt first, then across prompts, so both groups have identical prompt composition.

Metric (mean across 47 prompts)APIInterface
Sources per answer4.943.38
Answers with no sources14.9%9.6%
Answer length, words≈470≈370
Brand's competitors named per answer4.394.00
Answers mentioning the tracked brand31.4%29.8%

The API cited more sources on 36 of 47 prompts and gave a longer answer on 43 of 47. The length gap holds after stripping markup and links. At the same time, the API also had more answers with no sources at all: it either did not search or brought back a long list.

Baseline. To separate the effect of the collection method from ordinary model noise, we compared pairs of answers to the same prompt: API vs interface, interface vs interface on different days (1,769 pairs across 269 prompts), and API vs API (69 pairs across 24 prompts — a small sample).

Pair of answers to one promptDomain overlapPage overlapCompetitor set overlapSame "brand mentioned" status
API vs interface19.3%14.5%25.1%79.8%
Interface vs interface28.7%19.1%58.6%85.4%
API vs API38.3%31.0%39.3%87.0%

Interface-vs-interface overlap barely depends on the gap between answers: 29.5% of domains within a week, 27% beyond three weeks. So the drop to 19% for API vs interface is not explained by the API answers being a month older.

What this means:

  1. Averages look deceptively similar. The share of answers mentioning the brand is almost the same, 31% vs 30%. An API-based report would show a "correct" number, yet on individual answers the brand status disagrees more often: 20% of API–interface pairs versus 15% for the interface against itself.
  2. Different competitors. The interface names roughly the same competitors day to day (59%); API and interface agree on only 25%. For competitive analysis this is the key gap: you end up fighting brands that are not next to you in real answers.
  3. Different sources. The API cites more, but not the same pages: of the domains the API cited for a prompt, only 29% ever appeared in interface answers to that prompt over the following month. Outreach built on that list mostly misses.
  4. It matches Surfer. Surfer's Gemini numbers: API answers 1.27× longer, 1.10× more brands, 6.5 vs 3.4 sources. Ours, on different prompts, in a different language and a different month: 1.27×, 1.10×, 4.9 vs 3.4. Same direction, and the Gemini interface lands at about 3.4 sources per answer in both studies.

Limitations. The sample is small: 47 prompts, 13 brands. The periods are offset by about a month, with news and site changes in between. The API ran Flash-Lite, while Google does not disclose which model the app serves, so the model difference is part of the effect, not pure noise. Our brand and competitor recognition pipeline was also improving during those weeks. Treat these numbers as confirmation of direction, not as correction factors.

We wanted to repeat the comparison for Grok, which moved to the interface on September 18, but few customers monitor Grok, and there were not enough answer pairs. For Grok, the qualitative observation about sub-queries is what we have.

Where the API fits and where it does not

The API is not useless. It is good for learning what a model knows about a brand "from memory", for cheaply testing hundreds of phrasings, and for validating a hypothesis before adding prompts to monitoring. But it answers "what could the model say", not "what did my buyer see".

For GEO work that distinction decides a lot:

  • share of voice and the competitor list from an API describe a different competitive landscape;
  • sources for outreach from an API point to pages users do not see;
  • carousels, ads and cards are invisible to the API entirely;
  • longer API answers with more brands inflate presence: the mention shows up in the report but not on the buyer's screen.

On Claude, to be straightforward: it is the only engine GEO Scout currently reads through the official Anthropic API, with web search enabled and the user's country set. That is closer to the app than a bare model, but it is not claude.ai, and Claude data deserves the same caveat. The other 11 — ChatGPT, Perplexity, Gemini, Google AI Mode, Google AI Overview, Microsoft Copilot, DeepSeek, GigaChat, Grok, Alice AI and Yandex Search with Alice — are captured from the live interface. Details are in the data collection methodology and in why AI visibility matters. Monitoring starts from $24/mo (pricing); the free plan covers up to 5 prompts across 6 AI engines (ChatGPT, Gemini, DeepSeek, Perplexity, GigaChat, Alice AI), refreshed every 7 days.

Checklist: how to tell where a tool gets its answers

Questions for the vendor

  1. API or interface, per engine? Ask for a table, not "we use real answers". A mixed setup is fine as long as it is disclosed.
  2. Which model and mode? For API: which model, whether web search is on, with which settings. For interface: which mode (fast, reasoning), logged in or not.
  3. Where is the request sent from? Country, city, interface language. For local businesses this can change the answer substantially.
  4. Do you capture Alice AI product carousels, Yandex Direct ads and business cards? If none of the three, it is most likely an API.
  5. Which engines return sub-queries? If none, that is another API signal.
  6. Can I open the full stored answer with sources and compare it to my own screenshot?
  7. What happens to empty answers and answers "from memory"? Are they excluded, re-asked, or counted as "brand not mentioned"?

Signals in the data

  • Answers in the report are noticeably longer and more uniform in format than what you see in the app.
  • Answers consistently name more brands than you find in manual checks.
  • Commercial prompts in Alice AI and ChatGPT never show a single card, carousel or ad block.
  • Gemini sources include Google redirect links instead of page addresses.
  • The cited domains in the report barely overlap with what you see yourself.

A three-day mini test

Take 10 prompts and 3 AI engines. For three days in a row, ask them manually in the app from the same city configured for your brand, in a clean browser profile — the approach is covered in alternatives to manual ChatGPT monitoring. Write down the brands and cited domains. Then compare:

  • how much your manual answers agree with each other across days — this is the background noise;
  • how much the tool's answers agree with yours.

You will not get a full match even from the interface against itself: in our Gemini data, about 29% of domains and 59% of competitors. But if the tool agrees with your checks noticeably less than your checks agree with each other, the tool is not seeing the answer your buyer sees.

Частые вопросы

Why does an AI answer from the API differ from the answer in the app?
The API and the app share a model family, but they are different systems. The app runs its own system prompt, its own web search and source selection logic, knows the user's region and language, and adds cards, product carousels and ads on top of the text. Through the API you get the model you picked, with the parameters you set, and none of the interface blocks. So brand lists and sources diverge even for the exact same question.
How large is the gap between API and interface answers?
In Surfer's study (1,000 prompts, 5 AI products, 13,779 answers, collected on August 4, 2026), brand lists overlapped by 15.5–23.8% between API and interface, and by 21.3–31.6% after merging brand name variants. Cited domains overlapped from 4.8% on ChatGPT to 26.7% on Perplexity. In our Gemini comparison on 47 identical prompts, the set of competitors named overlapped by 25% between API and interface, versus 59% for the interface compared with itself on different days.
Is the API useless for AI visibility monitoring?
No. It is fine for answering "what does the model know about my brand": it is cheap and scales well. It does not answer "what does my buyer see". Share of voice, the competitors next to your brand and the sources to target for outreach all come out different through the API. If a tool is API-based, you should know that upfront and read its reports accordingly.
How can I tell whether a monitoring tool uses the API or the live interface?
Ask the vendor for a per-engine table: collection method, model or mode, and request location. Then check the data: do Alice AI answers include Yandex Direct ads, product carousels and business cards, are sub-queries (fan-out) captured, and do answer length and sources look like what you see yourself in the app from the same city. Compare several days, not a single answer, because the interface is noisy on its own.
Why is Claude collected via API in GEO Scout?
Claude is the only one of the 12 engines GEO Scout currently reads through the official Anthropic API, with web search enabled and the user's country set. That is as close to the app as the API gets, but it is not claude.ai itself. We state this in the methodology, and Claude data should be read with the same caveat this article describes.