Choosing an AI-visibility agency in 2026 is harder than choosing an SEO agency was in 2016. The market is young — nobody has a ten-year track record. Yet the snake-oil pattern is already visible. Every agency website now says "GEO" or "AEO" on the homepage. Few can show you a measurement.
The gap matters more than it did two years ago. Gartner predicted in early 2024 that search-engine volume would drop 25% by 2026 as users move to AI answer engines [1]. By April 2026, Fractl's analysis of over one million keywords found 29% of search volume already in decline — the shift arrived on schedule [2]. SparkToro's 2026 click-stream data shows fewer than one third of Google searches still send a click to the open web [3]. Buyers now ask ChatGPT, Perplexity, or Gemini instead of typing into a search box. The answer they get is the recommendation they act on.
We ended up on both sides of this. We evaluate partners for clients, and we compete in the same category. This is the checklist we use to separate agencies that measure from agencies that sell confidence.
Section 01
What is AI visibility, exactly?
AI visibility (also called GEO — Generative Engine Optimization, or AEO — Answer Engine Optimization) is the practice of raising the odds that AI answer engines cite your brand. The Princeton/Georgia Tech GEO study (Aggarwal et al., 2023) showed that adding citations, quotations, and statistics to content measurably improves its visibility — by up to 40% on their benchmark [4].
The important word is measurably. An agency that cannot say what it will measure, on which queries, and how often, is not doing GEO. It is doing SEO with a new label.
Section 02
Red flags — walk away
"We guarantee placement in ChatGPT answers." Nobody controls a model's output. Anyone guaranteeing citation placement is lying, or is selling you a definition of "placement" you would not recognize. The honest version: "we raise the probability, and we measure the result on a frozen query panel." Run.
No baseline, no panel. If the proposal jumps straight to activity (X articles, Y posts) without first measuring where you stand — citation share-of-voice, which queries matter, which competitors get cited — you are buying activity, not outcomes. A baseline is the only honest starting point.
Traffic metrics only. "We will grow your organic traffic" is an SEO answer to a GEO question. AI visibility is measured in citations, share-of-voice, sentiment, and assisted conversions. An agency that reports only sessions and rankings has not changed its model since 2019.
Proprietary "AI score" with no methodology. Several tools and agencies sell a black-box 0–100 score. If they cannot explain what the score counts, it exists to make the report look technical, not to make decisions better. Ask for the formula. Watch the pause.
Case studies without numbers. "We helped a leading SaaS brand improve its AI presence" tells you nothing. Real case studies show a baseline, a timeframe, and a delta — like "cited in 4 of 30 tracked queries in January, 11 of 30 in March." Vague adjectives are not evidence.
Section 03
Green flags — shortlist
They show you a frozen query panel before they sell you anything. Good shops start by measuring: 20–50 queries your buyers actually ask, your citations, and your competitors'. This takes a day and costs nothing to show. If they will not do it, ask why.
They talk about content structure, not just content volume. Answer-first passages, entity consistency, original data, quotable facts. The GEO research is clear that how a passage is written changes whether engines cite it [4]. If the plan is "more blog posts," you are looking at an SEO agency wearing a GEO hat.
They are honest about what does not work. Ask "what have you tried that failed?" An agency with real experiments has a list. An agency selling a template has a pause.
Sentiment and accuracy are in scope. Getting cited is only half the problem — what the engine says about you matters. Good agencies track whether citations are accurate and favorable, and have a plan for when they are not.
They tell you when you do not need them. For some companies — purely local services, ultra-niche industrial B2B — AI-visibility spend is premature. An agency that says "your money is better spent elsewhere first" is the one to hire later.
Section 04
The five questions that expose the difference in one call
Ask these verbatim. A measurement-first agency answers all five without reaching for a slide deck. A template agency redirects to "our proven process."
- "Show me a before/after measurement from a real client — with the query panel."
- "What exactly do you count as a citation, and how do you dedupe across engines?"
- "What is your plan for the first 30 days — and what will the day-30 report show?"
- "What did you try in the last year that did not work?"
- "How do you handle engines citing a competitor for queries we should own?"
The third question is the sharpest. An agency that has never re-measured a frozen panel cannot describe what a day-30 report looks like, because they have never produced one.
Section 05
What a fair pilot looks like
Do not sign a six-month retainer as a first step. A sane pilot runs 4–6 weeks: baseline measurement on a fixed query set, 3–5 high-leverage content changes, a re-measure, and a report on what moved and why.
You should pay for the pilot — free pilots produce theater, not work. But it should be priced like an experiment, not a retainer. If the pilot doubles your citation share-of-voice, you have evidence to scale. If it does not, you bought a cheap, honest answer. That still beats six months of activity-based invoicing.
Section 06
Why this works: the measurement loop is the product
Here is what most agencies will not tell you. The content is not the hard part. The hard part is discipline: define the panel, measure the baseline, ship structured work, re-measure weekly, and tell the client the truth when nothing moved.
AI answer engines reward a specific set of signals — citations to credible sources, direct quotations, statistical grounding, and answer-first structure [4]. Those signals are repeatable. But they only compound if you re-measure on the same panel. Without a frozen query set, a "before/after" is just two unrelated screenshots. With one, a 0.9%→21.5% swing is a claim you can verify.
That is why we publish our own numbers below. Not because we are unusually good — because this is what evidence looks like.
"If an agency can't show you the query panel and the baseline before they invoice, you're buying confidence, not measurement." — Gregory Shevchenko, founder, Humanswith.AI
Section 07
What measured work actually looks like
The difference between an agency that measures and one that performs is visible in their case studies. Below are our own published cases — every one has a baseline, a locked query panel, a timeframe, and a delta. We publish them because this is the standard we hold ourselves to, and it is the standard you should hold any agency to.
| Client | Baseline | Result | Timeframe |
|---|---|---|---|
| Humanswith.AI | 2 citations | 1,000+ citations, 819 measured AI mentions, 15.4% share-of-model | 3 months |
| Birdview PSA | 0.9% ChatGPT mention rate | 21.5% mention rate, 103 unique queries with presence | 8 weeks |
| GAC auto retailer | 1 AI mention | 9,042 reads, cited on all 9 platforms | 6 weeks |
| CodHob (RU) | 35 citations | 259 citations, 7.4× lift, #2 in niche | 7 weeks |
| Whitewill Dubai | 0 citations / 121 queries | Baseline audit complete, plan scoped | 12-week plan |
Full methodology and all eight cases: AI Visibility Case Hub.
Section 08
FAQ
What is the difference between SEO and GEO?
SEO optimizes for ranked links on a search results page. GEO optimizes for being cited or recommended inside an AI-generated answer. The metrics differ — SEO tracks clicks and positions; GEO tracks citations, share-of-voice, and sentiment across engines like ChatGPT, Perplexity, and Gemini.
How long does it take to see AI-visibility results?
On a fixed panel with weekly measurement, a real signal typically appears in 6–10 weeks. Our Birdview engagement moved ChatGPT mention rate from 0.9% to 21.5% in 8 weeks. Faster claims without a measured baseline are marketing, not measurement.
What should an AI-visibility report actually contain?
A frozen query panel, citation share-of-voice per engine, which queries you appear for, competitor citations on the same panel, sentiment of mentions, and a week-over-week delta. If the report has only traffic and rankings, it is an SEO report.
Can any agency guarantee ChatGPT will recommend us?
No. Nobody controls a model's output. A legitimate agency increases the probability of citation through content structure, entity consistency, and credible-source alignment — then measures the result. A guarantee is the single clearest red flag.
How much should an AI-visibility pilot cost?
A 4–6 week pilot should be priced like an experiment — baseline, 3–5 content changes, a re-measure, and a written report. It should not cost a six-month retainer. Our own measured pilots run in that 4–6 week band. Anything demanding a long lock-in before showing a baseline is structured to bill, not to prove.
We already rank 1 on Google. Do we still need GEO?
Yes, and the gap is widening. SparkToro's 2026 click-stream data shows fewer than one third of Google searches still send a click to the open web [3]. Fractl found 29% of search volume already in decline [2]. Ranking #1 on a results page does not mean you are the brand ChatGPT names when a buyer asks it directly.
Section 09
Sources
[1] Gartner — "Gartner Predicts Search Engine Volume Will Drop 25% by 2026, Due to AI Chatbots and Other Virtual Agents" (press release, February 19, 2024) — https://www.gartner.com/en/newsroom/press-releases/2024-02-19-gartner-predicts-search-engine-volume-will-drop-25-percent-by-2026-due-to-ai-chatbots-and-other-virtual-agents
[2] Fractl — "Search Isn't Dying, It's Redistributing" (April 2026) — across 1,010,848 high-volume keywords, 29% of search volume is in measurable decline, surpassing Gartner's predicted shift — https://www.frac.tl/ai-search-study/
[3] Rand Fishkin / SparkToro — "In 2026, Less than One Third of Google Searches Still Send a Click" (2026 zero-click click-stream study) — https://sparktoro.com/blog/zero-click-search-study-2024/
[4] Aggarwal et al. — "GEO: Generative Engine Optimization" (Princeton / Georgia Tech / IIT Delhi, 2023, arXiv:2311.09735) — citations, quotations, and statistics improve generative-engine visibility by up to 40% — https://arxiv.org/abs/2311.09735
For your team
Stop hiring agencies and freelancers
Hire not agencies and freelancers — but Marketing AI Agents for the AI Search.
- Per-engine citation map across 9 AI engines
- Content + schema work that earns the citation
- Honest 30-min strategy call before you commit
Cited across
- ChatGPT
- Claude
- Perplexity
- Gemini
- Grok
- DeepSeek
- Kimi
- Google AIO
- Copilot