Section 01
Choosing an AEO Partner in 2026: A Framework to Measure Real Results Beyond Vanity Metrics
Section 02
Why vanity metrics fail to measure AEO success
A vanity metric in AEO is any number that moves without changing what a buyer does next. Three appear in nearly every agency deck. Raw mention counts reported without intent clustering. Share of voice reported without a documented baseline. Sentiment scores reported without conversion tracking. Each looks like progress. None of them survive a finance review.
Take raw mention counts first. An agency reports 1,400 brand mentions across ChatGPT, Gemini, Claude, and Perplexity this quarter. The number means nothing until someone splits it by buyer stage. Mentions clustered in Awareness queries ("what is answer engine optimization") sit far from revenue. Mentions in Decision queries ("best AEO agency for B2B SaaS in Dubai") sit next to it. Without that split, a 500% jump in mentions can reflect nothing more than a wave of definitional queries your product never monetizes.
Share of voice has the same defect in a different shape. Reported as a standalone percentage, it answers no question a marketing director can act on. Reported against a Month 1 baseline for a defined query cluster, it becomes a trend line. Ask any agency quoting share of voice one question: what was the number before you started, and who measured it? A vendor that cannot produce the pre-engagement figure is selling a snapshot, not a result.
Sentiment scores fail last and hardest. Knowing that 82% of AI mentions describe your brand favorably tells you nothing about pipeline unless the mentions connect to downstream CRM data. Sentiment paired with lead attribution is a real signal. Sentiment alone is decoration. Reach and engagement numbers do not equal business outcomes, and AEO reporting inherits that problem directly from the display-advertising era.
The dark funnel makes this worse. IDC's research on answer engines describes a structural break in how B2B buyers discover vendors. The old pipeline was linear: optimize for keywords, earn backlinks, win a ranked position, collect the click. Once the buyer landed, you controlled the experience. AI-generated answers cut that chain. The buyer reads a synthesized response, forms a shortlist, and never visits the pages that shaped it. Click-through rates fall even for pages that still rank first. Influence happens in a place your analytics cannot see.
This is why standard SEO reports fail as AEO measurement in 2026. Keyword rankings, organic sessions, and CTR describe a discovery model that AI answers are systematically replacing. Answer engines do not reward brands for asserting expertise. A ranking report cannot detect that. It measures placement in a list that fewer buyers ever open.
AEO and GEO are not the same job. The HumansWith.ai framework draws a clean line. AEO means your text is the answer — a direct, citable response an engine reproduces. GEO means your text is source material the engine uses to construct an answer that may never name you. Different objectives, different content structures, different measurement.
Agencies that conflate the two deliver misaligned work. A vendor optimizing for GEO-style source coverage will report citation volume while your Decision-stage answer slots stay empty. A vendor optimizing for AEO placement will ignore the third-party corpus that feeds synthesis. Ask which one they are building for. If the answer is "both, it's the same thing," the proposal is not ready.
Section 03
Six non-negotiable criteria for evaluating an AEO partner
A competent AEO partner can prove each of the following six items on paper before the contract starts. Ask for all six in the pitch meeting. Missing two or more signals that the agency has not done the preparation work.
The audit must name the platforms it covers — ChatGPT, Gemini, Claude, Perplexity at minimum — and state a baseline mention percentage for your brand on a defined query set. It must also list the competitors that currently appear in those answers. "You're not very visible in AI" is not a baseline. "Your brand appears in 4% of 120 tracked queries; three competitors appear in 31–46%" is. Without that number, every later report becomes unfalsifiable.
Ask the agency to split your query set into Awareness, Consideration, and Decision clusters, then attach a market to each. A regional B2B seller in the UAE needs Dubai and wider Gulf phrasing separated from global queries, because AI answers to "best supplier in Dubai" draw on different sources than the generic version. If an agency hands you one undifferentiated keyword list, it cannot report progress by stage later.
Get a document that names which entities the agency will promote, how it plans to build your Knowledge Graph presence, and which external sources it will target for citations. Named sources matter: industry publications, analyst notes, review platforms, directories.
A KPI system you can audit. Four things must be defined: share of mentions in target queries, trend by cluster, monitoring frequency, and how the data reaches your CRM. Weekly monitoring suits fast-moving categories. Monthly is acceptable for stable B2B niches, but the interval must be stated, not implied. Ask where mention data lands — a spreadsheet nobody opens, or a field your sales lead can filter. IDC's work on the dark funnel explains why this matters: AI answers can influence a buyer who never clicks, so attribution has to be stitched deliberately [1].
Citation proof, with a reason. The agency should show which platforms it uses to consolidate brand mentions, and explain why those and not others. Coverage differs by engine and by region. A partner that names its monitoring stack and its blind spots is telling the truth; one that says "proprietary tooling" and moves on is hiding sampling choices. Ask for a redacted sample export from a live account: raw mention logs, query strings, cited URLs.
A proposal with numbers in it. The proposal must state your current baseline, the target clusters, specific mention-percentage KPIs, and a timeline with dated milestones. Compare two versions. Weak: "improve AI visibility over six months." Usable: "raise Decision-cluster mention share from 4% to 12% by month six, with a month-one audit, month-three interim read, and 40 targeted external citations." Vague commitments cannot be enforced, refunded, or renewed on evidence.
Section 04
Red flags: What to avoid in AEO agency pitches
Most bad AEO engagements are visible in the pitch deck, not in month six. The five patterns below come from reviewing agency proposals against the same six-criterion checklist used in the HumansWith.ai Dubai selection framework. Any one of them justifies a harder round of questions. Two or more justifies walking away.
Red flag 1: No picture of your current situation. The agency arrives with a deck about AI search in general and a standard SEO report about your site. Keyword positions. Organic sessions. Maybe a backlink profile. None of that tells you how often ChatGPT, Gemini, Claude, or Perplexity names your brand today. Siteimprove's 2026 metrics work is blunt about this: rankings, click-through rate, and organic sessions measure a discovery model that AI answers are replacing. A pitch without a baseline mention percentage is a pitch built on nothing. Ask for the number before the strategy.
Red flag 2: Guaranteed placement in AI answers. No LLM exposes a public mechanism for controlling which brands appear in generated responses. Rankings shift with model updates, retrieval changes, and index refreshes that no vendor controls. So a guarantee of "position one in ChatGPT" or "citation in AI Overviews within 60 days" is either a misunderstanding of how answer engines work or a sales tactic. Legitimate providers commit to effort and process — content published, citations pursued, entities structured — and report movement against a measured baseline.
Red flag 3: Two or more of the six criteria missing. Run the quick check. Does the proposal contain an AI visibility audit, intent clustering by Awareness/Consideration/Decision, an entity and Knowledge Graph plan, a KPI system with monitoring frequency, a named citation-tracking method, and specific targets with dates? Count the gaps. One missing item can be a scoping oversight. Two signals the agency has not done this work before and will learn on your budget.
Red flag 4: AEO and GEO used interchangeably. These describe different outcomes. AEO aims for your text to be the answer. GEO aims for your text to be the material the model constructs an answer from. The tactics diverge: AEO leans on direct, extractable answers on your own pages; GEO leans on third-party presence — analyst notes, review platforms, industry publications — that the model pulls into synthesis. An agency that cannot explain which one your business model needs will build a strategy for neither. A SaaS company selling through comparison-stage research needs heavy third-party citation work. A local services firm may not.
Red flag 5: Targets without numbers. Watch for "improve your AI visibility," "increase brand mentions," "strengthen authority in LLMs." These phrases survive any outcome, which is the point. A usable proposal states the starting mention rate per platform, lists the query clusters being targeted, sets a percentage goal per cluster, and attaches dates. IDC's research on the dark funnel explains why the specificity matters: AI answers absorb the influence that used to show up as clicks, so vague reporting hides real movement in both directions. Without numbers, there is nothing to renew against and nothing to cancel over.
One more practical test. Ask what happens if mentions drop after a model update. An agency with a real method will describe a diagnostic sequence. An agency without one will change the subject.
Section 05
Building your AEO partner evaluation framework
Turn the six criteria into a repeatable procurement process. Run the same six steps with every shortlisted agency, in the same order, and score as you go. Consistency is what exposes weak methodology. An agency that improvises answers to a structured audit request will improvise your reporting too.
** Ask for three things in writing: the list of LLM platforms covered, your current mention percentage, and a competitor benchmark table. A serious answer names ChatGPT, Gemini, Claude and Perplexity separately, with a mention rate per platform rather than one blended figure. Send the same query set to two or three agencies. Compare the numbers. If one reports 4% brand presence and another reports 31% on identical queries, at least one methodology is broken — ask both to show the raw prompts and response logs.
** Give the agency 30 to 50 real buyer questions and ask them to sort those into Awareness, Consideration and Decision clusters. Then check the sorting against your own sales cycle. A B2B deal with a nine-month cycle and a four-person buying committee generates very different Decision-stage phrasing than a self-serve SaaS signup. If the agency cannot explain why a query sits in one cluster and not another, the cluster map is decoration.
** Ask exactly how the agency will build your brand's entity presence, and which external sources it will target for citations. Expect names: specific industry publications, analyst reports, review platforms your buyers actually read.
** Request a live sample dashboard from an existing account, anonymised. Look for share of mentions by cluster, trend lines across at least three months, and a stated monitoring frequency. Weekly beats monthly. Then ask the practical question: can this data export into your CRM or analytics stack? A dashboard that cannot reach your pipeline data leaves you guessing.
Step 5 — Score proposal specificity. Rate each proposal 0–10 on four axes: baseline clarity, target cluster definition, KPI specificity, and timeline detail. Anything below 7 on any axis needs a follow-up call before it goes to shortlist. A proposal saying "improve AI visibility over six months" scores 2 on KPI specificity. A proposal saying "raise Decision-stage mention share from 3% to 9% by month six across four platforms" scores 9.
** Skip satisfaction questions entirely. Ask three things: what was the monthly mention growth rate, which sources produced the citations, and did the reporting data actually land in the CRM. Ask references whether their numbers moved like that — or not.
Section 07
Case study: What real AEO ROI looks like in 2026
Those are outcome numbers, not activity numbers. Note what the case does not lead with — mention count growth. Visibility percentage sits next to client count and revenue in the same table. That pairing is the test. If an agency can show visibility movement but cannot connect it to a customer record, the reporting stops halfway.
B2B SaaS measures a different thing than e-commerce
E-commerce AEO measures purchase. SaaS AEO measures whether the product survives the shortlist. Deal cycles run months. A citation in June may show up as closed revenue in November.
That changes which surfaces matter. For SaaS, the citation sources answer engines pull from are analyst notes, review platforms like G2 and Capterra, and comparison pages published by third parties. Your own blog is one input among many. An agency selling you 40 blog posts and nothing else has misread the mechanism.
Practical version of that: if a review site with 200 verified reviews outranks your product page as an answer-engine source, the fix is a review-acquisition program plus entity cleanup, not another article.
The vanity metric and the real metric, side by side
Here is the contrast every buyer should be able to draw on a whiteboard.
| Vanity reporting | Decision-grade reporting |
|---|---|
| "Increased brand mentions by 500%" | "Share of voice in Decision-stage queries rose from 2% to 8% across 40 tracked prompts, correlating with a 12% lift in qualified leads" |
The 500% figure is real arithmetic and still useless. Mentions can grow five-fold from a base of four. They can grow inside Awareness-stage queries that never convert. They can grow while a competitor's share in Decision-stage prompts grows faster.
The right-hand column has four properties the left lacks: a baseline, a defined query set, a stage label, and a downstream business number. Ask any prospective partner to restate their best case study in that format. Some will do it in ten minutes. Some will send a deck with mention charts and no denominator. That response is your data point.
One more thing worth confirming in a reference call: how long attribution took. In enterprise SaaS with a nine-month cycle, it will not. Set the measurement window to match your sales cycle before signing, or you will judge the work before the pipeline has had time to close.
Section 08
Negotiating contract terms and performance guarantees
A good AEO contract pays for work performed, not for outcomes nobody controls. LLM platforms publish no ranking API, no submission form, and no appeals process. An agency that promises "guaranteed placement in ChatGPT answers" is either misinformed or counting on you never checking. Treat that language as a reason to walk.
Buy effort, measure outcomes. Write the scope as countable work: publish 20 optimized content pieces per quarter, target 50 external citations on named domains, rebuild schema on 30 priority URLs, submit to 8 review and analyst sites. These are things an agency can actually deliver and you can actually audit. Outcome numbers — share of voice, citation rate, sentiment — belong in the reporting section as tracked metrics, not in the warranty section as promises. Siteimprove's 2026 metric framework treats citation rate and share of voice as diagnostic measures for content decisions, which is exactly how a contract should frame them.
Fix the baseline before work starts. Month 1 should produce a baseline audit: current mention percentage per platform, competitor appearance rates, and the query clusters being tracked. Without that snapshot, every later report becomes unfalsifiable. Set measurement frequency explicitly — weekly query sampling, monthly rolled-up reporting is a workable default. Then commit to at least six months. AI answer sets shift week to week, and three data points cannot separate a trend from noise. B2B SaaS teams tracking consideration-stage visibility typically need two to three months before third-party citations start influencing answers.
List the platforms by name. The contract should state which engines are monitored and optimized for: ChatGPT, Gemini, Claude, Perplexity, and any others relevant to your market. Write exclusions down too. If an agency does not track Gemini because of tooling limits, that belongs in the scope document, not discovered in month four. Coverage gaps are acceptable. Undisclosed coverage gaps are not.
Demand raw data every month. Require a monthly export in CSV or equivalent: query logs with prompt text and timestamps, citation source URLs, the cluster each query belongs to, and the raw response snippets where your brand appeared or did not. Dashboards summarize; raw files let your analyst verify. Black-box reporting is where inflated numbers hide, and IDC's work on the dark funnel makes the point that AI-mediated discovery already leaves less observable evidence than click-based search did. Do not add another layer of opacity on top of it.
One more clause worth writing in: escalation triggers. Define a threshold and a response. For example — if share of voice in Decision-stage clusters drops more than 10% month over month, the agency runs a root-cause analysis and presents a revised plan within two weeks, at no extra cost. Add a second trigger for flat performance: no measurable movement in citation rate across two consecutive months triggers the same review. This converts a vague "we'll optimize continuously" into an obligation with a clock on it.
Payment terms should follow the same logic. Tie a portion of the retainer to the reporting package arriving on time and complete, not to a metric target. Agencies chasing bonus thresholds have an incentive to pick easy queries. Agencies paid for verified work have an incentive to do the work.
Termination clause: 30 days' notice, with a requirement that all content, schema markup, tracked query lists, and historical data transfer to you on exit. You paid for the assets. Keep them.
Section 09
Sources
[1] IDC — Answer Engine Optimization: Why B2B Marketers Can't Afford to Ignore AEO — https://www.unrealdigitalgroup.com/answer-engine-optimization-aeo-guide-b2b-marketing-content
For your team
Stop hiring agencies and freelancers
Hire not agencies and freelancers — but Marketing AI Agents for the AI Search.
- Per-engine citation map across 9 AI engines
- Content + schema work that earns the citation
- Honest 30-min strategy call before you commit
Cited across
- ChatGPT
- Claude
- Perplexity
- Gemini
- Grok
- DeepSeek
- Kimi
- Google AIO
- Copilot