The Complete Enterprise GEO Interview Question Bank: From Core Foundations to Commercial ROI & Attribution

✅ 35 Questions With Model Answers ✅ Built From Live Hiring Panels ✅ Cited 2025–2026 Sources
Enterprise GEO interview question bank covering AI search retrieval, citation metrics, attribution and ROI

🤖 What does an enterprise GEO interview actually test?

An enterprise GEO interview tests five capabilities in sequence: whether the candidate understands how generative engines retrieve and synthesise information (RAG pipelines, embeddings, crawler behaviour), whether they can execute technically (chunking, schema, robots.txt directives for AI agents, rendering), whether they can think strategically about competitive displacement across engines, whether they can connect AI citations to pipeline and defend a budget to a CFO, and whether they can diagnose a live incident such as a citation collapse after a model update. The questions below map to those five layers, with model answers and the weak-answer patterns that should end an interview early.

📌 What this page is — and what it isn't
This is a hiring and interview-prep resource: questions, evaluation signals, model answers, and red flags. It is not a how-to for doing GEO work itself. For the execution layer, these guides go deeper:

I started keeping notes on GEO interviews in early 2025, mostly out of frustration. I was sitting on hiring panels where nobody in the room — including me, at first — could agree on what a good answer even sounded like. One panel would pass a candidate who could recite crawler user-agent strings. Another would reject someone who'd actually moved citation share for a real brand, because they couldn't define "vector embedding" cleanly. We were testing vocabulary, not capability.

This question bank is what came out of fixing that. Thirty-five questions, arranged by seniority, each with the specific signal the question is meant to surface. Some of them are deliberately unfair — Q23 and Q31 in particular have no clean answer, and how a candidate handles that ambiguity tells you more than any of the definitional questions do.

4.4× Conversion rate advantage of AI search visitors over traditional organic visitors — the number every GEO candidate should be able to contextualise, and caveat Source: Semrush AI search traffic study, 2025–2026
~90% Of pages cited by ChatGPT ranked position 21 or worse in conventional Google results — evidence that citation is not a reward for ranking Source: Semrush LLM citation analysis, July 2025
20.3% Share of verified bot traffic coming from AI crawlers, with AI-search bots adding a further 6.5% — crawl policy is now an infrastructure decision Source: Cloudflare crawler data, May 2026

1. How to Use This Question Bank

Don't run all thirty-five. That's a five-hour interview and it will tell you less than a well-chosen twelve. The bank is built so you can pull a vertical slice by seniority and a horizontal slice by the competency you're least sure about.

Role you're hiringSections to pull fromSuggested question countWhat you're really deciding
GEO / AI Search Associate Section 1 in full, 2–3 from Section 2 9–10 Can they learn fast and do they understand the mechanism, not just the buzzwords?
GEO Manager / Senior Specialist 2 from Section 1, all of Section 2, 3–4 from Section 3 12–14 Can they ship changes without breaking the site, and reason about trade-offs?
GEO Lead / Head of AI Search Section 3 in full, 3 from Section 4, 2 from Section 5 12–13 Can they set direction across engines and say no to bad ideas?
Director / VP Organic Growth Section 4 in full, 2 from Section 3, 3 from Section 5 13–14 Can they survive a CFO conversation and own a number?
Agency / consultant vetting Q6, Q9, Q11, Q16, Q22, Q27, Q31 7 Are they selling a dashboard or a measurable outcome?
🗣️ From my experience

The single best predictor I've found isn't knowledge at all — it's whether the candidate volunteers a number without being asked. Ask someone what they did in their last GEO role, and the strong ones say "we took Share of Model from 11% to 34% across 180 prompts on four engines, and self-reported attribution went from 3 deals a quarter to 19." The weak ones say "we improved our AI visibility significantly."

The second predictor is whether they'll admit measurement uncertainty. GEO measurement in 2026 is genuinely messy — answers are non-deterministic, referrers get stripped, and the benchmarks in circulation disagree by an order of magnitude. A candidate who presents it as solved is either inexperienced or selling something.

2. The Four-Band Scoring Rubric

Every question below assumes this rubric. Score each answer 1–4 and weight by seniority — for a Director role, Section 4 answers carry roughly double the weight of Section 1 answers.

BandWhat it sounds likeTypical tell
1 — Vocabulary Repeats industry terms accurately but can't explain the mechanism underneath or what they'd do on Monday. Defines GEO correctly, then can't say how a passage gets retrieved.
2 — Procedure Knows the standard playbook and can execute it. Struggles when the scenario breaks the playbook. Lists tactics in order, but every answer is the same list regardless of the question.
3 — Judgement Reasons from first principles, names trade-offs, quantifies with real numbers from real work. Says "it depends" and then immediately says on what.
4 — Ownership Ties the work to a commercial number, states what they'd stop doing, and describes a time they were wrong. Volunteers the failure case before you ask for it.
A note on band 4: very few candidates hit band 4 on foundational questions, and that's fine — the definitional questions cap out at band 3 by design. Reserve band 4 scoring for Sections 3, 4 and 5, where there's actually room to demonstrate ownership. If a candidate scores 4 on "what is a vector database," you're probably being charmed rather than informed.

3. Section 1 — Foundational Concepts (Beginner / Associate Level)

What this section is for

These seven questions establish whether the candidate understands the machine they're optimising for. You're not looking for textbook precision — you're looking for someone who can describe the path from a user's prompt to a brand mention without hand-waving.

Score generously on phrasing and strictly on mechanism. A candidate who says "the model looks up chunks that are semantically close to the question and writes an answer from them" has understood more than someone who recites the phrase "retrieval-augmented generation" three times without explaining the retrieval part.

Definition: Generative Engine Optimization (GEO)

GEO is the practice of making a brand retrievable, citable and recommendable inside AI-generated answers — in ChatGPT, Google AI Mode and AI Overviews, Perplexity, Gemini, Copilot and Claude. Where SEO optimises for a ranked link and AEO optimises for a single extracted answer, GEO optimises for synthesis: the model reads multiple sources, forms a position, and decides which brands to name.

Q1AssociateFocus: Core definitions

Explain the difference between SEO, AEO and GEO to someone who runs our paid media team. Then tell me where they overlap.

What the interviewer is evaluating

Conceptual clarity, ability to teach a non-specialist, and — most importantly — whether they treat GEO as a replacement for SEO or as a layer on top of it. The overlap half of the question is the real test; anyone can recite three definitions.

Ideal high-scoring answer

"SEO earns a ranked link on a results page — the user still does the clicking and comparing. AEO earns the single extracted answer: the featured snippet, the voice response, the box at the top. GEO earns a mention inside a generated answer, where the model has already read six or eight sources, synthesised them, and is now telling the user which two products to consider."

"The unit of optimisation is what changes. SEO optimises a page. AEO optimises an answer. GEO optimises a passage, plus the consensus about your brand across the wider web — because the model isn't only reading your site, it's reading Reddit, G2, Wikipedia and whatever review roundup happens to rank."

"The overlap is bigger than the discourse suggests. Crawlability, clean HTML, fast rendering, structured data, and topical depth serve all three. Where they genuinely diverge: SEO rewards domain strength and link equity heavily, GEO rewards passage-level clarity and third-party corroboration. Semrush found nearly 90% of ChatGPT-cited pages ranked position 21 or worse in Google — so citation isn't a prize for ranking, and you can't assume your SEO wins transfer."

Red flags / weak answer signals
  • "SEO is dead, GEO is the new SEO." Nobody who has looked at a 2026 analytics account believes this.
  • Treats AEO and GEO as identical, or can't articulate any difference between an extracted answer and a synthesised one.
  • Describes GEO purely as "writing content for AI" with no mention of retrieval, structure or off-site signals.
  • Can't name a single thing that serves all three disciplines at once.
Q2AssociateFocus: RAG architecture

A user types "best enterprise password manager for a 500-person company" into Perplexity. Walk me through everything that happens before our brand does or doesn't appear in that answer.

What the interviewer is evaluating

Whether they understand retrieval-augmented generation as a pipeline with distinct, separately-optimisable stages — and whether they know which stages they can influence. This question separates people who've read about RAG from people who've thought about where the leverage sits.

Ideal high-scoring answer

"The prompt gets interpreted and usually decomposed — 'best enterprise password manager for 500 people' likely fans out into several sub-queries: enterprise password manager comparison, SSO and SCIM support, per-seat pricing at mid-market volume. Perplexity runs those against its own index and live search."

"Retrieval happens at passage level, not page level. Candidate chunks get embedded and compared against the query embedding, so what's competing isn't my homepage against a competitor's homepage — it's a 90-word passage from my docs against a 90-word passage from a G2 comparison page. Then there's a re-ranking pass that weighs relevance, recency, and source authority."

"The top-ranked chunks go into the context window with an instruction to answer and cite. The model synthesises, and the answer reflects both what was retrieved and what the model already believes from pre-training. That's why brand consensus matters — if the model's priors say the category leaders are X and Y, my chunk has to be unambiguous enough to break that."

"Where I can act: retrievability (can the agent fetch and parse the page at all), chunk quality (is the passage self-contained and specific), corroboration (do third-party sources say the same thing), and freshness signals. What I can't act on: the ranking function, the decomposition logic, or the model's pre-training weights."

Red flags / weak answer signals
  • Describes it as "the AI searches Google and summarises the top 10" — a model of retrieval that's several years out of date.
  • No mention of chunking or passage-level competition; talks only about pages and domains.
  • Doesn't distinguish between what's retrieved at query time and what the model absorbed in pre-training.
  • Claims to be able to influence the ranking algorithm directly, or hints at manipulating the retrieval layer.
Q3AssociateFocus: Embeddings & chunking

What is a vector embedding, and why does the way we structure a page affect whether it gets retrieved?

What the interviewer is evaluating

Whether the technical vocabulary connects to an editorial decision. The answer I want ends with a content instruction, not a definition. Plenty of candidates can define embeddings; far fewer can tell a writer what to change on Monday because of them.

Ideal high-scoring answer

"An embedding turns a block of text into a list of numbers that positions it in a semantic space, so text about similar concepts lands close together regardless of whether it shares any keywords. A vector database stores those positions and lets a system find the nearest matches for a query embedding very fast."

"The practical consequence is that retrieval operates on chunks, and a chunk gets embedded on its own — with none of the surrounding context. So if a section header says 'Pricing' and the paragraph under it starts 'It costs $12 per seat,' that chunk is semantically weak. Nothing in it says what 'it' is. Rewrite it as 'Acme Vault costs $12 per user per month on the Business plan, with volume discounts above 250 seats' and the same information becomes retrievable on its own."

"So the instructions I give writers are: make every section self-contained, restate the subject rather than leaning on pronouns, keep one idea per passage, and put the direct answer in the first 40–60 words of a section before the nuance. That's also why a wall-of-text page underperforms a well-sectioned one even when the information is identical — bad chunk boundaries split a good answer in half."

Red flags / weak answer signals
  • Pure definition with no link to content structure or an editorial rule.
  • Confuses embeddings with keyword matching — "it's basically synonyms."
  • Suggests stuffing pages with entity names to "improve the vectors."
  • Never mentions that chunks are evaluated without surrounding context.
Q4AssociateFocus: Knowledge graphs & entities

Our brand is described inconsistently across the web — three different founding dates, two different category descriptions. Why does that matter to a generative engine, and what would you fix first?

What the interviewer is evaluating

Entity thinking. Does the candidate understand that models resolve a brand to an entity and then reason about it — and that conflicting facts don't average out, they produce hedged or wrong answers? Also tests prioritisation: which source do you fix first when you can't fix everything?

Ideal high-scoring answer

"Generative engines resolve a brand name to an entity — a node with attributes like category, founding date, headquarters, leadership, products — and then answer from those attributes plus retrieved text. When sources conflict, you get one of three bad outcomes: the model hedges, the model picks the most-repeated version regardless of whether it's correct, or the model declines to make a claim about you at all. In a comparison prompt, hedging is effectively losing."

"I'd fix in order of corroboration weight. First the sources models lean on most heavily for entity facts: Wikipedia and Wikidata if the brand qualifies, Crunchbase, LinkedIn, the company's own About page with Organization schema carrying foundingDate, sameAs and description. Then the high-traffic third-party profiles — G2, Capterra, industry directories. Then press and partner pages."

"The goal is one canonical sentence about what the company is, repeated verbatim wherever we control the copy. Not five well-written variations — one sentence. Consistency beats elegance here. I'd pair it with entity-level schema and sameAs linking so the association is explicit rather than inferred."

Red flags / weak answer signals
  • Only talks about on-site fixes; doesn't mention third-party sources at all.
  • No concept of entity resolution — treats the brand name as a keyword.
  • Proposes rewriting everything at once, with no prioritisation logic.
  • Suggests editing Wikipedia directly as a brand representative, without acknowledging the conflict-of-interest rules.
Q5AssociateFocus: Crawler mechanics

Name the AI crawlers you'd expect to see in our server logs, and tell me what blocking each one would actually cost us.

What the interviewer is evaluating

This is the highest-consequence factual question in Section 1. Getting it wrong in production removes a brand from live AI answers. I want to know whether the candidate distinguishes training crawlers from retrieval agents from user-initiated fetches — because the business impact of blocking each is completely different.

Ideal high-scoring answer

"Three functional categories, and the distinction is what matters. Training crawlers: GPTBot, ClaudeBot, CCBot, Meta-ExternalAgent, Applebot-Extended. Blocking these limits whether your content shapes future model weights — a slow, compounding cost with no immediate traffic impact."

"Live retrieval agents: OAI-SearchBot for ChatGPT Search, PerplexityBot, and Google-Extended, which governs Gemini and AI Overview usage without affecting Googlebot indexing. Blocking these is the expensive mistake — you disappear from answers immediately, not eventually."

"User-action fetchers: ChatGPT-User, Claude-User, Perplexity-User. These fire when a person explicitly asks the assistant to open your page. Blocking them means a prospect who's actively trying to read your pricing page gets an error, which is about the worst outcome on the list."

"I'd also flag that this is now an infrastructure conversation, not just a visibility one. Cloudflare's May 2026 data put AI crawlers at 20.3% of verified bot traffic with AI-search bots adding 6.5% more, and GPTBot alone at around 11.5% of AI bot requests. On a large site that's real origin load, so the policy discussion involves engineering, not just marketing. I'd write the policy per-agent in robots.txt, verify with reverse DNS rather than trusting the user-agent string, and audit the WAF and bot-mitigation rules separately — most 'we're blocked' incidents I've seen were a bot-management rule, not robots.txt. The full AI crawler directive reference covers the per-agent syntax."

Red flags / weak answer signals
  • Treats all AI bots as one category — "we allow AI crawlers" with no per-agent distinction.
  • Thinks blocking Google-Extended affects Google Search rankings.
  • Doesn't know that robots.txt is advisory and that enforcement lives at the WAF/CDN layer.
  • Never mentions server logs or verification; assumes the user-agent header is trustworthy.
Q6AssociateFocus: Core metrics

Define Share of Model and Share of Recommendation. How would you actually measure them for us next week?

What the interviewer is evaluating

Measurement discipline. The definitions are easy; the methodology is where people fall apart. I'm listening for sampling, repetition, and an acknowledgement that these are estimates with error bars — not counts.

Ideal high-scoring answer

"Share of Model is the percentage of runs, across a fixed prompt set, where the brand appears at all. Share of Recommendation is the narrower subset where the brand is actively recommended rather than just mentioned in passing — 'Acme is one of several options' counts for SoM but not SoR."

"To measure it, I'd build a prompt set of 150–200 prompts covering the real buyer journey: category discovery, comparison, objection handling, alternatives-to-competitor, and implementation questions. Freeze that set. Then run each prompt on each engine at least five times, because answers are non-deterministic — one run tells you nothing. Log mention, recommendation, position in the answer, sentiment, and which sources were cited."

"Report a rolling four-week average with a confidence band, not a point estimate. And I'd always report the cited-source list alongside the score, because that's the actionable half — knowing that we're absent from 40% of answers is interesting, knowing that a specific Reddit thread and two G2 roundups are being cited instead is a work plan."

"For week one specifically: I'd want the existing prompt set if one exists, a baseline run before I change anything, and clarity on whether we're buying a tool or scripting it. Citation rate as a standalone metric is worth defining up front too, so we're not arguing about denominators in month three."

Red flags / weak answer signals
  • Proposes a single run per prompt, or doesn't mention non-determinism at all.
  • No fixed prompt set — "I'd just ask ChatGPT some questions about the category."
  • Can't distinguish a mention from a recommendation.
  • Names a vendor tool as the entire answer, with no methodology of their own.
  • Reports the metric without the cited-source detail, which makes it undiagnosable.
Q7AssociateFocus: Citation types

We get cited directly in 12% of target answers. In another 30%, the answer recommends us but cites a third-party review site instead of us. Which number matters more, and why?

What the interviewer is evaluating

Whether the candidate optimises for vanity (our link is in the answer) or outcome (the buyer chooses us). It also surfaces how they think about the trade-off between traffic and influence — which is the central tension in the whole discipline.

Ideal high-scoring answer

"The 30% matters more commercially, and the 12% matters more operationally. Indirect citation — where the model recommends us but sources a third party — means the buyer still leaves with our name and usually still converts; they'll search the brand or go direct. Direct citation gives us the click and, critically, the measurable session."

"So I'd treat them as two different jobs. Indirect citation share is the brand-influence metric and it drives the off-site programme: review platforms, community presence, comparison content on sites we don't own. Direct citation share is the traffic and attribution metric and it drives the on-site programme: passage structure, schema, freshness, retrievability."

"The reason not to collapse them into one number: if you only chase direct citations, you'll under-invest in exactly the third-party consensus that determines whether the model recommends you at all. And if you only chase indirect, you'll have no measurable traffic and a very hard budget conversation. I'd report both, with branded search lift as the bridge metric that shows indirect citation is doing something real."

Red flags / weak answer signals
  • Dismisses indirect citations as worthless because "there's no click."
  • Treats the two as the same metric, or can't explain why you'd measure them separately.
  • No mention of branded search or any bridge metric linking influence to measurable behaviour.
  • Answers purely in traffic terms, with no reference to the buying decision.

4. Section 2 — Technical & On-Page GEO (Intermediate Level)

What this section is for

Section 1 tests understanding. Section 2 tests execution: can this person open a codebase, a CMS and a robots.txt file and make changes that hold up under review? These eight questions are where most mid-level candidates separate — the playbook answers are easy to spot, and so are the people who've actually shipped.

Push for specifics relentlessly here. "We added schema" is not an answer. "We added FAQPage to 340 support articles and saw citation rate on those URLs move from 4% to 17% over six weeks, while the control group didn't move" is an answer.

Q8IntermediateFocus: Content structuring

Here's a 3,000-word product page that gets good Google traffic and zero AI citations. Restructure it. What specifically changes?

What the interviewer is evaluating

Concrete editorial judgement. Can they translate retrieval mechanics into a document outline? I also want to see whether they preserve the existing SEO performance while restructuring — a candidate who'd happily torch a page that's ranking well has a risk problem.

Ideal high-scoring answer

"First I'd check what the page is actually being asked to do, then restructure without changing the URL, title or primary heading — the Google performance is an asset I'm not going to gamble."

"The changes, in order of impact:

  • Answer-first sections. Every H2 gets a 40–60 word direct answer immediately beneath it, before any narrative. That's the chunk most likely to be lifted.
  • Self-contained passages. Replace pronoun chains and 'as mentioned above' references. Each section restates its subject so it survives being read in isolation.
  • Information density. Strip the throat-clearing. Models reward specificity — numbers, named constraints, versions, limits. 'Scales well' becomes 'tested to 40,000 concurrent sessions per node.'
  • Modular chunking. Break the 3,000 words into 8–12 sections of 200–400 words with descriptive headings phrased the way people ask. 'Pricing' becomes 'How much does Acme cost per user?'
  • Extractable comparison data. Convert prose comparisons into a real HTML table. Tables chunk cleanly and get quoted.
  • Schema. Product, Offer, FAQPage for the Q&A block, with values that match the visible text exactly.
  • Prompt alignment. Add sections that answer the questions buyers actually type into an assistant — 'Acme vs [competitor] for teams over 200', 'does Acme support SCIM provisioning' — which the current page doesn't cover because nobody searches those as keywords.

"Then I'd measure it as a cohort: this page plus 20 similar ones restructured, 20 comparable ones left alone, citation rate tracked on both for six weeks. If the treated cohort doesn't move, my model of the problem is wrong and I'd rather find that out on 20 pages than 2,000."

Red flags / weak answer signals
  • "Add an FAQ section and schema" as the whole answer.
  • Proposes rewriting from scratch with no regard for existing rankings.
  • No mention of passage self-containment or answer-first structure.
  • Recommends adding word count. More words is almost never the fix.
  • No measurement plan, or a plan with no control group.
Q9IntermediateFocus: Structured data

Which Schema.org types actually earn their implementation cost for GEO, and which are theatre? Defend your list.

What the interviewer is evaluating

Whether they can prioritise under constraint and whether they understand what schema does for a generative engine versus what it does for a search engine. Bonus signal: do they know schema is a disambiguation aid, not a ranking lever?

Ideal high-scoring answer

"Schema doesn't make a model cite you. It makes the facts on your page unambiguous, which reduces the chance the model mis-attributes or skips them. So the types worth the cost are the ones that resolve ambiguity about entities, claims and provenance."

"High value: Organization with sameAs and a stable @id — this is the entity anchor for everything else. Product plus Offer for anything transactional, because price and availability are exactly the facts models get wrong. FAQPage where the Q&A genuinely exists on the page. Article with real author, datePublished and dateModified — provenance and freshness both feed re-ranking. HowTo for procedural content. Dataset if you publish original research, which is disproportionately citable."

"Lower value or theatre: speakable beyond a couple of selectors, deeply nested BreadcrumbList on a flat site, Review markup on self-hosted testimonials, and the habit of marking up everything WebPage with no relationships between nodes. A disconnected @graph where @id references don't resolve is worse than no schema — it signals sloppiness and the parser drops nodes silently."

"The non-negotiable rule: schema must match visible content exactly. Marked-up prices that differ from displayed prices is the fastest way to get your structured data distrusted. The full schema implementation guide covers the @graph patterns I'd use."

Red flags / weak answer signals
  • "Add all the schema you can" — no prioritisation, no cost awareness.
  • Claims schema is a direct ranking or citation factor.
  • Recommends marking up content that isn't visible on the page.
  • Doesn't know what an @graph is or why @id references need to resolve.
Q10IntermediateFocus: Crawl policy

We have a public marketing site, a gated resource library, and paid research reports behind a paywall. Write me the AI crawler policy. What goes in robots.txt, and what doesn't belong there at all?

What the interviewer is evaluating

Practical judgement about the training-versus-retrieval trade-off, plus an understanding that robots.txt is a request and not a security control. This is a question where "it depends" is the correct opening, as long as they say on what.

Ideal high-scoring answer

"Different content, different policy. The marketing site should be fully open to everything — training and retrieval both. That content exists to be repeated. The gated library: I'd allow retrieval agents on the landing pages and abstracts, because those are what should surface in an answer, and disallow the asset paths. The paid research: disallow training crawlers, and enforce it at the application layer, because robots.txt only stops well-behaved agents."

# Marketing site — fully open
User-agent: GPTBot
Allow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: Google-Extended
Allow: /

# Gated assets and paid research
User-agent: GPTBot
Disallow: /research/full/
Disallow: /downloads/

User-agent: ClaudeBot
Disallow: /research/full/
Disallow: /downloads/

Sitemap: https://example.com/sitemap.xml

"What doesn't belong in robots.txt: anything you actually need enforced. Paywalled PDFs need auth, not a directive. And I'd never disallow a path in robots.txt as a way of hiding it — you've just published a list of interesting URLs."

"I'd also raise the commercial question rather than deciding it alone: for a research business, allowing training crawlers on premium content may be giving away the product. That's a licensing conversation for legal and the content P&L, not a technical SEO decision. My job is to make sure whoever decides it understands that blocking retrieval agents removes us from live answers, while blocking training crawlers costs us slowly and invisibly. I'd also publish an llms.txt as a pointer to our best canonical sources — it's low cost, and adoption is uneven but growing."

Red flags / weak answer signals
  • Blanket allow or blanket block with no segmentation by content type.
  • Treats robots.txt as a security mechanism for paid content.
  • Blocks retrieval agents alongside training crawlers without flagging the visibility cost.
  • Makes the licensing call unilaterally instead of escalating it.
  • Presents llms.txt as a universally-honoured standard rather than an emerging convention.
Q11IntermediateFocus: Rendering & delivery

How would you prove whether an AI agent can actually see the content on our React-based pricing page?

What the interviewer is evaluating

Diagnostic method. Anyone can say "AI crawlers don't run JavaScript." I want the test they'd run, and I want to know whether they'd verify rather than assume. This is also the question where a hands-on candidate is instantly distinguishable from a strategy-deck candidate.

Ideal high-scoring answer

"I'd test rather than theorise. Fetch the URL with JavaScript disabled and with the relevant user agents — curl with a GPTBot or PerplexityBot UA string — and diff the raw HTML against the rendered DOM. If the price only exists in the rendered DOM, we have a problem for every agent that doesn't execute JS, which is most of them outside Google's pipeline."

"Second test: ask the engines directly. Prompt ChatGPT and Perplexity with 'what does [product] cost' and see what they return. If they return an outdated price, a competitor's price, or nothing, that's live confirmation. It's crude, but it's the ground truth we're actually optimising for."

"Third: server logs. Are GPTBot, OAI-SearchBot, PerplexityBot and ClaudeBot hitting the URL at all, and what status codes are they getting? I've found more 403s from bot-mitigation rules than genuine rendering problems. Rate limiting that returns 429 to AI agents is another common one and it's invisible in every SEO tool."

"The fix ladder, cheapest first: server-side render or statically pre-render the commercially critical facts; if the framework makes that hard, inject the key facts into the initial HTML payload or the JSON-LD; and as a stopgap, make sure a plain-HTML canonical version of the pricing data exists somewhere crawlable. I'd resist building a separate markdown-delivery layer for bots unless there's a clear case — it's close enough to cloaking to make me uncomfortable, and it doubles the maintenance surface. The JavaScript SEO guide covers the rendering trade-offs in more depth."

Red flags / weak answer signals
  • Asserts that AI crawlers can't render JS without proposing any verification.
  • Never mentions server logs or status codes.
  • Proposes serving different content to bots than to users with no hesitation about cloaking.
  • Jumps straight to "we need to rebuild the site in Next.js" as the first answer.
Q12IntermediateFocus: Third-party footprint

The models keep citing a Reddit thread and two review roundups that don't mention us. You have 90 days and no paid placement budget. What's the plan?

What the interviewer is evaluating

Whether they understand that most GEO leverage sits off-site — and whether they'd do it ethically. This question has a trapdoor: a candidate who proposes astroturfing Reddit has just told you they'd put the brand at risk.

Ideal high-scoring answer

"Start by mapping the actual citation graph — for the 40 prompts that matter most, which domains get cited and how often. That gives a ranked target list instead of a guess. In most B2B categories I've mapped, a small number of sources carry a disproportionate share: Reddit, one or two review platforms, a couple of independent blogs, and the category's Wikipedia page."

"Then three workstreams over 90 days:

  • Review platforms (weeks 1–6). Structured review generation on G2, Capterra or the vertical equivalent — a real campaign to existing happy customers, not incentivised fakes. Volume and recency both matter, and review text is heavily quoted in comparison answers.
  • Community presence (ongoing). Real accounts, disclosed affiliation, answering questions in the subreddits and forums where our buyers already are. Slow, and the only version that works. Reddit's weight in citations makes this high-leverage, and it's also the workstream most likely to backfire if faked.
  • Earned inclusion in roundups (weeks 2–12). Find the comparison and 'best X tools' pages that are actually being cited, and pitch the authors with something genuinely useful — original data, a free tool, a correction to outdated information about the category. Most of those pages are maintained and their owners want accuracy.

"I'd add one asset of our own: a piece of original research with numbers nobody else has. Original data gets cited across all of the above and is the only content type I've seen reliably pull citations without a distribution budget."

"What I would not do: create fake accounts, buy reviews, or pay for undisclosed placements. Beyond the ethics, platform enforcement is getting better and a brand caught doing it loses the community channel permanently."

Red flags / weak answer signals
  • Any proposal involving fake accounts, incentivised reviews, or undisclosed affiliation.
  • Focuses entirely on owned content, ignoring that the cited sources aren't ours.
  • No method for identifying which third-party sources actually get cited.
  • Treats Reddit as a place to post promotional content rather than participate.
Q13IntermediateFocus: Freshness pipelines

We have 4,200 indexable pages. How do you keep them fresh enough for retrieval without a content team of thirty people?

What the interviewer is evaluating

Systems thinking and prioritisation at scale. Also whether they know the difference between genuine freshness and date-stamp manipulation — and whether they understand that constant rewriting can destabilise citations that are already working.

Ideal high-scoring answer

"Freshness isn't a site-wide property, it's a per-page decision, so the first job is segmentation. I'd bucket the 4,200 into three: time-sensitive pages where facts decay (pricing, integrations, anything with a year in it, regulatory content), stable reference pages where the information doesn't change, and low-value pages that should probably be pruned rather than maintained."

"Then build a decay detector rather than a calendar. The signals I'd monitor: citation rate dropping on a URL that previously got cited, competitor content published more recently on the same query cluster, internal data changing — a price change or feature launch that should have triggered updates on 30 pages and didn't — and Search Console impression decline. That's a queue, ranked, refreshed weekly. It usually surfaces 20–40 pages a month, which one or two people can handle."

"Two guardrails. First, dateModified only moves when the content actually changed materially — bumping dates on untouched pages is a trust risk with both search and AI systems, and it destroys your own ability to correlate changes with outcomes. Second, don't rewrite a page that's currently being cited unless you have a specific reason. I've watched a team lose a stable citation position by 'improving' a page and accidentally deleting the exact passage that was being quoted. Diff before and after, and track which passage is being lifted."

"For the low-value bucket, pruning often does more than refreshing. Thin pages dilute topical clarity and consume crawl budget that AI agents also spend."

Red flags / weak answer signals
  • Proposes updating all 4,200 pages on a rolling schedule with no prioritisation.
  • Suggests bumping dateModified without content changes.
  • No detection mechanism — freshness handled purely by calendar.
  • Doesn't consider that rewriting can break existing citations.
  • Proposes mass AI-generated refreshes with no human review step.
Q14IntermediateFocus: Prompt alignment

Our keyword research says "project management software" gets 165,000 searches a month. Why is that number close to useless for GEO planning, and what would you use instead?

What the interviewer is evaluating

Whether they've genuinely updated their research methodology or just renamed keyword research. The good answer acknowledges that prompt volume data barely exists and describes how to work around that honestly.

Ideal high-scoring answer

"That number tells us about a query pattern that doesn't exist in an assistant. Nobody types 'project management software' into ChatGPT — they type 'we're a 40-person agency drowning in client work, Asana feels like overkill, what should we actually use.' The volume metric measures the wrong shape of input."

"What I'd use instead, in order of reliability:

  • Sales call transcripts. The single best prompt source I've found. The questions prospects ask on discovery calls are almost verbatim the prompts they type. Mine Gong or Chorus for question phrasing.
  • Support tickets and chat logs. Same logic, post-purchase, and they surface the objection-handling prompts.
  • Long-tail conversational queries in Search Console. Question-form queries with low volume are a decent proxy for prompt phrasing.
  • Reddit and community threads. People ask the internet in roughly the same voice they ask an assistant.
  • The engines themselves. Ask an engine what follow-up questions people typically have about the category, then use its own expansion as the cluster.

"Then cluster by decision stage rather than by keyword: discovery, shortlisting, comparison, objection, implementation. That cluster map becomes both the content plan and the measurement prompt set, which is a nice property — you're measured on the same questions you optimised for. Conversational query research goes deeper on the clustering method."

"I'd be honest with stakeholders that prompt volume is estimated, not measured. Anyone selling you precise prompt volume data in 2026 is modelling, not counting."

Red flags / weak answer signals
  • Treats GEO research as ordinary keyword research with longer keywords.
  • Cites a tool's "prompt volume" figure as hard data with no caveat.
  • Never mentions first-party sources — sales calls, support, CRM.
  • No clustering logic, just a longer keyword list.
Q15IntermediateFocus: Technical audit

Walk me through the first eight checks of a GEO technical audit. What's on the list that wouldn't be on a traditional SEO audit?

What the interviewer is evaluating

Audit discipline and whether their mental checklist has genuinely evolved. The second half of the question matters most — if their GEO audit is an SEO audit with schema added, they haven't done this work.

Ideal high-scoring answer

"In order, because each one can invalidate the next:

  • 1. Agent access. robots.txt per-agent directives, plus WAF, CDN and rate-limit rules. Verified against server logs, not assumed.
  • 2. Status codes for AI user agents. Are GPTBot, PerplexityBot, ClaudeBot and OAI-SearchBot getting 200s on commercially important URLs?
  • 3. Raw-HTML content parity. Diff raw HTML against rendered DOM on the top 50 revenue pages.
  • 4. Passage structure. Sample 30 pages and check for answer-first sections, self-contained passages, and sane chunk boundaries.
  • 5. Entity consistency. Brand facts across owned properties, schema, and the top third-party profiles. Conflicts logged.
  • 6. Structured data validity. @graph integrity, unresolved @id references, mismatches between markup and visible text.
  • 7. Citation baseline. Fixed prompt set run across four engines, five runs each, with the cited-source list captured.
  • 8. Third-party citation graph. Which external domains are winning our answers, and whether we appear on them.

"Items 1, 2, 3, 7 and 8 either don't appear on a traditional SEO audit or appear in a very different form. And I'd add speed as a shared concern rather than a GEO-specific one — retrieval agents time out, and a slow origin means a fetch that never completes. Core Web Vitals work pays into both disciplines, though for AI agents it's time-to-first-byte and server response that matter far more than layout shift."

"The order is deliberate. There's no point auditing passage structure on pages the agent is getting a 403 on."

Red flags / weak answer signals
  • Recites a standard SEO audit checklist with "add schema" bolted on.
  • No server-log or access-verification step anywhere in the list.
  • No citation baseline — audits the site but never checks the engines.
  • Applies Core Web Vitals reasoning to AI agents as if they were browsers rendering layouts.

5. Section 3 — Advanced Strategy & Competitive Displacement (Senior / Lead Level)

What this section is for

Senior GEO work is mostly about sequencing and refusal: deciding what to do first when everything looks urgent, and saying no to the tactics that would work briefly and cost you later. These seven questions have no single right answer, which is the point — you're scoring the reasoning, not the conclusion.

Ask follow-ups aggressively here. "What would change your mind?" and "what's the version of this that fails?" are the two most useful probes in the whole bank.

Q16SeniorFocus: Prompt reverse engineering

Reverse engineer the prompt set for our category. Talk me through your method, and tell me how you'd know when the set is good enough to stop.

What the interviewer is evaluating

Methodological rigour and a stopping rule. Candidates who can't say when they'd stop tend to build 800-prompt monitoring sets that cost a fortune to run and produce noise instead of signal.

Ideal high-scoring answer

"I build it in four passes. Pass one is extraction from first-party sources — sales transcripts, support tickets, won/lost interview notes, the search bar on our own site. Pass two is expansion: take each extracted question and generate the realistic variants a buyer would phrase, including the ones that name competitors, because those are the highest-commercial-value prompts and they never come from keyword tools."

"Pass three is journey mapping. Every prompt gets tagged with a decision stage and an intent type: discovery, shortlist, comparison, objection, validation, implementation. I want coverage across all of them, weighted toward the stages where a model's recommendation actually changes the outcome — which in most B2B categories is shortlisting and comparison, not discovery."

"Pass four is engine-driven expansion. Run the set, look at what follow-up questions the engines themselves suggest, and add the ones that represent real journeys. This catches phrasing we'd never have guessed."

"The stopping rule: saturation. When a new pass of 30 prompts produces no new cited domains and no new competitor set, the map is stable. Practically that's landed between 120 and 220 prompts in every category I've done it for. Then I split it — a stable core of about 60 for weekly monitoring, and a rotating remainder run monthly. Monitoring all 200 weekly across four engines with five runs each is 4,000 API calls a week for marginal extra information."

"I'd also version the prompt set and treat changes to it as a measurement event. Adding 20 prompts mid-quarter will move Share of Model for reasons that have nothing to do with our work, and if you don't annotate it, you'll spend a week explaining a fake trend."

Red flags / weak answer signals
  • No stopping rule — would keep adding prompts indefinitely.
  • Prompt set built entirely from keyword tools with no first-party input.
  • No stage or intent tagging, so every prompt is weighted equally.
  • Doesn't version the set or account for how changing it distorts trend data.
  • Ignores competitor-named prompts, which are usually the commercially decisive ones.
Q17SeniorFocus: Competitive displacement

A well-funded incumbent is named in 80% of answers for our category. We're named in 9%. How do you close that, and how long does it take?

What the interviewer is evaluating

Strategic realism. There's a right answer here and it starts with "not by attacking them head-on." I'm also listening for a defensible timeline — a candidate who promises parity in a quarter is either naive or managing up.

Ideal high-scoring answer

"Head-on displacement on generic category prompts is the slowest, most expensive path, because the incumbent's position is built on years of accumulated third-party consensus that I can't outspend in a quarter. So I'd flank."

"First, segment the prompt set by contestability. Generic prompts — 'best CRM' — are where the incumbent is strongest. Qualified prompts — 'best CRM for a 30-person field services team that needs offline mobile access' — are where the model has weaker priors and has to rely on retrieved text. That's where we can win in weeks rather than years. In every displacement programme I've run, 60–70% of early wins came from qualified prompts."

"Second, find the incumbent's weak flank by reading what the model says about them. If the answers consistently mention 'steep learning curve' or 'expensive for smaller teams,' that's the model reporting the consensus — and it tells me exactly which comparison content and which customer proof points to build. I'd build genuinely useful comparison content that concedes what they're better at, because models cite balanced sources more readily than promotional ones, and an obviously one-sided comparison page reads as marketing to both readers and re-rankers."

"Third, attack the citation graph rather than the answer. If three review roundups and a Reddit thread drive 40% of the citations in this category, getting accurately represented in those four places moves more than a hundred pages on our own site."

"Timeline: measurable movement on qualified prompts in 6–10 weeks. Meaningful movement on generic category prompts in 9–18 months, and only if brand-building happens alongside — because at that level it stops being a GEO problem and becomes a brand authority problem. I'd rather commit to 9% → 35% on qualified prompts this year than promise category parity and miss."

Red flags / weak answer signals
  • Promises to displace an entrenched incumbent on generic prompts within a quarter.
  • No segmentation between contestable and uncontestable prompts.
  • Proposes only on-site content as the lever.
  • Suggests negative or disparaging content about the competitor.
  • Never mentions reading what the model already says about the incumbent.
Q18SeniorFocus: Multi-engine nuance

We're strong in Perplexity and invisible in ChatGPT. What's your hypothesis, and how does your optimisation differ by engine?

What the interviewer is evaluating

Whether they've actually worked across engines or treat "AI search" as one monolith. The specific failure mode I'm probing for: someone who optimises for Perplexity's live-retrieval behaviour and assumes it transfers to a model leaning harder on pre-training.

Ideal high-scoring answer

"The most likely explanation is that we've optimised for retrieval and not for consensus. Perplexity is retrieval-heavy — it searches, cites transparently, and rewards fresh, well-structured, crawlable pages. That's a solvable technical problem and it responds within weeks. ChatGPT leans more on pre-training plus selective retrieval, so being named there depends much more on how widely and consistently the brand is discussed across the corpus. You can't rewrite your way into pre-training weights."

"So the diagnosis path: check whether OAI-SearchBot and ChatGPT-User can actually reach us — access issues explain more 'invisibility' than people expect. Then test whether we appear when we force retrieval by naming our brand in the prompt. If ChatGPT describes us accurately when asked directly but never surfaces us unprompted, that's a consensus problem, not a technical one, and the fix is third-party footprint and time."

EngineRetrieval behaviourWhat moves the needleRealistic response time
PerplexitySearch-first, transparent inline citations, heavy community-source weightingCrawlability, passage structure, freshness, presence on Reddit and review platforms2–6 weeks
Google AI Mode / AI OverviewsGrounded in Google's index; strong correlation with conventional rankingClassic SEO strength plus extractable passages and structured data4–12 weeks
ChatGPT SearchPre-training priors plus selective live retrievalBreadth of third-party mention, entity consistency, being named in widely-copied sources3–9 months
ClaudeConservative retrieval, cautious about unsupported claimsWell-sourced, balanced, factually precise content; hedged or promotional copy underperforms2–6 months
CopilotBing-grounded, enterprise and document contextBing indexation health, which is frequently neglected entirely4–10 weeks

"Practically that means I'd sequence differently per engine: technical and structural work buys Perplexity and AI Mode gains this quarter, while the ChatGPT gap gets a longer off-site programme with monthly checkpoints and no promise of a fast turn. The engine-by-engine optimisation guide breaks down the tactical differences, and Bing indexation is worth a dedicated check because Copilot visibility is often a Bing coverage problem in disguise."

Red flags / weak answer signals
  • Treats all engines as interchangeable — "optimise once, works everywhere."
  • Promises the same timeline across engines.
  • Doesn't check access before diagnosing strategy.
  • Unaware that Copilot visibility depends on Bing indexation.
  • Claims to be able to influence pre-training directly.
Q19SeniorFocus: Model bias & sentiment

The engines consistently describe us as "a budget option" when we're positioned as premium. How do you change what the model believes?

What the interviewer is evaluating

Understanding that models report aggregate sentiment rather than your marketing, and that the fix is a corpus problem with a long time constant. Also tests whether they'd loop in product and pricing rather than treating it as a content task.

Ideal high-scoring answer

"The model isn't wrong about the corpus — it's summarising what the web says. So step one is finding out where 'budget option' comes from. Usually it's a handful of high-authority sources: a review roundup that slots us into a cheap-alternatives listicle, a pricing comparison table, a Reddit thread from three years ago, or our own historical positioning that we changed in product but never changed in the places that get cited."

"Step two is honest self-assessment. If we're the cheapest in the category and our own homepage leads with price, the model is reflecting reality and this is a positioning problem for the exec team, not a GEO problem. I'd say that out loud rather than promising to fix perception around an unchanged product."

"Assuming the positioning has genuinely moved, the work is corpus correction: update our own properties to lead with the premium proof points — outcomes, enterprise capability, support SLAs — rather than price; get the outdated third-party pages corrected, which mostly works if you approach editors with facts; commission or publish evidence that supports the new framing, like benchmark data or named customer outcomes; and generate recent reviews that reflect the current product, because recency is weighted and a three-year-old consensus is a three-year-old product being described."

"Timeline: 6–12 months, and it's not linear. I'd track it as a sentiment metric — classify every answer mentioning us on a premium-to-budget axis — and report the distribution shift rather than a single score. Brand sentiment as an AI ranking input is the fuller version of this argument."

Red flags / weak answer signals
  • Proposes fixing it by rewriting the homepage and nothing else.
  • Suggests prompt-injection style tricks, hidden text, or instructions aimed at the model.
  • Doesn't consider that the characterisation might be accurate.
  • Promises a fast fix for a corpus-level perception problem.
  • No measurement approach for sentiment, only for mention count.
Q20SeniorFocus: Experimental rigour

You changed 200 pages and Share of Model went from 22% to 27%. Convince me that was you and not noise.

What the interviewer is evaluating

Intellectual honesty under pressure. Generative answers are volatile enough that a five-point swing can be pure variance. The best candidates get slightly uncomfortable at this question and then explain exactly how they'd defend or disprove the claim — the worst ones double down.

Ideal high-scoring answer

"Honestly, on those numbers alone I can't — and I'd say that before I said anything else. A five-point move on a sampled, non-deterministic metric is well inside the range that answer volatility produces on its own. Identical prompts return materially different source sets on repeat runs, so single-shot comparisons are close to meaningless."

"What would make the claim defensible: a holdout. Split comparable pages into treated and control cohorts matched on traffic, topic and existing citation rate, then compare the difference in differences. If treated moved +5 and control moved +4.5, I've proven nothing. If treated moved +5 and control moved -1, that's a real signal."

"Then sample depth. Five runs per prompt per engine minimum, ideally ten on the core set, reported as a mean with a confidence interval. And URL-level attribution: did the specific pages I changed start appearing as cited sources, or did Share of Model rise because of something unrelated — a PR hit, a competitor outage, an engine changing its retrieval mix?"

"Finally, timing hygiene. Annotate the measurement timeline with every known external event: model updates, our own releases, competitor launches. Half the 'GEO wins' I've seen presented internally were foundation model updates that lifted everyone in the category."

Red flags / weak answer signals
  • Defends the result confidently with no control group.
  • Unaware of answer volatility across repeat runs.
  • No URL-level evidence connecting changed pages to citations.
  • Attributes every positive movement to their own work by default.
  • Gets defensive rather than engaging with the methodological challenge.
Q21SeniorFocus: Operating model & tooling

You have £180k a year for GEO — headcount, tools, content, everything. Allocate it, and tell me what you're deliberately not buying.

What the interviewer is evaluating

Resource judgement and vendor scepticism. The AI visibility tool market in 2026 is crowded and expensive, and a candidate who'd spend a third of the budget on dashboards has misunderstood where the work is.

Ideal high-scoring answer

"Rough allocation: about 55–60% to people, 20–25% to content and original research, 10–12% to tooling, and the rest held back for opportunistic spend — a data partnership, a conference, a correction campaign."

"People first because GEO output is mostly editorial and technical work that someone has to actually do. I'd rather have one strong senior practitioner and a good writer than a junior plus four subscriptions."

"On tooling, I'd start with scripted monitoring rather than a platform. A prompt set, scheduled API runs across engines, results in a warehouse, visualised in whatever BI tool we already pay for. That's a week of engineering and it gives us data we own and can join to the CRM — which is the thing that makes it valuable. Most commercial AI visibility tools give you a prettier version of the same numbers and a monthly fee, with a rigid prompt structure and no join to pipeline. I'd revisit that once monitoring exceeds what scripting can maintain. The open-source GEO tooling roundup covers what can be assembled without licence costs."

"What I'm deliberately not buying: a full-suite AI visibility platform in year one, agency retainers for 'AI SEO' that duplicate in-house work, and any tool that promises to get us cited. Nothing can promise that, and the ones that claim to are usually doing something I wouldn't want associated with the brand."

"The 20–25% on content skews heavily toward original research, because it's the only content type I've consistently seen earn citations across engines without distribution spend."

Red flags / weak answer signals
  • Spends 30%+ on tooling before any measurement exists.
  • No allocation to content production at all.
  • Can't name anything they'd decline to buy.
  • Recommends a specific vendor immediately without asking what's already in the stack.
Q22SeniorFocus: Risk & ethics

What GEO tactics would you refuse to run, even if they worked and even if I asked you to?

What the interviewer is evaluating

Whether the candidate has a line and can articulate it to a superior. In a discipline this young, the person you hire will face tactics with no established norms. You want someone whose judgement you'd trust unsupervised — and someone who'd push back on you.

Ideal high-scoring answer

"A few, and I'd want to say them out loud in the interview rather than discover the disagreement later."

"Prompt injection — hidden text, invisible instructions, content designed to manipulate a model rather than inform a reader. It's deceptive, it's increasingly detectable, and the downside is having the brand publicly named in someone's research paper about manipulation."

"Serving substantively different content to AI agents than to humans. There's a legitimate grey area around format — a clean text endpoint isn't inherently deceptive — but different substance is cloaking and I'd treat it as such."

"Fabricated social proof: purchased reviews, sockpuppet accounts, fake community threads. Beyond ethics, it's a single-point-of-failure risk — one exposure and the community channel is gone permanently, which in GEO terms is a large share of your citation graph."

"And publishing claims we can't substantiate to win comparison prompts. If we claim a benchmark number, it needs a methodology page behind it. Models increasingly weight sourced claims, so the unsourced version underperforms anyway, but I'd refuse it even if it worked."

"On the grey areas I'd escalate rather than decide alone: aggressive competitor comparison content, and how hard to push for corrections on third-party pages. Those are legal and brand conversations."

Red flags / weak answer signals
  • Can't name a single tactic they'd refuse.
  • Describes prompt injection or hidden text as a legitimate technique.
  • Frames the answer entirely around "getting caught" rather than whether it's right.
  • Says they'd do whatever leadership asked — which is the wrong answer to a question about refusal.

6. Section 4 — Revenue, Attribution & Commercial ROI (Director / VP / C-Level)

What this section is for

This is the section that decides senior hires. Plenty of people can run a citation audit; very few can sit across from a CFO who has just been told the channel drives 1% of sessions and come out with a budget. These eight questions test whether the candidate can own a commercial number rather than report an activity metric.

Weight these answers at roughly double for any role at Director level or above. And listen for the phrase "influenced pipeline" — its absence is usually diagnostic.

Definition: GEO attribution

GEO attribution is the practice of connecting AI-generated brand exposure to commercial outcomes using converging evidence rather than a single tracked click. A working model combines four sources: self-reported attribution at form fill and at sales qualification, AI referral sessions isolated in analytics, server-log evidence of AI agent retrieval, and branded search lift. No one source is sufficient, because most AI exposure produces no referrer.

Q23Director / VPFocus: Attribution architecture

Design our GEO attribution model end to end. Assume GA4, Salesforce, and a six-month enterprise sales cycle.

What the interviewer is evaluating

Systems design across marketing and revenue operations, and whether they accept the measurement limits honestly. The tell for a strong candidate is that they design for converging evidence rather than trying to force a single deterministic path.

Ideal high-scoring answer

"Four layers, because no single one survives contact with reality.

  • Layer 1 — Referral capture. A custom channel group in GA4 for AI sources: chatgpt.com, openai.com, perplexity.ai, gemini.google.com, claude.ai, copilot.microsoft.com and their variants, separated from both Organic and Direct. This catches the minority of AI-driven visits that preserve a referrer. It's the smallest layer and the one most people mistake for the whole model.
  • Layer 2 — Self-reported attribution. A required "How did you first hear about us?" field on demo and trial forms with an explicit AI assistant option, plus a structured field in Salesforce that sales fills during qualification. This is the highest-signal layer for enterprise deals and the only one that captures the buyer who heard about us in ChatGPT three weeks ago and later typed the brand into Google.
  • Layer 3 — Server-side evidence. Log analysis of AI agent fetches by URL, so we can see which pages are being retrieved and when. Combined with a reverse-proxy or edge rule we can enrich and store agent requests without depending on client-side tracking.
  • Layer 4 — Aggregate lift. Branded query impressions in Search Console and direct-traffic trend, tracked against citation share. If Share of Model rises and branded search rises with it on the same curve, that's corroboration even without deterministic attribution.

"Then the plumbing: GA4 client ID or a first-party identifier pushed into Salesforce on form fill, self-reported source stored as a field on both Lead and Opportunity, and reporting on influenced pipeline — any opportunity with an AI touch in any layer — rather than last-click. With a six-month cycle, first-touch and last-touch are both misleading, so multi-touch influence is the only honest framing."

"And I'd write down the known gaps at the start: referrer stripping, users who never click, and self-report recall bias. Publishing the limitations up front is what stops the model being quietly discredited in month eight when someone notices the numbers don't reconcile. The GA4 configuration guide covers the channel-group setup."

Red flags / weak answer signals
  • Relies on GA4 referral data as the single source of truth.
  • No self-reported attribution layer.
  • Proposes last-click attribution for a six-month enterprise cycle.
  • Claims they can attribute AI influence deterministically.
  • No CRM integration — stops at the marketing analytics layer.
Q24Director / VPFocus: Self-reported attribution

Self-reported attribution is notoriously unreliable. Why would you use it anyway, and how would you make it less bad?

What the interviewer is evaluating

Whether they can work with imperfect data rather than dismissing it. Also survey design competence, which is more rare than it should be — most implementations I've reviewed are actively misleading because of how the question is worded.

Ideal high-scoring answer

"Because the alternative is worse. For a channel where most exposure leaves no referrer, self-report is often the only direct evidence that AI influenced a deal. It's noisy, but noisy directional data beats confident data about the wrong thing."

"How to make it less bad:

  • Ask about first discovery, not last touch. 'How did you first hear about us?' Otherwise everyone answers 'Google' because that's the last thing they typed.
  • Make the AI option explicit and specific. 'ChatGPT / Claude / Perplexity / another AI assistant' as a named option. If it's buried in 'Other', it will be undercounted dramatically.
  • Free-text with structured coding. Open field plus a classification pass beats a dropdown that forces a wrong choice.
  • Ask twice. Once at form fill and once in sales qualification. Agreement between the two is your confidence signal, and the gap is measurable rather than assumed.
  • Never make it required with a default. A required dropdown defaulting to the first option produces garbage at scale.

"Then calibrate rather than trust. Compare self-report rates against the referral and server-log layers over a quarter — if 14% of closed-won deals self-report AI discovery while AI referrals are 1.5% of sessions, that ratio itself is the interesting finding, and it's the number I'd take to a CFO. The gap is the channel's invisible contribution."

"I'd also report it as a range, not a point. 'Between 9% and 16% of Q3 pipeline had a credible AI discovery touch' is defensible. '13.4%' is not."

Red flags / weak answer signals
  • Dismisses self-report entirely as unusable.
  • Presents self-reported figures as precise without a confidence range.
  • Asks "how did you hear about us" at last touch and doesn't see the problem.
  • No cross-validation against any other data source.
Q25Director / VPFocus: Server-side tracking

Our GA4 shows 0.4% of sessions from AI sources. Sales says half their pipeline mentions ChatGPT. Reconcile that for me.

What the interviewer is evaluating

Analytical honesty when two data sources disagree — and whether they know the specific mechanisms that cause AI traffic to be undercounted. The weak answer picks a side. The strong answer explains why both numbers can be roughly right.

Ideal high-scoring answer

"Both numbers are probably close to true, and the gap is structural rather than anyone being wrong."

"Three mechanisms. First, referrer stripping — a large share of assistant-referred sessions arrive with no referrer and land in Direct. Published benchmarks put the misclassified share very high; one 2026 B2B analysis found roughly three-quarters of ChatGPT-referred sessions ending up in GA4 Direct. Second, and bigger: the majority of AI exposure never produces a click at all. Someone reads a recommendation, remembers the name, and searches the brand two days later. That session is branded organic. Third, in-app fetches and agent retrieval don't appear as sessions in any client-side analytics."

"To reconcile it I'd triangulate rather than adjudicate. Pull the sales-mentioned accounts and check whether their first recorded session was branded organic or direct — if so, that's consistent with AI discovery followed by brand search. Check server logs for AI agent retrieval on the URLs those accounts later visited. And track branded search volume against Share of Model over time; if they move together, the mechanism is confirmed at the aggregate level."

"The reconciled framing I'd take to leadership: 'AI search drives 0.4% of measured sessions and an estimated 8–15% of pipeline influence. Those aren't contradictory — they're measuring different things, and the second is the one that matters.' Then I'd fix the measurement gap where it's fixable: channel group configuration, landing-page-pattern detection, self-report fields."

Red flags / weak answer signals
  • Declares GA4 correct and sales anecdotal, or vice versa.
  • Doesn't know that AI referrals are frequently misclassified as Direct.
  • Misses the zero-click mechanism entirely — assumes influence requires a session.
  • Proposes no reconciliation method, only an opinion.
Q26Director / VPFocus: CAC & cycle velocity

How would you prove GEO affects customer acquisition cost or sales cycle length, rather than just asserting it?

What the interviewer is evaluating

Whether they can construct a cohort analysis and whether they understand selection bias. This is the question that separates people who've presented to a CFO from people who've presented to a CMO.

Ideal high-scoring answer

"Cohort comparison, with the selection bias stated openly. Split closed-won deals into those with a recorded AI discovery touch and those without, matched as closely as possible on segment, deal size and region. Then compare: days from first touch to close, number of sales touches required, discount depth, win rate from first meeting, and fully-loaded acquisition cost."

"My hypothesis, and what I've seen hold in two B2B programmes, is that AI-sourced deals arrive further along — the buyer has already done comparison research with an assistant that named us alongside two alternatives, so discovery calls start at evaluation rather than education. That shows up as fewer touches and a shorter cycle, which is a CAC reduction even when the marketing spend is unchanged."

"The honest caveat, which I'd volunteer: this is correlational. Buyers who use AI assistants for vendor research may simply be more prepared buyers generally. So I'd frame the finding as 'AI-sourced deals close 22% faster with 30% fewer sales touches' and explicitly note that we can't cleanly separate the channel effect from the buyer-type effect. Overclaiming causality here is how a measurement programme loses credibility permanently."

"Semrush's cross-industry work put AI search visitors at roughly 4.4× the conversion value of traditional organic. Ahrefs, reporting on their own site, found 0.5% of traffic driving 12.1% of signups — about 23×. Similarweb's panel came in around 2.15×. That spread isn't contradiction, it's methodology: one site with a technical audience versus a cross-industry average versus a conservative panel. Which is why I'd insist on measuring our own multiple rather than importing anyone's benchmark into a business case."

Sources referenced: Semrush, We Studied the Impact of AI Search on SEO Traffic; Ahrefs self-reported single-site data, 2025; Similarweb cross-site panel, 2026.

Red flags / weak answer signals
  • Asserts a CAC improvement with no cohort method.
  • Claims causality from correlational cohort data with no caveat.
  • Quotes an industry conversion multiple as if it applies to this business.
  • No matching between cohorts, so the comparison is confounded by segment or deal size.
Q27Director / VPFocus: Executive dashboard design

Design the one-page GEO dashboard our exec team sees monthly. What's on it, and what did you leave off?

What the interviewer is evaluating

Audience discipline. Most GEO practitioners want to show citation charts because they're fun. Executives need to see a commercial number and a decision. The exclusions matter as much as the inclusions.

Ideal high-scoring answer

"Six elements, one page, same layout every month so the trend is readable at a glance:

  • AI-influenced pipeline — value and deal count, trended over four quarters. This is the headline and the only number most of the room will remember.
  • Share of Recommendation on the core prompt set, with the competitor set beside it. Competitive context is what makes a visibility number meaningful to an exec.
  • Sentiment distribution — how engines characterise us, shown as a shift over time rather than a score.
  • Branded search trend as the corroborating signal.
  • One paragraph of narrative: what moved, why, what we did.
  • One recommendation with an estimated impact and the resource it needs.

"What I'd leave off: prompt-level results, engine-by-engine breakdowns, citation counts, crawler statistics, page-level performance, and anything requiring explanation of what a vector database is. Those live in the operational dashboard the team uses weekly."

"Data plumbing behind it: Salesforce opportunity data joined to the AI-influence flag, GA4 channel data, Search Console branded query data, and our citation monitoring output, all landing in a warehouse and surfaced in whatever BI tool the business already uses. The join key is the opportunity record, not the session — that's the design decision that makes the dashboard credible to finance."

"One structural point: I'd present GEO inside the organic reporting narrative rather than as a separate deck. A standalone AI dashboard invites the question 'is this replacing SEO?' every single month. The broader reporting framework covers the audience-matching logic."

Red flags / weak answer signals
  • Leads with citation counts or Share of Model rather than a commercial metric.
  • Includes fifteen metrics because they're all available.
  • No recommendation or narrative — data only.
  • No CRM data in the dashboard at all.
  • Can't name anything they'd exclude.
Q28C-LevelFocus: CFO budget defence

I'm the CFO. You want £400k for GEO next year. The channel drove 1.1% of sessions last year. Go.

What the interviewer is evaluating

Commercial fluency under adversarial questioning. Does the candidate reframe from traffic to unit economics without sounding evasive? Do they bring a downside case? This is the single most predictive question in the bank for Director-and-above hires.

Ideal high-scoring answer

"Sessions is the wrong denominator, and I'd rather argue about the right one than defend the wrong one."

"Here's the case. That 1.1% of sessions converted at several times the rate of the rest of organic — I'd bring our own measured multiple, not an industry benchmark. On top of that, 11% of closed-won deals last year self-reported first hearing about us through an AI assistant, and those deals closed faster with fewer sales touches. So the channel's revenue contribution is materially larger than its session share, and the gap is a measurement artefact, not an economics one."

"The unit economics I'd put on one slide: cost per influenced opportunity versus our blended number, and the paid-equivalent cost of the demand we're capturing. Then the forward case — AI-assisted research is growing as a share of B2B buying behaviour in every dataset available, and the cost of building citation position is lower now than it will be once the category consolidates. Third-party consensus compounds; we're buying a position that gets more expensive every quarter we wait."

"Three things I'd concede without being asked. One: measurement is imprecise and I'll report ranges, not false precision. Two: a meaningful share of the value is brand influence that won't show in a click-based model, so if the board requires deterministic attribution for every pound, this isn't the right investment and I'd say so. Three: here's the downside case — if Share of Recommendation doesn't move 15 points in nine months and influenced pipeline doesn't grow, the programme should be cut back to a maintenance budget, and I'll bring that recommendation myself."

"And I'd offer the structured alternative: £400k as proposed, or £180k for a narrowed programme covering three product lines instead of seven, with a defined decision gate at month six. Giving a CFO a scoped option rather than a binary tends to get the conversation to yes."

Red flags / weak answer signals
  • Argues that traffic will eventually come — defending the wrong metric.
  • Quotes industry statistics as if they were this company's numbers.
  • No downside case, no kill criteria, no scoped alternative.
  • Appeals to fear of missing out as the primary argument.
  • Gets defensive or evasive when the premise is challenged.
Q29C-LevelFocus: Portfolio allocation

We have one organic budget. How much of it moves from SEO to GEO next year, and how did you arrive at that split?

What the interviewer is evaluating

Portfolio thinking and resistance to hype. A candidate who wants to move half the budget to GEO in a business where organic search still drives the majority of pipeline is telling you they'd chase novelty with someone else's money.

Ideal high-scoring answer

"I'd resist framing it as a transfer, because a large share of the work is shared infrastructure. Crawlability, page speed, structured data, topical depth, entity consistency — those serve both, and the finance framing of 'moving budget from SEO to GEO' obscures that most of the spend is dual-purpose."

"The part that is genuinely distinct — citation monitoring, off-site consensus work, prompt-set research, engine-specific structuring — I'd size relative to where the buying behaviour actually is for our category. If our own data shows 11% of deals with AI discovery touches and that's grown from 4%, a GEO-specific allocation somewhere around 15–20% of the organic budget is proportionate with a bias toward the trend. Not 50%, and not 3%."

"I'd set it as a rule rather than a number: GEO-specific allocation tracks slightly ahead of measured AI influence on pipeline, reviewed quarterly. That way the split is evidence-led and I'm not relitigating it every planning cycle."

"The thing I'd protect: don't gut technical SEO to fund GEO. The retrieval layer depends on the same crawl and rendering health, so cutting there degrades both. If I had to find the money somewhere, I'd take it from low-yield link acquisition and thin content production before I touched infrastructure."

Red flags / weak answer signals
  • Proposes a dramatic reallocation with no reference to where the business's buyers actually are.
  • Doesn't recognise the shared infrastructure between SEO and GEO.
  • Would fund GEO by cutting technical SEO.
  • No mechanism for revisiting the split as evidence changes.
Q30C-LevelFocus: Forecasting

Build me a 12-month GEO forecast. What assumptions are you stating, and which one is most likely to be wrong?

What the interviewer is evaluating

Forecasting honesty. The second half of the question is the whole point — a candidate who can identify the weakest assumption in their own model is someone whose forecasts you can actually plan around.

Ideal high-scoring answer

"I'd build it bottom-up from four inputs rather than top-down from a market growth number. Input one: target Share of Recommendation by quarter on the core prompt set, based on measured movement from a pilot rather than aspiration. Input two: estimated prompt volume for those clusters — clearly labelled as an estimate. Input three: our measured relationship between recommendation share and influenced-pipeline volume, from historical cohorts. Input four: average deal value and win rate by segment, which finance already owns."

"Three scenarios, because a single-line forecast for a channel this volatile is false precision. Conservative assumes flat AI adoption and modest share gains. Base assumes continued adoption growth and the pilot's rate of share gain. Aggressive assumes one engine becomes a default surface in our market."

"The assumption most likely to be wrong is the stability of the relationship between citation share and pipeline. It's derived from a small sample over a short window, and it could break if an engine changes how prominently it surfaces recommendations, or introduces sponsored placement — which is now an active possibility rather than a hypothetical. If that relationship shifts by a factor of two, the whole forecast moves with it."

"Second most fragile: prompt volume estimates. Nobody has real data there and anyone claiming otherwise is modelling. Third: that our competitors stay passive, which they won't."

"So I'd build in a quarterly re-forecast, with the assumption set explicitly restated each time. I'd rather be re-forecast four times and honest than precise once and wrong."

Red flags / weak answer signals
  • Single-line forecast with no scenarios or ranges.
  • Top-down forecast built from market growth statistics rather than the company's own data.
  • Can't identify the weakest assumption in their own model.
  • No re-forecast cadence.
  • Assumes competitors remain static.

7. Section 5 — Practical Case Studies & Scenario Questions

What this section is for

Scenario questions are where rehearsed answers stop working. Give the candidate the situation, let them ask you clarifying questions — and pay attention to which questions they ask, because that's usually more revealing than the answer they eventually give.

Score the diagnostic sequence, not the conclusion. Someone who checks infrastructure before rewriting content has the right instincts even if they land somewhere you disagree with.

Q31ScenarioFocus: B2B SaaS deal cycle

You join a B2B SaaS company with a nine-month sales cycle. Deals consistently reach the shortlist stage and then lose to the same two competitors. Sales says prospects arrive saying "ChatGPT mentioned you as an alternative." What do you do in your first 30 days?

What the interviewer is evaluating

Whether they can read the diagnostic signal hidden in the scenario. "Mentioned as an alternative" is the key phrase — the brand has mention share but not recommendation share, and losing at shortlist stage points at comparison prompts rather than discovery prompts. A strong candidate hears that immediately.

Ideal high-scoring answer

"The phrasing tells me a lot. We're in the answer as a footnote — 'alternatives include' — not as a recommendation. And we're losing at shortlist, not at discovery, so this isn't an awareness problem. It's a positioning problem inside the comparison layer."

"Days 1–10: baseline and listen. Build the prompt set from sales call transcripts and lost-deal notes, weighted heavily toward comparison and objection prompts, because that's where we're losing. Run it across four engines, five runs each, and capture not just whether we're mentioned but how — position in the answer, framing, and which of our attributes the model cites. Then read 15 lost-deal notes properly and ask three AEs what objections they hear that they can't answer."

"Days 10–20: find the source of the framing. If engines describe us as 'a lighter-weight alternative to X,' find where that language comes from — usually a specific review site category, an old positioning page of ours, or a comparison article. Map the citation graph for the comparison prompts specifically; I'd expect a small number of sources doing most of the work."

"Days 20–30: the plan and one quick win. The plan is likely three-part — correct the third-party framing at its highest-weighted sources, build genuinely balanced comparison content that concedes what the competitors do better, and get recent reviews reflecting the current product. The quick win is usually fixing our own comparison pages, which in most SaaS companies are outdated, one-sided, and structured as marketing copy rather than as extractable comparison data."

"What I'd bring to the 30-day review: the baseline numbers, the specific sources driving the framing, a quarterly target for Share of Recommendation on comparison prompts, and the self-reported attribution field added to forms so we stop relying on anecdote by month three."

Red flags / weak answer signals
  • Misses that mention share and recommendation share are different problems.
  • Starts publishing content before establishing a baseline.
  • Never talks to sales or reads lost-deal notes.
  • Focuses on discovery-stage prompts when the loss is at shortlist.
  • Promises pipeline impact within 30 days on a nine-month cycle.
Q32ScenarioFocus: Incident diagnosis

A major foundation model update ships. Within two weeks our brand citations drop by roughly 60% across the prompt set. Walk me through your diagnosis and your fix, hour by hour if you need to.

What the interviewer is evaluating

Incident discipline. The temptation is to attribute the drop to the model update and start rewriting content. The correct instinct is to rule out cheaper explanations first — and to establish whether the drop is even real before doing anything. This is the question I'd ask every senior candidate without exception.

Ideal high-scoring answer

"I'd work outward from the cheapest, most likely explanation, and I'd resist touching content until step four.

  • Step 1 — Is it real? Re-run the full prompt set at higher sample depth, ten runs per prompt rather than five. Answer volatility alone produces large swings, and I've seen 'drops' that were entirely sampling artefacts. Check whether the prompt set or the measurement script changed. Check whether the drop is uniform or concentrated in specific engines and clusters — a drop on one engine is a different problem from a drop on four.
  • Step 2 — Can they still reach us? Server logs for GPTBot, OAI-SearchBot, PerplexityBot, ClaudeBot: request volume, status codes, response times. Then robots.txt history, WAF and bot-management rule changes, CDN configuration, certificate issues, and rate limiting. In my experience this is where the cause sits more often than anyone expects — a security team tightens bot rules in the same fortnight as a model update and the correlation looks causal.
  • Step 3 — Did our content change? Diff the previously-cited URLs. Were they edited, consolidated, redirected, migrated? Did a CMS release change the rendering path so the key passages now load client-side? Check whether the specific passages that were being quoted still exist verbatim.
  • Step 4 — Did the competitive picture change? Who's being cited instead of us? If a competitor published something significant, or a new review roundup started getting cited, that's displacement, not a model problem. Also check whether the citation mix shifted structurally — for example, toward community sources — which changes the work required.
  • Step 5 — Only now, the model. If access is clean, content is unchanged, and competitors didn't move, then it's the update: different retrieval weighting, different source preferences, or different synthesis behaviour. Confirm by checking whether category peers moved too — if everyone in the category dropped, it's the engine.

"The fix depends on which step it stopped at, and the honest answer is that steps 2 and 3 are fixable in days while step 5 is a rebuild, not a patch. If it's genuinely the model, I'd resist the urge to rewrite everything on a hunch. I'd take the answers that replaced ours, analyse what the newly-cited sources have in common — length, structure, source type, recency — and run a controlled test on a subset of pages before committing the whole library."

"Communication matters as much as the diagnosis. Day one I'd tell stakeholders what I know, what I don't, and when they'll get the next update — because the alternative is leadership hearing about a 60% drop through a vendor dashboard and drawing their own conclusions."

Red flags / weak answer signals
  • Immediately attributes the drop to the model update without ruling anything out.
  • Starts rewriting content as the first action.
  • Never checks server logs, robots.txt history, or WAF changes.
  • Doesn't verify the drop is real before acting on it.
  • No stakeholder communication plan.
  • Claims the situation is unfixable and recommends abandoning the channel.
Q33ScenarioFocus: Regulated verticals (Healthcare / Pharma)

We're a healthcare company. Medical affairs has to approve every public claim, legal reviews everything, and engines are answering patient questions about our therapeutic area right now. How do you run GEO here?

What the interviewer is evaluating

Whether they can operate inside real constraints instead of treating compliance as an obstacle. Also tests whether they understand that YMYL content is held to a different standard by both search and generative systems — and whether they'd flag patient-safety risk proactively.

Ideal high-scoring answer

"The constraints actually work in our favour here, which is worth saying up front. Generative systems weight authority, provenance and sourcing heavily on health topics, and a company with medical review, named clinical authors and citable primary sources has structural advantages over content farms. The problem isn't the constraints, it's the review cycle time."

"So I'd redesign the workflow rather than fight it. Pre-approved modular content blocks — mechanism of action, indication summaries, safety information — reviewed once, reusable across pages, versioned. That turns a six-week approval per asset into a two-day assembly job. I'd build that with medical affairs rather than presenting it to them."

"On structure: heavy emphasis on provenance. Named authors with credentials, medical reviewer with date, MedicalWebPage and Article schema, explicit citations to primary literature, clear last-reviewed dates. That's both compliance and GEO — the things that satisfy a medical reviewer are the things that make a passage citable."

"The critical piece most people miss: monitoring for harm, not just for visibility. I'd run a prompt set specifically checking whether engines are stating anything inaccurate or unsafe about our products — wrong dosing, missing contraindications, off-label framing. That's a patient safety and regulatory issue before it's a marketing one, and I'd want a defined escalation path to medical affairs and regulatory within 24 hours of finding something, plus a documented process for reporting to the engine provider."

"What I wouldn't do: chase consumer-symptom prompts that verge on medical advice, or optimise for prompts where the right answer is 'talk to your doctor.' Beyond regulatory exposure, engines handle those conservatively and we'd be spending effort to lose."

Red flags / weak answer signals
  • Treats medical and legal review as a blocker to be circumvented.
  • No monitoring for inaccurate or unsafe statements about the products.
  • Proposes tactics that would put regulated claims into uncontrolled channels.
  • Doesn't recognise the authority advantage regulated organisations have in YMYL topics.
  • Suggests AI-generated medical content without a clinical review step.
Q34ScenarioFocus: D2C high-ticket conversion

We sell £3,500 mattresses direct to consumer. Buyers research for weeks. Engines are answering "best mattress for side sleepers with back pain" and we're nowhere. Where do you start, and what's different about D2C versus B2B here?

What the interviewer is evaluating

Category adaptation. D2C high-consideration has different dynamics from B2B: review platforms and community sentiment dominate, the affiliate-review ecosystem is heavily commercialised, and product data feeds matter in ways they don't for software. A candidate who runs the same playbook regardless of category is a band-2 hire.

Ideal high-scoring answer

"The big structural difference: in B2B the citation graph is analyst content, review platforms and communities. In D2C mattresses it's dominated by affiliate review sites with commercial relationships, plus Reddit, which is unusually influential precisely because consumers distrust the affiliate layer. Engines read both, and increasingly weight the community source."

"Start with the citation map for the specific prompt cluster — side sleepers, back pain, firmness, trial periods. I'd expect a handful of review sites and two or three subreddits to carry most of it. That's the target list, and it's mostly not our website."

"Three workstreams. First, product data integrity: firmness ratings, materials, dimensions, trial and warranty terms marked up with Product and Offer schema and consistent across every channel including retail partners and feeds. Engines answering 'which mattresses have a 100-night trial' are reading structured data, and inconsistency between our site and a retailer's listing produces hedged answers."

"Second, the review ecosystem: get accurately represented on the sites that actually get cited, through whatever legitimate route exists — product submission, correcting outdated specs, providing test units. Plus first-party review volume and recency, because review text gets quoted directly."

"Third, genuinely useful owned content that answers the clinical-adjacent questions: what firmness suits side sleepers with lower back pain, and why. Specific, sourced, not promotional. This is where most D2C brands fail — they write product pages and expect them to win informational prompts."

"On measurement, D2C has an advantage: shorter cycles mean self-reported attribution at checkout gives usable signal within weeks rather than quarters. I'd add it to the post-purchase flow immediately. The customer journey mapping matters more here than in B2B because the research phase spans many touchpoints and one assistant conversation may sit at the very start of it."

Red flags / weak answer signals
  • Applies a B2B playbook unchanged.
  • Ignores product feeds, structured product data, and retailer listing consistency.
  • Proposes paying affiliate sites for favourable placement without disclosure.
  • Focuses only on the owned site when the citation graph is mostly third-party.
  • Misses that shorter D2C cycles make attribution measurement easier, not harder.
Q35ScenarioFocus: Misinformation crisis

Three engines are telling prospects that our product doesn't support SSO. It has for two years. Sales lost two deals to it last month. Fix it.

What the interviewer is evaluating

Crisis response and an understanding that you cannot edit a model. The strong answer separates immediate damage control from the slower corpus correction, and includes a sales enablement step — because the deals are being lost this quarter, not whenever the corpus updates.

Ideal high-scoring answer

"Three timelines running at once: today, this month, this quarter.

Today — stop the bleeding. Arm sales with a direct response and a link to a definitive SSO documentation page. Deals are being lost now and no corpus fix arrives in time for them. I'd also check whether any engine offers a feedback or correction mechanism and use it, though I wouldn't count on it."

This week — find the source. The model got this somewhere. Usually it's one of four things: our own documentation says SSO is 'coming soon' on a page nobody updated, a comparison table on a review site is out of date, an old Reddit thread or support article says it's unsupported, or a competitor's comparison page claims it. I'd search for the claim in the engines' own cited sources — they'll often tell you where they got it, which is the fastest route to the origin."

This month — correct the corpus. Fix every owned instance first, including release notes, changelogs, old blog posts and help centre articles. Then request corrections on third-party sources: review platforms usually update feature matrices when a vendor submits evidence, and comparison-page authors generally correct factual errors when approached properly. Publish an unambiguous, well-structured SSO page — supported protocols, providers, setup steps, plan availability — because the correction needs something citable to replace the wrong claim with."

This quarter — verify and monitor. Add the claim to the monitoring prompt set and track how long it takes each engine to update. In my experience retrieval-heavy engines like Perplexity correct within weeks of the source changing, while pre-training-weighted answers can persist for months. Then add a standing check: a small prompt set that verifies engines state our core capabilities correctly, run monthly, owned by whoever owns product marketing. Feature-claim drift is a permanent condition now, not a one-off incident."

"The organisational fix I'd push for: product launches get a corpus checklist. When a feature ships, someone updates the docs, the comparison pages, the review platform listings and the help centre — not just the release notes."

Red flags / weak answer signals
  • Believes the model can be contacted and corrected directly, like a data feed.
  • No immediate sales enablement — treats it purely as a content problem.
  • Doesn't investigate where the claim originated.
  • Fixes only owned properties and assumes third-party sources will follow.
  • No monitoring to verify the correction propagated.
  • Promises all engines will be corrected within days.

8. The Interviewer Scorecard

Scoring after the fact, from memory, is how panels end up arguing about vibes. Fill this in during the interview, one row per competency, and have every panellist submit before the debrief starts.

CompetencyQuestions that test itWeight (IC)Weight (Director+)Minimum bar
Retrieval fundamentalsQ2, Q3, Q525%10%Band 3 for IC, band 2 for Director
Technical executionQ8–Q11, Q1530%10%Band 3 for IC
Off-site & consensus strategyQ12, Q17, Q1915%20%Band 3 across the board
Measurement rigourQ6, Q20, Q24, Q2515%25%Band 3, and band 4 for any measurement-owning role
Commercial ownershipQ26–Q305%30%Band 4 for Director+, band 2 acceptable for IC
Judgement under ambiguityQ22, Q32, Q3510%5%Band 3 minimum for every level, no exceptions

🚩 Automatic no-hire signals, regardless of score elsewhere

  • Proposes prompt injection, hidden text, or any tactic designed to deceive a model rather than inform a reader.
  • Proposes fake reviews, sockpuppet accounts, or undisclosed paid placement.
  • Presents GEO measurement as solved and precise.
  • Cannot name a single thing they got wrong in a previous role.
  • Claims guaranteed citation outcomes for a defined timeline.
  • Answers every question with the same list of tactics regardless of what was asked — usually a memorised playbook rather than practice.
🗣️ From my experience

The most useful interview I ever ran used four questions instead of twelve. We asked Q20 (prove it wasn't noise), Q22 (what would you refuse), Q28 (the CFO), and Q32 (the citation drop), and gave the candidate forty minutes with permission to ask us anything. She spent the first six minutes asking about our sales cycle, our CRM hygiene, and whether marketing and sales agreed on what a qualified lead was — before answering a single question.

That was the hire. Not because her answers were better than anyone else's, but because she understood that GEO in an enterprise is 30% retrieval mechanics and 70% getting other teams to change what they do. We'd been screening for the 30%.

9. If You're the Candidate: How to Prepare

Most of this page is written for the interviewer, but the same material works in reverse. A few things that consistently separate strong candidates in these interviews, based on the panels I've sat on:

1
Bring three numbers you own

Not industry statistics — your numbers. A before-and-after on citation share, a pipeline figure, a page-cohort result. If you've never measured your own work, say so directly and explain how you'd set it up now. That reads far better than borrowed benchmarks.

2
Prepare a failure, properly

A real one, with the diagnosis and what you changed afterwards. "We rewrote 400 pages before establishing a baseline and couldn't prove any of it worked" is a strong answer. Interviewers at band 4 are specifically listening for this and almost nobody volunteers it.

3
Know the measurement limits cold

Non-determinism, referrer stripping, the zero-click majority, the wide spread between published conversion benchmarks. Being able to explain why Ahrefs reported 23× and Similarweb reported 2.15× — and why neither is your number — signals more expertise than any tactic you could list.

4
Have a position on ethics before you're asked

Q22 catches people flat-footed. Decide in advance what you'd refuse and why, and be ready to say it to someone senior who might disagree.

5
Ask about their data before you answer

CRM hygiene, whether sales and marketing agree on definitions, whether anyone's measuring citations today. It's not a tactic — the answers genuinely change what you'd do, and asking demonstrates it.

📚 Worth working through before an interview: the AEO / SEO / GEO implementation checklist for the execution layer, the technical SEO MCQ to pressure-test fundamentals that still come up in GEO interviews, and the AI SEO tools guide so you can discuss the tooling landscape without sounding vendor-captured.

The question bank tests knowledge. These guides build it — each one covers a competency block from the sections above in full depth.

🔗 Build the Knowledge Behind the Answers
🤖 GEO · AEO · Citations How to Rank in AI Overviews & LLM Answers

The core GEO pillar — how retrieval and citation actually work, and what to change on a page to earn a mention. The execution layer behind Sections 1 and 2 of this question bank.

Read guide →
🕷️ Crawlers · robots.txt · Access Robots.txt for AI Crawlers: The Complete Directive Reference

Per-agent directives for GPTBot, ClaudeBot, PerplexityBot, Google-Extended and the rest — including what blocking each one actually costs. The reference behind Q5 and Q10.

Read guide →
📊 Metrics · Citation Rate · Reporting AI Citation Rate: Defining and Measuring the Metric

How to define citation rate, Share of Model and Share of Recommendation in a way that survives scrutiny — sampling, denominators, and confidence bands. Directly supports Q6, Q20 and Q27.

Read guide →
🧩 Schema · Entities · Structured Data Schema Markup & Structured Data Guide 2026

The @graph patterns, entity anchoring and sameAs linking that make brand facts unambiguous to generative engines — the implementation detail behind Q4 and Q9.

Read guide →
🔍 Multi-Engine · Perplexity · ChatGPT How to Optimise for Perplexity, ChatGPT & Gemini

Engine-by-engine retrieval behaviour and what moves each one — the material a candidate needs to answer Q18 without generalising across platforms that behave very differently.

Read guide →
📈 Reporting · KPIs · Stakeholders SEO Reporting Guide 2026: KPIs, Dashboards & Stakeholder Reports

Audience-matched reporting structure, vanity versus value metrics, and the executive one-pager format — the framework the dashboard answer in Q27 is built on.

Read guide →

11. Frequently Asked Questions About GEO Interviews

What is GEO and how is it different from SEO and AEO?

GEO (Generative Engine Optimization) is the practice of making a brand retrievable, citable and recommendable inside generated answers from systems like ChatGPT, Google AI Mode, Perplexity, Gemini and Claude. SEO optimises for ranked links. AEO optimises for a single extracted answer such as a featured snippet. GEO optimises for synthesis — the model reads many sources, forms a position, and decides which brands to name.

The practical difference is the unit of optimisation: SEO optimises pages, AEO optimises answers, GEO optimises passages plus the third-party consensus surrounding a brand. Semrush found nearly 90% of ChatGPT-cited pages ranked position 21 or worse in Google, which is the clearest evidence that citation isn't simply a reward for ranking well.

What questions are asked in an enterprise GEO interview?

Enterprise GEO interviews typically cover five areas: retrieval foundations (how RAG works, crawler behaviour, citation metrics), technical execution (chunking, schema, robots.txt directives, JavaScript rendering), advanced strategy (prompt reverse engineering, competitive displacement, multi-engine differences), commercial ownership (attribution design, CAC impact, budget defence), and scenario diagnosis (citation drops, regulated verticals, misinformation incidents).

Seniority determines the weighting. An associate interview is mostly Sections 1 and 2; a Director interview is mostly Sections 4 and 5. The competency that carries a minimum bar at every level is judgement under ambiguity.

What is Share of Model (SoM) in generative engine optimization?

Share of Model is the percentage of runs, across a fixed and repeatable prompt set, in which a brand is mentioned by a generative engine at all. Share of Recommendation is the narrower subset where the brand is actively recommended rather than merely named.

Both are sampled estimates, not counts. Because answers are non-deterministic, a credible measurement programme freezes the prompt set, runs each prompt at least five times per engine, logs position and sentiment alongside the mention, and reports a rolling average with a confidence band. A single-run number is close to meaningless.

How do you measure ROI from GEO when AI search sends so few clicks?

You measure influenced pipeline rather than sessions. The working model combines four inputs: self-reported attribution captured at form fill and again at sales qualification, AI referral sessions isolated in GA4, server-log evidence of AI agent retrieval on key pages, and branded search lift in Search Console.

Semrush's analysis put AI search visitors at roughly 4.4× the conversion value of traditional organic visitors, while Ahrefs reported 23× on their own site and Similarweb's panel found 2.15×. That spread is methodological, not contradictory — which is precisely why you measure your own multiple rather than importing a benchmark into a business case.

How do you track AI referral traffic in GA4?

Create a custom channel group capturing referrals from chatgpt.com, openai.com, perplexity.ai, gemini.google.com, claude.ai and copilot.microsoft.com, separated from both Organic and Direct. Then accept that it will undercount substantially.

A large share of AI-referred sessions arrive without a referrer and land in Direct — one 2026 B2B analysis put roughly three-quarters of ChatGPT-referred sessions in that bucket. More importantly, most AI exposure produces no click at all: the buyer reads a recommendation and searches the brand two days later. Pair the channel group with self-reported attribution, server-log analysis and branded search tracking, and treat GA4 as one signal of four.

How do you diagnose a sudden drop in brand citations after a model update?

Work outward from the cheapest explanation. First confirm the drop is real by re-running the prompt set at higher sample depth, since answer volatility alone produces large swings. Second, check infrastructure: robots.txt history, WAF and bot-mitigation rules, CDN configuration, and server-log evidence that AI agents are still fetching key URLs successfully. Third, check whether the previously-cited content was edited, consolidated or redirected. Fourth, check whether a competitor or a new third-party source displaced you.

Only after all four should you attribute the drop to the model update itself — and if it genuinely is the model, analyse what the newly-cited sources have in common and test on a page cohort before rewriting the library.

What should a GEO candidate know about robots.txt and AI crawlers?

They should distinguish three categories, because blocking the wrong one has very different costs. Training crawlers — GPTBot, ClaudeBot, CCBot, Applebot-Extended — affect whether content shapes future model weights. Retrieval agents — OAI-SearchBot, PerplexityBot, Google-Extended — affect whether you appear in live answers today. User-action fetchers — ChatGPT-User, Claude-User — fire when a person explicitly asks the assistant to open your page.

They should also know robots.txt is advisory, that enforcement lives at the WAF or CDN layer, and that most "we've been blocked" incidents trace to bot-mitigation rules rather than the robots file. Cloudflare's May 2026 data put AI crawlers at 20.3% of verified bot traffic, so crawl policy is now an infrastructure cost conversation as well as a visibility one.

Three things to do before your next GEO interview — on either side of the table: (1) Pick four questions from this bank rather than twelve, and give the candidate permission to ask you things — the questions they ask will tell you more than the answers they give. (2) Write down your automatic no-hire signals before the first interview, not after the third, because panels drift toward whoever interviewed best rather than whoever would do the job. (3) If you're the candidate, measure something on your own site this week — one prompt set, one engine, five runs — so you walk in with a number that's yours.
Compiled by Rohit Kunal Technical SEO Specialist & AI Search Researcher · IndexCraft

Rohit Kunal is a Technical SEO Specialist and AI Search Researcher at IndexCraft. He's spent 13+ years on hands-on organic growth work — technical audits, GA4 and attribution implementation, Core Web Vitals, and since 2024, AI search citation strategy. He has sat on hiring panels for GEO, technical SEO and organic growth roles at agencies and in-house teams, and built the scoring rubric on this page after watching too many panels disagree about what a good answer sounded like.

Since Google AI Overviews launched globally in May 2024 he's tracked citation patterns across ChatGPT, Google AI Mode, Perplexity and Claude, running fixed prompt sets against live answers and documenting how citation share moves — and doesn't move — in response to on-site changes. He writes IndexCraft's ongoing GEO and AEO research and has worked with 150+ websites on measurable organic growth.