Generative engine optimisation for enterprise retail

Published:   
August 13, 2026
Updated:  
August 13, 2026
Generative engine optimisation for enterprise retail
Article highlights
  • Being named inside the answer now matters more than ranking beneath it. Pew Research Center's browsing-panel study of 900 US adults found users clicked a standard search result in 8% of visits where an AI summary appeared, against 15% where none did — and clicked a link inside the summary in roughly 1% of visits.
  • The academic evidence on GEO content tactics is contested, not settled. The original KDD 2024 paper measured visibility gains of up to 40% from content changes; a later peer-reviewed benchmark covering product recommendation found the same class of tactics largely ineffective, sometimes harmful, and subject to diminishing returns as more competitors adopt them.
  • Position in the model's context beats prose polish. That same benchmark found the rank a document holds when it is handed to the model outweighs sophisticated rewriting — which means conventional technical SEO remains the load-bearing input, not a legacy concern.
  • Your store network is the GEO asset nobody writes about. The ABS measured domestic online sales at 12.7% of Australian retail turnover in its final Retail Trade release, so the clear majority of transactions still complete in a store — while store locator pages score below homepages on machine readability.
  • Australian shoppers are adding AI to their research, not substituting it. IAB Australia's survey of 1,079 online shoppers found around 60% use AI tools when shopping, but 74% use it as one source among several, and eight in ten hold some concern about relying on it.

Introduction

Something changed in the click data, and it changed in a direction that is hard to argue with. Pew Research Center tracked the browsing of 900 US adults across 68,879 unique Google searches in March 2025 and found that when an AI summary appeared, users clicked a standard result in 8% of visits — against 15% when no summary appeared. Clicks on the citations inside the summary ran at about 1%. Meanwhile Adobe, drawing on more than a trillion visits to US retail sites, reports that the AI-referred visitors who do arrive now convert better than any other channel it measures. Fewer clicks, better clicks. Both things are true at once, and together they change what a retail website is for.

Almost every guide on generative engine optimisation is written for software marketers, whose entire business is a website. Yours is not. This guide covers how generative engines select a retailer, why your product feed, store pages and category pages are the three assets that decide the outcome, what to measure, how to sequence the work over two quarters, and — the part most guides skip — where the evidence is thin enough that you should hedge your budget.

How we got here — from ranked lists to a single recommendation

The term was coined in November 2023, when a group of researchers from Princeton, Georgia Tech, IIT Delhi and the Allen Institute for AI published GEO: Generative Engine Optimization, later accepted to KDD 2024. They built a benchmark of queries across nine domains, tested nine content strategies, and reported visibility improvements of up to 40% in generative engine responses. Two tactics did most of the work: adding relevant statistics, and adding quotations from credible sources. Keyword stuffing, the oldest trick in search, performed poorly and sometimes went negative.

That paper set the terms of the conversation, and the vendor category that followed has been repeating its headline number ever since. Its central claim has since been directly challenged — we return to that below, because it changes how you should spend.

What is not disputed is the shift in surface area. Google's own guidance, published in May 2026, states plainly that its generative features are built on its core Search ranking and quality systems, using retrieval-augmented generation — a technique, also called grounding, where the model retrieves pages from an index before it writes an answer — and that there are no special technical requirements to appear. At the same time, an entirely separate channel opened that has nothing to do with crawling at all. OpenAI now ingests merchant-supplied product feeds directly, so that ChatGPT works from a structured file you push rather than a page it guesses at.

Two doors, then. One is the index you have always optimised for. The other is a data pipeline that looks far more like Google Merchant Center than like content marketing. Most retail teams have someone accountable for the first door and nobody accountable for the second.

How a generative engine actually chooses a retailer

Strip away the vocabulary and there are three gates. Working out which one you are failing is most of the diagnostic work.

Gate one: can the machine read the page. This is more of a problem than most teams assume. Adobe scored pages across the US retail sector for machine readability and published sector averages in April 2026: homepages came in at 75%, meaning roughly a quarter of homepage content was not readable by a language model. Category pages scored 74%, store locators 73%, and individual product pages 66% — the worst-performing page type in the set. These figures come from Adobe, which sells a product to fix the problem and has not published its scoring methodology, so treat the absolute numbers with caution. The ranking across page types is the useful part, and it points somewhere uncomfortable: the pages carrying your commercial intent are the least legible to the systems now doing the recommending.

Gate two: does the retrieval step surface you. This is where the most interesting finding of the last two years sits. The peer-reviewed C-SEO Bench, presented at NeurIPS 2025, tested conversational optimisation methods across two tasks — question answering and product recommendation — with three domains each. Most methods were largely ineffective and frequently pushed rankings the wrong way. What mattered far more was the position a document held when it was handed to the model, with the first two positions conferring significant advantage in every domain tested. In other words: what gets you into the answer is mostly what got you to the top of the retrieval set — conventional search performance.

Gate three: does the model trust you enough to name you. Retrieval gets you into the candidate pool; selection is a separate judgement, and the evidence suggests it leans on signals outside your website. A comparative analysis from University of Toronto researchers found AI search exhibited a pronounced bias toward earned media over brand-owned content, against Google's more balanced mix. This is a preprint, not a peer-reviewed result, so hold it loosely — but if it holds, part of your GEO budget belongs to PR and review platforms rather than to your CMS.

The three GEO assets generic guides never touch

Here is where retail diverges sharply from the software playbook every other guide is written from. A SaaS company has a website, a blog and a pricing page. You have a catalogue, a store network and a taxonomy — three assets with no real equivalent in the guides currently ranking for this term.

Product data — the feed is the new shelf

The most consequential change for a retailer is not stylistic. It is that a second, parallel distribution channel now exists which bypasses your website entirely. OpenAI's product feed specification defines the fields, types and validation rules a merchant must supply for products to be indexed and displayed accurately inside ChatGPT, and OpenAI is explicit that the merchant-provided feed — not passive crawling — is the structured source of truth for price and availability. Google's guidance points the same way, naming Merchant Center feed quality as an ecommerce input to its AI responses.

The consequence is a reporting-line problem more than a technical one. Feed quality has historically sat with paid media or merchandising, judged on shopping-ads performance. It is now a discovery input for organic AI surfaces — so the person accountable for AI visibility and the person accountable for the feed are usually different people with different targets. Where a specification is published, treat it as a requirements document and audit field by field rather than assuming your Google feed maps across.

There is a harder question underneath. Harvard's Aounon Kumar and Himabindu Lakkaraju demonstrated that inserting an optimised strategic text sequence into a product's information page could move a product from not being recommended at all to being the model's top recommendation — a technique they showed could also circumvent model guardrails. Their catalogue was fictitious coffee machines rather than a live retailer, so read this as a demonstrated vulnerability, not a measured market effect. The same behaviour is being formalised in E-GEO, an MIT and Columbia dataset of 13,747 multi-sentence consumer product queries paired with retrieved listings. The strategic read is not that you should do this. It is that product text is a ranking surface, that it is manipulable, and that platforms will eventually police it — so build your advantage on attribute depth, which nobody can patch away.

Store and location pages — the asset with no software equivalent

This is the section absent from every GEO guide written for SaaS, and where the largest unclaimed advantage sits for a multi-store retailer.

Start with the size of the prize, because it gets consistently understated. In its final Retail Trade release for June 2025, the ABS put domestic online sales at 12.7% of total Australian retail turnover, up from 11.6% a year earlier and from 1.8% when it first measured the series in 2013. Australia Post, drawing on CommBank iQ data, reports that 24% of all retail spend is now online against a record $82.6 billion spent online in 2025. Those two figures use different bases and we could not reconcile them from the published methodology, so treat neither as the number. On both, the same conclusion holds: most Australian retail transactions still complete in a store, so the store network is the majority of what a shopper is trying to find.

Generative engines are far more selective about local recommendations than the results they are displacing. SOCi's 2026 Local Visibility Index, covering more than 350,000 business locations across 2,751 brands, reports that ChatGPT recommended 1.2% of brand locations against a 35.9% appearance rate in Google's local three-pack. SOCi sells multi-location visibility software and its methodology is not fully published, so this is a vendor benchmark, not an independent measurement. Its most useful finding is structural: in retail, SOCi found only 45% overlap between the brands most visible in traditional local search and those most recommended by AI platforms. If that holds even approximately, a strong local search program is not evidence you are covered.

What this means practically. Your location pages have to answer the question a shopper actually asks an assistant, which is almost never "where is your Chatswood store". It is closer to "which store near me has this in a 10 in stock tonight". That needs three things a typical store page does not carry: per-location stock visibility, per-location service detail — click-and-collect windows, fitting rooms, repairs, returns handling — and opening hours that survive a public holiday. Australia has eight jurisdictions with different holiday calendars and trading-hour rules, so a nationally templated hours field is wrong somewhere in the country most months of the year.

Category pages — the answer to the comparison question

Category pages are the most undervalued GEO asset a retailer owns, for a specific and measurable reason. Pew found that longer, more conversational queries are far more likely to trigger a generative answer: 8% of one- and two-word searches produced an AI summary, against 53% of searches of ten words or more, and 60% of searches beginning with a question word. Comparison and constraint queries are long by nature. "Best cordless vacuum for a small apartment with pets under $400" is a category-page question, not a product-page question.

A product page can only argue for one product. A category page carries the comparison logic, the selection criteria, the trade-offs and the price bands — the shape of content a generative engine needs to synthesise a recommendation. The original GEO paper's finding that statistics and citations lifted visibility most is, read carefully, a finding about explanatory content rather than transactional content. Category pages are where explanatory content lives on a retail site — and on most retail sites they are a filtered grid with a hundred words above it.

Two constraints. Google's guidance is explicit that creating pages at scale primarily to manipulate rankings is not rewarded, so a generated page per long-tail permutation is the wrong reading. And the honest version of this work is expensive: it needs someone with category knowledge writing real buying guidance. That is a merchandising resource question, not a content-ops one.

What good looks like — and what nobody can benchmark yet

Measurement is where GEO is weakest, and you should budget for that rather than pretend otherwise.

Share of answer is the closest thing to a primary metric: across a fixed set of buyer questions run weekly, how often are you named, and how often is a competitor named instead. Fix the prompt set and the schedule before you start — model outputs vary between runs, and an unfixed prompt set produces noise you will misread as movement. We could not find a reliable published Australian benchmark for share of answer in any retail category, so this measures against your own baseline and named competitors, not an industry figure.

Machine readability is the one input you fully control and can audit yourself. Render your product, category and store pages the way a crawler does, without JavaScript execution, and check what survives. Adobe's page-type ranking is a sequencing hint, not a target.

Crawler access is a decision you may have already made without knowing. Cloudflare publishes crawl-to-refer ratios by AI platform — pages crawled per referral sent back — and notes that native-app referrals often arrive without a referer header, which can make the imbalance look worse than it is. The point is not the ratio. It is that if your CDN or robots.txt blocks the retrieval crawlers, no amount of content work will register — and this is usually a default someone inherited rather than a decision anyone made.

AI-referred conversion needs its own analytics segment before you can argue for budget. Adobe reports AI-referred retail visitors in May 2026 converted 54% higher and generated 53% more revenue per visit than non-AI traffic, having converted at roughly half the rate a year earlier. This is Adobe's own analytics data, published alongside a product launch, and it is US retail. Australian logistics costs, store density and payment mix differ enough that the multiple should not transfer directly. Measure your own.

Implementation — two quarters, in order

1. Confirm you are not blocked. Audit robots.txt, CDN bot rules and WAF configuration against the current retrieval and user-triggered crawlers, then reconcile what robots.txt permits against what actually reaches origin. These disagree more often than teams expect, and everything downstream is wasted if this gate is shut.

2. Fix product page readability before you write anything new. Product pages are the weakest page type on Adobe's benchmark and carry your commercial intent. Server-render the specification tables, price, availability and review summary. This is engineering work, not content work, and it returns the most for the least.

3. Treat the feed as a discovery asset and give it an owner. Audit against the published specifications field by field. Prioritise attribute completeness on top revenue categories over uniform coverage across the catalogue, and make sure whoever owns the feed knows organic AI visibility now depends on it.

4. Rebuild location data as structured data, not a page template. Per-store hours with jurisdiction-correct holiday handling, per-store services, and per-store stock where your systems support it. This has the longest lead time, because it usually depends on store-level inventory accuracy you do not yet have.

5. Write real buying guidance on your top ten categories. Not an SEO introduction. Selection criteria, trade-offs, price bands and the comparison a shopper is actually making — specific enough that a model can lift a defensible claim from it.

6. Stand up measurement last, and hold it steady. Fixed prompt set, fixed cadence, named competitors, AI referral segmented in analytics. Six months of consistent measurement beats a sophisticated dashboard rebuilt every quarter.

Prerequisites, stated plainly. Steps four and five do not work without store-level inventory accuracy and category expertise respectively. If you have neither, steps one to three still return value on their own, and are where you should stop.

Common failure modes

Buying a monitoring tool and calling it a program. Visibility dashboards tell you that you are absent, not why. The why is usually gate one or gate two, both engineering problems.

Treating GEO as a content brief. The C-SEO Bench result cuts against this directly. If retrieval position outweighs content sophistication, a content-only response addresses the wrong gate.

Generating pages at scale. Google has named this specifically as ineffective and potentially in breach of its spam policies. The permutation logic that once worked for long-tail search does not transfer.

Optimising the website and ignoring the store network. The most common failure for a multi-store retailer, and the one with the largest opportunity cost, given where the majority of transactions still complete.

Letting a US benchmark set an Australian target. Australian shoppers are demonstrably more cautious. IAB Australia and Pureprofile's nationally representative survey of 1,079 Australian online shoppers, fielded in May 2026, found around 60% use AI when shopping, rising to 75% of those aged 18 to 39 — but that 74% use it as one source among several, and eight in ten hold some concern about using AI to research products. Search engines and retailer sites remain central, with 92% using one or the other to discover and compare. AI is entering the Australian research phase without displacing anything yet.

Where this evidence is weak

Four things you should know are unresolved before you commit budget.

The founding claim is contested. The KDD 2024 paper reported up to 40% visibility gains from content optimisation. C-SEO Bench, peer-reviewed at NeurIPS a year later, found most such methods largely ineffective and sometimes harmful, and explicitly framed its results as challenging the earlier paper's assumption that traditional SEO would be superseded. Both are credible. The disagreement is genuine and unresolved.

The gains may be self-cancelling. C-SEO Bench also found average gain per adopter fell as the number of adopters rose, converging toward zero at full adoption, and described the problem as congested and zero-sum. If that generalises, early movers capture a temporary advantage and late movers pay to stand still. No published work we could find has measured how fast that window closes in retail.

Almost all the commercial evidence is vendor-published. The traffic multiples, conversion lifts and readability scores driving this category's urgency come from companies selling the remedy, with methodologies that are not public. Failed programs are not written up. Treat every published uplift figure as a ceiling, not an average.

There are no Australian benchmarks. The behavioural data underpinning nearly every number above is US. The Australian data that exists is consumer survey data on intent, not measured outcome data on retailer visibility. Anyone quoting you an Australian GEO benchmark is extrapolating.

Where this is heading — labelled as forecast, not fact

Two developments look likely enough to plan around, and both are predictions rather than findings.

The first is that the feed becomes the primary channel and the page becomes secondary. Both major platforms are converging on merchant-supplied structured data as the source of truth for commerce, with OpenAI's specification supporting refreshes measured in minutes rather than days. If that holds, catalogue data quality becomes a discovery capability rather than an operations chore, and retailers with clean attribute coverage across their full range hold an advantage content work cannot close.

The second is that physical availability becomes the differentiator. Australia Post reports that six in ten Australian shoppers now use AI, with agentic commerce expected to influence up to 30% of ecommerce transactions by 2030 — a widely repeated projection that should be read as an estimate, not a validated forecast, and one drawn from consultancy modelling rather than measured behaviour. What follows is straightforward. When an assistant can answer "who has this near me, today", the retailer that can answer accurately at store level wins the query. The one that cannot is not in the answer.

The bottom line

The evidence supports a narrow, confident conclusion and does not support a broad one. What is well established is that AI summaries reduce clicks, that being named in the answer is worth more than ranking below it, and that retrieval position — ordinary technical SEO — remains the dominant input. What is genuinely contested is whether content-level GEO tactics work at all once competitors adopt them.

So spend accordingly. Put budget into what pays off under either reading of the evidence: crawler access, machine-readable product and category pages, complete feed attributes, and location data accurate at store level. Those are not GEO bets. They are retail data hygiene that happens to be the entry requirement for AI visibility, and they hold their value if the current tactical consensus is overturned next year.

Treat everything else as an experiment. Fixed prompt set, weekly cadence, your own conversion data, six months minimum before concluding anything. If a vendor offers you an Australian retail GEO benchmark, ask for the sample size and the date, and expect not to get them.

The retailer who wins this is not the one with the best AI content strategy. It is the one whose product, category and store data was already accurate enough to be worth citing.

Where to start

Most of the gap between a retailer that gets recommended and one that does not comes down to whether store-level stock and cart state are accurate enough to answer a shopper's question in real time. If you are unsure whether your store network could answer "who has this near me, today", that is the audit worth running first. Talk to Awayco about unifying stock and cart state across your stores, web and mobile — the data layer everything above depends on.

Newsletter

Subscribe for cutting-edge AI updates

Get the latest thinking on AI-powered retail — from product personalisation to in-store innovation — delivered to your inbox once a month.

Thanks for subscribing to our newsletter!
Oops! Something went wrong while submitting the form.
Only one email per month — No spam!