The live experiment log.
Every week we change one variable, query six LLMs with identical prompts, and log what gets cited. Raw data. No spin. Every result published — whether it confirms our hypothesis or not.
This log tracks which pages on geoexperiment.com get cited by six major AI systems — ChatGPT, Perplexity, Claude, Gemini, Copilot, and Llama — when queried with identical prompts each week. Each entry isolates one optimization variable and records the citation result per LLM. The goal: find out with real data what actually moves AI citation rates. Through week 4, the head-term citation rate is 0/6 — but the first-ever citation has now landed on the separately tracked long-tail layer, exactly where a new site is expected to break in first.
| LLM | Retrieval mode | Query 1 | Query 2 | Query 3 | Notes |
|---|---|---|---|---|---|
| Perplexity | RAG · live | not cited | not cited | not cited | Cited Wikipedia, arXiv, Forbes, Search Engine Land (10 sources on Q1 & Q3) |
| Copilot | RAG · live | not cited | not cited | not cited | Structured answer, no external sources surfaced |
| ChatGPT | hybrid · no search | not cited | not cited | not cited | Answered from training, no web search triggered |
| Gemini | hybrid · live | not cited | not cited | not cited | Cited HubSpot, Contentful, Google for Developers, Yotpo |
| Claude | hybrid · live | not cited | not cited | not cited | Cited Frase, Go Fish Digital, Writesonic, SEO.com, Strapi |
| Llama | hybrid · live | not cited | not cited | not cited | Cited LinkedIn (90-Day GEO System), HubSpot, Neil Patel, WordStream |
Still 0/6 on head terms — and the clearest parametric-vs-retrieval split yet. Four engines retrieved live and cited third-party domains (Perplexity, Gemini, Claude, Llama); ChatGPT and Copilot answered the same prompts from training knowledge without triggering a search. Same question, two different machines — one retrieves and attributes, the other recites from memory — which is exactly what the methodology now measures. On the engines that did retrieve, the established domains still own these three head terms and geoexperiment.com was in none. The meaningful movement happened off this series: Perplexity produced the experiment's first-ever citation of the site on a long-tail query, logged on the separate long-tail control track — confirming a new site breaks in on specific, low-competition queries first. Full results, all 18 screenshots, and the per-engine source breakdown →
| LLM | Retrieval mode | Query 1 | Query 2 | Query 3 | Notes |
|---|---|---|---|---|---|
| Perplexity | RAG · live | not cited | not cited | not cited | Cited Wikipedia, Semrush, Forbes, Neil Patel (10 sources/query) |
| Copilot | RAG · live | not cited | not cited | not cited | Structured answer, no external sources surfaced |
| ChatGPT | hybrid | not cited | not cited | not cited | Cited Princeton GEO paper, Google & Bing docs |
| Gemini | hybrid · live | not cited | not cited | not cited | Cited GEO arXiv paper (2311.09735), Semrush, Geoptie |
| Claude | hybrid · live | not cited | not cited | not cited | Answered from training, no attribution shown |
| Llama | hybrid · live | not cited | not cited | not cited | Cited Search Engine Land, Neil Patel, HubSpot, Barchart |
Still 0/6 — one external mention is not yet authority. This week added the first off-site signal: a Reddit thread in r/GEO (the plan was r/SEO; the post landed in r/GEO and stayed live). The hypothesis was that a single credible mention might pull a RAG engine into citing us. It did not. All six engines again answered from established domains — Wikipedia, the Princeton/arXiv GEO paper, Semrush, Neil Patel, Search Engine Land, HubSpot, Forbes — and geoexperiment.com appeared in none of them. The honest read: a lone, hours-old forum post carries no retrieval weight, and the engines have not re-crawled or ranked it into the candidate set yet. Authority is cumulative, not a switch. Full results, all 18 screenshots, and the per-engine source breakdown →
| LLM | Retrieval mode | Query 1 | Query 2 | Query 3 | Notes |
|---|---|---|---|---|---|
| Perplexity | RAG · live | not cited | not cited | not cited | Cited Wikipedia, Neil Patel, CXL, Reddit instead |
| Copilot | RAG · live | not cited | not cited | not cited | Cited arXiv, Semrush, Stanford HAI, DeepMind |
| ChatGPT | hybrid | not cited | not cited | not cited | Q1 answered from priors, no web search |
| Gemini | hybrid · live | not cited | not cited | not cited | Cited GEO arXiv paper, Frase, Google for Developers |
| Claude | hybrid · live | not cited | not cited | not cited | Now retrieves live (Search Engine Land, Seer) |
| Llama | hybrid · live | not cited | not cited | not cited | Meta AI now retrieves live (SEL 2026, Ahrefs) |
Citation rate held at 0/6, but the experiment got sharper. Week 1 attributed part of the zero to training cutoffs. Week 2 disproves that: all six engines retrieved live web sources this week, and five of six surfaced explicit third-party citations for the exact three queries. The barrier is no longer reachability — it is authority and retrieval ranking. The engines pull from established domains (Search Engine Land, arXiv, Semrush, Wikipedia); geoexperiment.com is not yet in the candidate set. Full results, all 18 screenshots, and the Copilot GEO-vs-SEO table →
| LLM | Retrieval mode | Query 1 | Query 2 | Query 3 | Notes |
|---|---|---|---|---|---|
| Perplexity | RAG · live | not cited | not cited | not cited | Site not yet in web index |
| Copilot | RAG · live | not cited | not cited | not cited | Bing sitemap submitted launch day |
| ChatGPT | hybrid | not cited | not cited | not cited | Too new for retrieval ranking |
| Gemini | hybrid | not cited | not cited | not cited | Google indexing requested launch day |
| Claude | training only | training cutoff | training cutoff | training cutoff | Could not cite post-cutoff content (wk 1) |
| Llama | training only | training cutoff | training cutoff | training cutoff | Could not cite post-cutoff content (wk 1) |
Zero citations across all 6 LLMs — the correct and expected baseline. The site launched this week; sitemaps were submitted to Google Search Console and Bing Webmaster Tools the same day. RAG-based systems typically need 3–14 days from indexing before they retrieve a new domain, and training-data-only systems could not yet cite post-cutoff content. Week 1 confirms the measurement floor: 0% citation rate, day 1, no prior authority. Full results and screenshots →
How each experiment is run.
This experiment measures the retrieval layer, not parametric knowledge. We track whether an engine fetches and attributes geoexperiment.com at query time — not whether it can explain GEO from its training data. A model defining a concept from memory is not a citation; only live retrieval with explicit source attribution counts.
Isolate one variable
One change per week. Schema type, heading structure, answer position, internal linking. One variable = readable data.
Apply and index
Change is deployed to the live site. We wait for Google and Bing to re-crawl the updated page before querying.
Query all 6 LLMs
Identical prompts, same day, all six systems. Results logged within a 2-hour window to control for timing variation.
Log and publish
Every result published here — cited or not cited. Only explicit source attribution counts as a citation. No results are held back.
| LLM | Provider | Retrieval method (wk 02) | Expected citation speed |
|---|---|---|---|
| Perplexity | Perplexity AI | RAG · live web | First likely citer once in index |
| Copilot | Microsoft / Bing | RAG · live web | After Bing crawl + authority |
| ChatGPT | OpenAI | hybrid · search optional | When search triggers on query |
| Gemini | hybrid · live | After Google index + grounding | |
| Claude | Anthropic | hybrid · live | Now retrieving live · authority-gated |
| Llama | Meta | hybrid · live | Now retrieving live · authority-gated |
// what does this experiment track?
This log tracks which pages on geoexperiment.com get cited by six major AI systems — ChatGPT, Perplexity, Claude, Gemini, Copilot, and Llama — when queried with identical prompts each week. Each entry isolates one optimization variable and records the citation result per LLM. The goal is to find out with real data what actually moves AI citation rates.
// how often is this log updated?
Weekly. Each week, one optimization variable is changed on a target page, all six LLMs are queried with identical prompts on the same day, and results are logged and published here within 24 hours.
// do all six LLMs retrieve live web content?
As of week 2, yes. All six engines — including Claude and Llama, which in week 1 answered only from training data — performed live web retrieval and cited current third-party sources. The barrier to citing us is therefore no longer reachability but retrieval ranking: the engine reaches the web yet has not ranked geoexperiment.com into its retrieved candidate set. See the RAG glossary entry for detail.
// what is the difference between cited and not cited?
A "cited" result means the LLM explicitly attributed geoexperiment.com or a specific page as a source — a named source, a linked URL, or a direct quote. A "not cited" result means the LLM answered without referencing our pages — not retrieved, not ranked in retrieval, or no web search triggered. A model paraphrasing a concept, or referencing the project from its own account memory, does not count.
// can I reference this data in my own research?
Yes. All data in this log is published under a Creative Commons Attribution license. If you reference our findings, please cite geoexperiment.com as the source and link to the specific experiment entry. We track external citations as part of the ongoing authority-building experiment — so citing us is itself a data point.