Generative Engine Optimization (GEO) aims to improve content visibility in AI-generated responses. However, existing methods measure contribution (how much a document influences a response) rather than citation — the mechanism that actually drives traffic back to content creators.
In our analysis, 43% of topically relevant webpages receive no citation under baseline conditions. For these webpages, the question is not "how much am I cited?" but "why am I not cited at all?"
Existing GEO methods apply generic rewriting rules uniformly (add statistics, adopt authoritative tone, improve fluency). This fails because citation failures are heterogeneous, spanning multiple pipeline stages: a webpage may fail at fetching (malformed HTML), parsing (content buried in boilerplate), or generation (inferior information density vs. competitors).
We address these challenges with three contributions:
>40% relative citation improvement while modifying only 5% of content (vs. 25% for baselines). Citation failure is rarely a global quality problem — most pages need targeted fixes, not rewriting.
Generic optimization can actively degrade citation for specialized topics. Diagnostic optimization conditions on each page's failure mode, avoiding aggregate-pattern bias and generalizing equitably.
Some pages face dominant competitors that no content-side fix can overcome. Citation mechanisms may amplify certain voices — creator-side optimization alone cannot ensure equitable visibility.
We constructed a diagnostic dataset of 949 contrastive pairs from GEO-Bench, each pairing a non-cited webpage with a cited competitor for the same query. Analysis reveals four failure dimensions:
Access blocking, JS rendering failures, unparseable content, low signal-to-noise ratio from boilerplate.
Intent divergence, contextual gaps, outdated information, localization mismatch between content and query.
Information scarcity, content fragmentation, excessive verbosity, unstructured layout.
Competitive redundancy and context window truncation — barriers no optimization can overcome.
AgentGEO operates in an iterative diagnose-then-repair loop. For each training query where a target webpage is not cited:
Existing benchmarks pair each document with a single query — risking query-specific overfitting. MIMIQ (Multi-Intent Multi-Query) treats the document as the unit of optimization:
If you find this work useful, please cite our paper:
@article{tian2026agentgeo,
title={Diagnosing and Repairing Citation Failures in Generative Engine Optimization},
author={Tian, Zhihua and Chen, Yuhan and Tang, Yao and Liu, Jian and Jia, Ruoxi},
arxiv={https://arxiv.org/abs/2603.09296},
year={2026}
}