Diagnosing and Repairing Citation Failures in Generative Engine Optimization


TLDR: We introduce AgentGEO, an agentic framework that diagnoses why webpages fail to be cited in AI-generated responses and applies targeted repairs. Combined with a citation failure taxonomy and a document-centric benchmark (MIMIQ), AgentGEO achieves >40% relative improvement in citation rates while modifying only 5% of content, compared to 25% for baselines.

Zhihua Tian*1, Yuhan Chen*1, Yao Tang*2, Jian Liu2, Ruoxi Jia1
1Virginia Tech  2Zhejiang University 

Overview

Generative Engine Optimization (GEO) aims to improve content visibility in AI-generated responses. However, existing methods measure contribution (how much a document influences a response) rather than citation — the mechanism that actually drives traffic back to content creators.

In our analysis, 43% of topically relevant webpages receive no citation under baseline conditions. For these webpages, the question is not "how much am I cited?" but "why am I not cited at all?"

Existing GEO methods apply generic rewriting rules uniformly (add statistics, adopt authoritative tone, improve fluency). This fails because citation failures are heterogeneous, spanning multiple pipeline stages: a webpage may fail at fetching (malformed HTML), parsing (content buried in boilerplate), or generation (inferior information density vs. competitors).

We address these challenges with three contributions:

  • Citation Failure Taxonomy — the first systematic categorization of why webpages fail to be cited, spanning fetching, parsing, and generation stages.
  • AgentGEO — an agentic system that diagnoses failures using this taxonomy, selects targeted repairs from a tool library, and iterates until citation is achieved.
  • MIMIQ Benchmark — a document-centric evaluation protocol with train/test query splits to test generalization, not overfitting.

Empirical Takeaways

Diagnosis Beats Generic Rules

>40% relative citation improvement while modifying only 5% of content (vs. 25% for baselines). Citation failure is rarely a global quality problem — most pages need targeted fixes, not rewriting.

Generic Rules Harm Long-Tail Content

Generic optimization can actively degrade citation for specialized topics. Diagnostic optimization conditions on each page's failure mode, avoiding aggregate-pattern bias and generalizing equitably.

Not All Failures Are Recoverable

Some pages face dominant competitors that no content-side fix can overcome. Citation mechanisms may amplify certain voices — creator-side optimization alone cannot ensure equitable visibility.

Method

We constructed a diagnostic dataset of 949 contrastive pairs from GEO-Bench, each pairing a non-cited webpage with a cited competitor for the same query. Analysis reveals four failure dimensions:

Taxonomy of citation failure modes in generative engines
1

Technical Integrity (10.1%)

Access blocking, JS rendering failures, unparseable content, low signal-to-noise ratio from boilerplate.

2

Semantic Alignment (62.2%)

Intent divergence, contextual gaps, outdated information, localization mismatch between content and query.

3

Content Quality (27.1%)

Information scarcity, content fragmentation, excessive verbosity, unstructured layout.

4

Systemic Exclusion (0.6%)

Competitive redundancy and context window truncation — barriers no optimization can overcome.

AgentGEO operates in an iterative diagnose-then-repair loop. For each training query where a target webpage is not cited:

  • Failure Diagnosis: Simulate the GE pipeline, compare against cited competitors, classify the vulnerability using our taxonomy.
  • Tool Selection with Memory: Select a repair tool based on the diagnosed vulnerability. A query-specific memory tracks prior attempts to avoid repeating ineffective repairs.
  • Iterative Refinement: Apply the tool, check if citation is achieved. If not, re-diagnose and iterate.
  • Batch Aggregation: Aggregate suggestions across training queries to produce updates that generalize across diverse query intents.
The workflow of AgentGEO

Existing benchmarks pair each document with a single query — risking query-specific overfitting. MIMIQ (Multi-Intent Multi-Query) treats the document as the unit of optimization:

  • Multiple queries per document spanning diverse intents, personas, and phrasings.
  • Methods optimize on training queries and are evaluated on held-out test queries.
  • This tests whether optimization produces genuinely more citable content — not just overfitting to specific query formulations.

Demo: Before vs. After AgentGEO Optimization

Health
Law & Gov.
Computers & Tech
Home & Garden
Business

Citation

If you find this work useful, please cite our paper:

@article{tian2026agentgeo,
  title={Diagnosing and Repairing Citation Failures in Generative Engine Optimization},
  author={Tian, Zhihua and Chen, Yuhan and Tang, Yao and Liu, Jian and Jia, Ruoxi},
  arxiv={https://arxiv.org/abs/2603.09296},
  year={2026}
}