# LumenGEO > LumenGEO is an AI citation monitoring and GEO (Generative Engine Optimization) platform. We help brands get cited by ChatGPT, Perplexity, Gemini, and other AI search engines. ## For agents - [llms-full.txt](https://lumengeo.co/llms-full.txt): The full LumenGEO guide corpus as plain markdown (expanded version of this file) - [MCP server](https://lumengeo.co/mcp): Query our first-party AI-citation datasets via Model Context Protocol (streamable HTTP; card at /.well-known/mcp/server-card.json) - [Agent skills](https://lumengeo.co/.well-known/agent-skills/index.json): Skill documents for querying our data, building GEO action plans, and running GEO audits - Markdown content negotiation: send `Accept: text/markdown` on `/`, `/blog`, or any `/blog/` page to get markdown instead of HTML ## Free Tools - [Free GEO Audit](https://lumengeo.co/audit): Get your GEO Score in 60 seconds — see if ChatGPT cites your brand - [GEO Action Plan Builder](https://lumengeo.co/tools/geo-action-plan): Answer 3 questions and get a personalized, prioritized GEO action plan — the exact tactics and steps to get cited, as a shareable plan or trackable checklist - [GEO Tools Hub](https://lumengeo.co/tools): All of LumenGEO's free GEO tools in one place - [AI Crawler Checker](https://lumengeo.co/tools/ai-crawler-check): Check if AI bots can access your website - [GEO Readiness Scanner](https://lumengeo.co/tools/geo-readiness): Scan any URL for the 14 signals that drive AI citations - [llms.txt Generator](https://lumengeo.co/tools/llms-txt-generator): Build a formatted llms.txt file so AI search engines can map your site - [GEO Glossary](https://lumengeo.co/glossary): 70 GEO terms defined with quotable definitions ## Blog — GEO Research & Guides - [What is GEO?](https://lumengeo.co/blog/what-is-geo): The complete guide to Generative Engine Optimization - [GEO vs SEO](https://lumengeo.co/blog/geo-vs-seo): What changes when AI answers the query - [How to Get Cited by ChatGPT](https://lumengeo.co/blog/how-to-get-cited-by-chatgpt): Data-driven guide to ChatGPT citations - [Why Your Brand Isn't Cited by ChatGPT](https://lumengeo.co/blog/why-your-brand-isnt-cited-by-chatgpt): 10 common causes and concrete fixes - [AI Search Engines Guide](https://lumengeo.co/blog/ai-search-engines-complete-guide): How each AI platform handles citations - [What is a GEO Score?](https://lumengeo.co/blog/what-is-a-geo-score): How to measure AI search visibility - [AI Search Optimization](https://lumengeo.co/blog/ai-search-optimization): Complete guide to getting cited (2026) - [LLM Optimization](https://lumengeo.co/blog/llm-optimization): How to make AI models cite your brand - [Perplexity SEO](https://lumengeo.co/blog/perplexity-seo): How to get cited in Perplexity AI - [ChatGPT Citation Pipeline](https://lumengeo.co/blog/chatgpt-citation-pipeline): The 6-stage pipeline ChatGPT uses to select sources - [AI Citation Signals](https://lumengeo.co/blog/ai-citation-signals): What content gets cited vs ignored (2026 data) - [Google AI Overviews Optimization](https://lumengeo.co/blog/google-ai-overviews-optimization): Complete guide to AIO optimization - [Bing SEO for AI Search](https://lumengeo.co/blog/bing-seo-for-ai-search): Why Bing indexing is load-bearing for ChatGPT and Copilot - [GEO Tools 2026 Buyer's Guide](https://lumengeo.co/blog/geo-tools-2026-buyers-guide): 8 GEO platforms compared on pricing, features, and use case fit - [Best Perplexity SEO Tracking Tools (2026)](https://lumengeo.co/blog/best-perplexity-seo-tracking-tools): Honest comparison of tools that track Perplexity AI visibility — monitoring vs fixing - [How to Track Brand Citations in ChatGPT, Claude & Perplexity (2026)](https://lumengeo.co/blog/track-brand-citations-in-ai): The manual per-engine method + what to actually track (share of answers, competitors, source URLs) - [How to Build Entity Authority for AI Search](https://lumengeo.co/blog/how-to-build-entity-authority-for-ai-search): The 6-mechanism framework for entity authority - [GEO for E-commerce Brands](https://lumengeo.co/blog/geo-for-ecommerce-brands): Vertical-specific 3-layer strategy for e-commerce - [GEO for B2B SaaS](https://lumengeo.co/blog/geo-for-b2b-saas): How to get your software into AI-generated shortlists - [Claude SEO](https://lumengeo.co/blog/claude-seo): How to get cited by Anthropic's Claude via Brave Search - [Microsoft Copilot SEO](https://lumengeo.co/blog/microsoft-copilot-seo): How to get cited by the most selective AI search engine - [AI Search vs Google Search 2026](https://lumengeo.co/blog/ai-search-vs-google-search-2026): The traffic shift explained - [Multi-modal GEO](https://lumengeo.co/blog/multi-modal-geo): Optimizing images and video for AI citation - [How to Run a GEO Audit](https://lumengeo.co/blog/how-to-run-a-geo-audit): The complete 7-step methodology - [Why AI Citations Decay](https://lumengeo.co/blog/why-ai-citations-decay): The 4.5-week half-life and the freshness program - [Is Your Site Agent-Ready?](https://lumengeo.co/blog/is-your-site-agent-ready): GEO for the age of AI agents (ChatGPT Agent, Comet, MCP) - [Cloudflare's Agent Readiness Score: 21 to 64 in One Session](https://lumengeo.co/blog/cloudflare-agent-readiness-score): The complete punch list for passing Cloudflare's agent-readiness scan — and the checks you should refuse to fake - [Does Schema Markup Help AI Citations?](https://lumengeo.co/blog/does-schema-markup-help-ai-citations): The 2026 causal evidence says no - [Why Your Site May Be Invisible to AI](https://lumengeo.co/blog/why-your-site-may-be-invisible-to-ai): The Cloudflare edge-block layer most sites miss - [You Can't Measure GEO With One Check](https://lumengeo.co/blog/you-cant-measure-geo-with-one-check): AI search is stochastic - [We Ran 1,000+ GEO Audits](https://lumengeo.co/blog/geo-audit-findings-2026): What actually gets brands cited by AI ## Comparisons - [ChatGPT vs Perplexity](https://lumengeo.co/compare/chatgpt-vs-perplexity) - [ChatGPT vs Gemini](https://lumengeo.co/compare/chatgpt-vs-gemini) - [Perplexity vs Gemini](https://lumengeo.co/compare/perplexity-vs-gemini) - [ChatGPT vs Copilot](https://lumengeo.co/compare/chatgpt-vs-copilot) - [ChatGPT vs Claude](https://lumengeo.co/compare/chatgpt-vs-claude) - [GEO vs AEO](https://lumengeo.co/compare/geo-vs-aeo) - [LumenGEO vs Otterly.AI](https://lumengeo.co/compare/lumengeo-vs-otterly) - [LumenGEO vs HubSpot AI Search Grader](https://lumengeo.co/compare/lumengeo-vs-hubspot) - [LumenGEO vs Semrush AI Visibility](https://lumengeo.co/compare/lumengeo-vs-semrush) ## AI Citation Index — who AI search surfaces, by vertical (first-party data) - [AI Citation Index (hub)](https://lumengeo.co/citations): The candidate pool AI search draws from across 30 verticals and 150 buyer queries — most-surfaced domains, brand vs aggregator vs editorial mix, and stability. A retrieval-pool proxy, not citation logging. - [The AI Citation Index Study (2026)](https://lumengeo.co/blog/ai-citation-index-2026): Full write-up — third-party sources dominate the pool, the universal authorities that recur across verticals, and what it means for GEO - [VPN](https://lumengeo.co/citations/vpn) · [Electric Cars](https://lumengeo.co/citations/electric-car) · [Mattresses](https://lumengeo.co/citations/mattress) · [Password Managers](https://lumengeo.co/citations/password-manager) · [Smart Home Security](https://lumengeo.co/citations/smart-home-security) · [CRM](https://lumengeo.co/citations/crm) · [Web Hosting](https://lumengeo.co/citations/web-hosting) · [Project Management](https://lumengeo.co/citations/project-management) · [Credit Cards](https://lumengeo.co/citations/credit-card) · [Life Insurance](https://lumengeo.co/citations/life-insurance) — and 20 more at the hub ## Product - [Pricing](https://lumengeo.co/pricing): Free audit, Playbook ($297), Starter ($49/mo), Pro ($99/mo) - [GEO Optimization Playbook](https://lumengeo.co/playbook): 47-page playbook based on 37 real GEO experiments - [Product Overview](https://lumengeo.co/product): AI citation monitoring SaaS ## Key Definitions - **GEO (Generative Engine Optimization)**: The practice of structuring content so AI search engines cite your brand in their answers - **GEO Score**: A 0-100 metric measuring brand visibility across AI search platforms (Citation Presence, Prominence, Quality, Density) - **AI Citation**: When an AI search engine explicitly names and links to your website as a source in its response - **Fan-Out Queries**: Machine-generated sub-queries that ChatGPT creates from user prompts (89.6% of prompts trigger 2+ sub-queries) - **The LumenGEO Citation Pipeline Model**: A 6-stage framework describing how ChatGPT selects sources to cite - **Citation Halo**: The indirect citation benefit a brand receives when authoritative third-party sources that cite the brand are themselves cited by AI search engines - **Entity Authority**: The strength and consistency of machine-readable signals about a brand across the web — including structured data, entity database presence, and brand mention patterns ## Key Research Findings - 6.5% of unique domains in source documents receive inline citations in AI answers (Georgia Tech, 2024) - Brand mention frequency correlates with AI citation at r=0.664 — 3x stronger than backlinks (r=0.218) (SE Ranking) - Pages with 15+ named entities earn citations at 4.8x the rate of pages with fewer than 8 entities (Wellows) - Original-data content earns 4.1x more citations than content that summarizes other people's data (Digital Bloom) - 46.7% of Perplexity's top-10 cited sources are Reddit pages (Indig/Gauge) - 99.5% of Google AI Overview citations come from pages already ranking on Google page 1 (BrightEdge) - 32.9% of all citations come from fan-out sub-queries with zero traditional search volume (Ekamoira) - Definitive phrasing earns citations at 36.2% vs 20.2% for hedged language — 1.8x advantage (Growth Marshal) - The most-surfaced source across 30 commercial verticals is a third party, not any brand: in LumenGEO's 150-query retrieval index, Zapier's blog appears in the candidate pool for 10 of 30 verticals, ahead of TechRadar, G2, and NerdWallet (LumenGEO, 2026) - 64.2% of all URLs in the AI-search retrieval pool for buyer queries are roundup/listicle pages ("best/top X"); brand-owned pages are only ~49% of distinct domains surfaced (LumenGEO Citation Index, 2026) --- # Full article corpus # How Long Is the Average Page Perplexity Cites? We Measured 954 Cited Pages (2026 Data) > First-party data: across 954 pages Perplexity cites (from 1,385 citations, 160 commercial queries, 8 industries), the median cited page is 1,748 words — not the 3,000-word pillar the folklore promises. The mean (2,471) is inflated by a long tail; 33.2% of cited pages are under 1,000 words and 21% under 500. Length isn't the gate. Full per-industry breakdown and honest methodology. Canonical: https://lumengeo.co/blog/perplexity-cited-page-word-count **The median page Perplexity cites is 1,748 words — not the 3,000-word pillar the folklore promises. The mean is higher, 2,471 words, but only because a long tail of very long pages drags it up; the median is the honest center. And a third of what Perplexity actually cites is short: 33.2% of measured cited pages run under 1,000 words, 21% under 500. At the other end, 12.5% are 5,000 words or more. We measured this directly across 954 cited pages pulled from 1,385 Perplexity citations spanning 160 commercial queries and 8 industries.** *Citations captured 2026-07-02. Page word counts measured 2026-07-16. First-party data — full methodology below.* > **Key takeaway:** "You need a 3,000-word pillar page to get cited" is folklore, and the data doesn't support it. The median page Perplexity cites is 1,748 words; the citation-weighted median — which lets heavily re-cited pages count more — is 1,754, essentially identical. A third of cited pages (33.2%) are under 1,000 words. Length isn't the gate. Perplexity cites pages of every length, from 300-word answer snippets to 6,000-word guides, and picks whichever most directly answers the query. Usefulness-per-query, not word count, is what gets a page pulled. This is a live gap in the public record. Ask an AI engine "what's the average word count of a page Perplexity cites" and you get hedging — "it varies," "aim for comprehensive content," "1,500 to 2,500 words is often recommended" — because no one had measured the pages Perplexity actually cites and published a number. We built the dataset to close it. ## Why this question exists in the first place The "long content wins" belief is one of the stickiest in SEO, and it got imported into GEO wholesale. The pitch is intuitive: a longer page covers more sub-questions, so an answer engine is more likely to find the passage it needs. But "longer pages surface more passages" and "the pages that actually get cited are long" are two different claims, and only the second is testable against real citations. So we tested it: we took every page Perplexity cited across a broad set of commercial queries and measured how long each one actually is. ## The headline distribution We ran 160 commercial queries through Perplexity across 8 industries and captured every citation. Then, on 2026-07-16, we fetched each unique cited URL and counted the words in its main content. Here's how the lengths fall out across the 954 pages we could measure. | Word-count band | Pages | Share of cited pages | |---|---|---| | **< 500 words** | 200 | **21.0%** | | 500 – 1,000 | 116 | 12.2% | | 1,000 – 2,000 | 211 | 22.1% | | 2,000 – 3,000 | 150 | 15.7% | | 3,000 – 5,000 | 158 | 16.6% | | **5,000+ words** | 119 | **12.5%** | | **All measured pages** | **954** | **100%** | The center of the distribution: | Statistic | Value (words) | |---|---| | **Median (unique pages)** | **1,748** | | Mean (unique pages) | 2,471 | | 25th percentile | 687 | | 75th percentile | 3,307 | | 90th percentile | 5,334 | The gap between the median (1,748) and the mean (2,471) is the whole story in two numbers. When a mean sits ~40% above the median, it means a minority of very long pages is pulling the average up while most cited pages sit well below it. The 90th percentile is 5,334 words — so the top 10% of cited pages are the multi-thousand-word guides everyone pictures, and they distort the average, but they are not the norm. Half of everything Perplexity cited was shorter than 1,748 words, and a quarter was shorter than 687. *My read: the 3,000-word target isn't wrong so much as aimed at the wrong percentile. A 3,000-word page sits around the 72nd percentile of what Perplexity cites here — above average, not required. You're competing against a median of 1,748, and a fifth of the pages that beat you to a citation are under 500 words.* ## Unique-page vs citation-weighted: they barely move There are two honest ways to average this. Count each cited page once (unique-page view), or weight each page by how many times it was cited so a page cited five times counts five times (citation-weighted view). If Perplexity systematically re-cited its *longer* pages, the citation-weighted median would jump above the unique-page median. It doesn't: unique-page median is 1,748, citation-weighted median is 1,754 — a six-word difference, statistical noise. That itself is a finding. The pages Perplexity leans on repeatedly are not systematically longer than the pages it cites once. Length doesn't buy you repeat citations any more than it buys you the first one. ## Length by industry Word count of cited pages varies more by *what's being searched* than by any universal length rule. Marketing/SEO queries pull long, technical explainers; home-services and legal queries pull much shorter pages. | Industry | Pages (n) | Median words | p25 – p75 | |---|---|---|---| | **Marketing / SEO** | 118 | **2,965** | 1,289 – 5,028 | | Travel | 113 | 1,880 | 858 – 3,053 | | SaaS / software | 131 | 1,874 | 923 – 3,581 | | Ecommerce / retail | 113 | 1,777 | 621 – 3,342 | | Finance / fintech | 120 | 1,711 | 781 – 3,286 | | Healthcare / wellness | 129 | 1,414 | 464 – 3,252 | | Legal services | 135 | 1,395 | 521 – 2,729 | | **Home services** | 109 | **1,256** | 354 – 2,767 | The spread is real: the median cited page in Marketing/SEO (2,965 words) is more than double the median in Home services (1,256). Marketing/SEO is the one vertical where the 3,000-word pillar is roughly the median rather than the exception — its readers and its cited pages both skew toward long-form how-to content. At the other end, home-services and legal queries ("how much does X cost," "do I need a lawyer for Y") get answered by short, direct pages, and their p25 values (354 and 521 words) show a big chunk of cited pages barely clearing a few hundred words. *What I take from this: there is no single word-count target — there's a target per query type. If your category behaves like home services or legal, a tight 900-word page that answers one question cleanly is competing at the median. If you're in Marketing/SEO, the bar for a comprehensive piece genuinely is higher. Match the length distribution of your vertical, not a blog-wide rule of thumb.* For which sources dominate each of these categories in the first place, see our companion breakdown of [Perplexity's citation patterns and how sources get selected](/blog/perplexity-seo), and the broader field numbers in our roundup of [AI search statistics for 2026](/blog/ai-search-statistics-2026). ## Do the top citation slots go to longer pages? Slightly — but less than you'd guess. We split the measured citations by position: the pages cited in an answer's top-3 slots versus everything from position 4 on. | Citation position | Pages (n) | Median words | Mean words | |---|---|---|---| | **Top 3** | 360 | **1,930** | 2,868 | | Positions 4+ | 608 | **1,625** | 2,252 | Top-3 cited pages run a median of 1,930 words versus 1,625 for the deeper slots — a real but modest edge of about 300 words. Longer pages are a little more likely to land in the prominent positions, but a 1,930-word median for the *best* slots is still nowhere near the pillar-page mythology. You do not need to be the longest page in the answer to be the first one cited. ## Methodology & limitations Built to be reproducible, and honest about where it's soft. - **Source citations.** 160 commercial queries (8 industries × 20 buyer-intent queries: "best [category]," "is [product] worth it," "[A] vs [B]") run through Perplexity with web search enabled, US locale, captured live via DataForSEO on **2026-07-02**. That produced **1,385 citations across 1,369 unique URLs**. - **Word-count pass.** On **2026-07-16** we fetched each unique cited URL and counted words from the server-rendered HTML using a disclosed heuristic: strip `