verisible AI visibility, measured honestly
← Field notes
2026-07-05 · geo · chatgpt · ai-search

How ChatGPT chooses which websites to cite

Every URL ChatGPT search considers citing was first discovered through Bing’s index (The Keyword). Not Google. If Bing has never indexed your page, ChatGPT cannot cite it, no matter how well you rank in Google. That single dependency decides more visibility outcomes than any content tactic.

ChatGPT chooses sources in a chain: it discovers candidate pages through Bing’s search index, favors recent and well-structured content, fetches the live page to read it, and cites the passages it can actually pull. Content under 30 days old earns roughly 3.2x more AI citations than older pages (study, apiSERPent). Bing indexation is the gate; freshness and structure are the filters.

That order matters. Marketers spend on the filters while ignoring the gate. Check Bing Webmaster Tools before you touch a word of copy.

Why does Bing sit upstream of everything?

OpenAI did not build a web index from scratch. Building one is a multi-billion-dollar undertaking, so ChatGPT search leans on Bing’s existing index for URL discovery (Yoast). Microsoft’s infrastructure handles crawling, indexing, and serving; ChatGPT queries it, then applies its own model to decide what to cite.

The practical consequence is blunt. Your Google rankings are irrelevant to ChatGPT except to the extent Bing agrees with them, and Bing often does not. A page Google loves and Bing has never crawled is invisible in ChatGPT. Bing Webmaster Tools is free, most SEO teams ignore it, and it is the single best ground-truth check for ChatGPT eligibility you have.

What makes the model favor one page over another?

Once candidates exist, recency does heavy lifting. Multiple 2026 analyses find the same pattern: freshly updated content gets cited far more often, with the sub-30-day advantage landing around 3.2x. Roughly half of content cited in AI answers is less than 13 weeks old. The model has a reason for this bias: if it recommends the “best tools in 2026” and cites a page last touched in 2024, the recommendation risks being wrong, and wrong is the one thing an answer engine cannot afford.

The signals the model reads for freshness are boringly concrete: the dateModified field in your Article schema, the visible last-updated date on the page, then the original publish date. Backdating a page you did not actually update is a short-lived trick and an easy one to get caught faking.

Structure is the other filter. The engine cites passages, so short paragraphs carrying one idea each beat long, hedged blocks. A clear heading followed by a direct two-sentence answer is retrieval-friendly. A 400-word wall that buries the point is not.

Does ChatGPT actually read my page, or just the snippet?

It fetches the live page. This is the part that separates AI answer engines from classic search, and it has a failure mode. Because engines like ChatGPT and Perplexity load your page in real time rather than reading a stored copy, a slow server can get abandoned mid-fetch. Mike King of iPullRank flagged the tell: HTTP 499 responses in your logs mean the engine timed out and gave up. Your normal SEO tools will not warn you about this, because they are not the ones fetching.

A page that ranks well but throws 499s under live fetch is a page that ChatGPT tried to cite and walked away from. Fixing server response time on those pages tends to show up in visibility within days, not months, because there is no slow index to wait on. The engine simply succeeds on the next fetch.

Can you verify whose search actually answered?

Yes, and this is where it gets useful. When we run questions through engines that use native provider search, the citations sometimes carry a watermark in the URL. On one paid-tier run, the returned links included utm_source=openai, OpenAI’s own search parameter, stamped right into the citation. That tag is a fingerprint: it tells you the answer came from OpenAI’s real search surface and not some substituted middleman index.

We now treat non-empty, correctly watermarked citations as a canary. If an engine claims to be running native search and the citations arrive blank or unbranded, the measurement is suspect and we flag it rather than report it. Most tools never check. They report whatever comes back as gospel, which is how you end up optimizing for a search surface that was never actually queried. The mechanics of that fan-out step sit one layer up, in query fan-out.

The concrete next step is smaller than it sounds. Confirm your key pages are in Bing’s index, check your logs for 499s on the pages you care about, and read what an engine actually cites when asked about you. Our free scan does the last part and shows the citations verbatim: run it.

Sources

See it on your own site. Five buyer questions to an AI, and we show you whether you're in the answers.
Start tracking free