AI Search

AI Engines Know Exactly Who I Am. They Still Won't Cite My Work

Some links here are affiliate links: if you buy through them I may earn a commission, at no extra cost to you. I only recommend tools I'd point a colleague to, and rankings are never paid for.

On this page
  1. The two numbers
  2. The branded result, and why it matters more than it looks
  3. Why entity ≠ source
  4. What I actually changed between the runs
  5. Methodology — and where this run is weaker than the first
  6. What the engines cited instead
  7. The finding I didn’t expect: the tool roster is churning fast
  8. So what actually gets you cited?
  9. What happens next

Ask Perplexity who I am and it cites my own site first, then describes me using my own About page. Ask any of the three big engines a question about my actual field and I appear nowhere — zero citations across eighteen answers, two runs, eight weeks apart.

Those two facts sitting next to each other are more useful than either alone. The engines can read my site, they have indexed it, they will quote it, and they treat it as authoritative — about me. On the questions I actually want to be found for, I do not exist.

That gap has a name worth being precise about: entity recognition is not source citation. They are different systems doing different jobs, and the GEO advice that treats them as one funnel is why people do a lot of structural work and see nothing move.

The two numbers

Two runs, two very different results
Query typeAnswers checkedmarcodiversi.com cited
Category prompts (2026-07-12)90
Category prompts (2026-09-03)90
Branded / entity query (2026-09-03)1Yes — cited first of six sources

Eighteen category answers. Zero. One branded query. First source.

The branded result, and why it matters more than it looks

I asked Perplexity who is Marco Diversi, in Italian, in a clean session with no account history.

It cited marcodiversi.com first, ahead of five other sources — a mix of third-party professional profiles and social properties. And the description it produced tracks my own About page closely: the engineering background, the years in SEO and affiliate, the scale of the site network, the company I founded. It reproduced details specific enough — down to the dialect listed in my languages line — that there is no ambiguity about where it came from.

That is the most encouraging data point in either run, and it kills a hypothesis I had been carrying. The site is not invisible. It is not blocked, not unreadable, not structurally disqualified. When the engines have a reason to reach for it, they reach for it, they rank it first, and they trust it enough to build a biography out of it.

So the zero on the category prompts is not a crawling problem or a formatting problem. It is a different problem entirely.

One honest footnote I will deal with separately: several of the third-party profiles cited alongside my site are years out of date and describe work I no longer do. That is a real problem, but it is a brand problem rather than a citation-mechanics one, and it deserves its own piece rather than a paragraph here.

Why entity ≠ source

The distinction, as plainly as I can put it.

Entity recognition answers “who or what is X?” The engine needs a canonical description of a thing. Your own site is the obvious authority on yourself, so it wins — close to a solved problem if your About page exists and is coherent.

Source citation answers “what’s the best X?” or “how do I do Y?” Now the engine is choosing evidence, from among everyone writing about the topic. Being a known entity buys you nothing here. Nobody gets cited in a tool roundup because the engine knows who they are.

The GEO checklists conflate these. Schema, llms.txt, an About page, consistent naming — that stack genuinely helps the first. It is close to irrelevant to the second. I did all of it and moved from zero to zero on the questions that matter commercially.

What I actually changed between the runs

Not an experiment — real work, all of it on the standard checklist:

Change When
Unblocked every AI crawler at the CDN edge Between runs
Migrated hosting infrastructure 2026-09-02
Published original datasets (pricing index, citability benchmark, crawler reference) Between runs
Shipped four free client-side tools Between runs
Rewrote the tool cluster answer-first, added schema and llms.txt Between runs
Made the reference set explicitly free to cite under CC BY 4.0 Between runs

The branded result proves that work landed — the site is readable and trusted. The category result proves it was not the lever for the thing I wanted.

For completeness: there was also a Product Hunt launch in July. It produced a handful of comments and essentially one visitor, so it is not part of this story — I mention it only so nobody assumes an unmeasured traffic spike explains anything below.

Methodology — and where this run is weaker than the first

Same three prompts, same three engines, browsing on, one clean pass:

  1. best GEO tools 2026
  2. how to get cited by ChatGPT
  3. AI visibility tools for small teams

Plus the branded query, run separately in a clean session.

An honest limitation, stated up front: this run’s per-engine attribution is weaker than July’s. In the first run I tagged every answer to its engine as I captured it. This time I ran all three prompts on all three engines but did not tag each response as I went. So I can report the union of what was named and cited with confidence, and I cannot attribute any specific category claim to a specific engine.

I am not going to reconstruct that from memory to make the piece tidier. The tracker still shows the July run for per-engine charts and says so; this run contributes the union, the branded result and the before/after.

The rest of the caveat is unchanged. One operator, one pass per engine, one day, browsing on — not a controlled study across accounts, regions and repeats. The branded query was in Italian, which may well affect which sources surface. The value is the dated series, not any single answer.

What the engines cited instead

Two categories, and the split is the lesson.

Third-party roundupsismybrandinai.com, geoptie.com. Pages whose entire purpose is to list and rank tools in this category.

Official and vendor documentationdevelopers.openai.com, help.openai.com, developers.google.com for the “how to get cited” prompt; otterly.ai, ahrefs.com, semrush.com for the tool prompts.

Same shape as July: roundups and vendor-owned pages, not independent operator write-ups. For “how to get cited by ChatGPT,” the engines went to OpenAI’s and Google’s own developer docs — which, on reflection, is exactly correct. The primary source outranks anyone’s commentary on it, and no amount of answer-first formatting changes that.

The uncomfortable implication: for a chunk of informational queries, the ceiling is not your structure, it is that a first-party source exists and you are not it.

The finding I didn’t expect: the tool roster is churning fast

Comparing the two runs turned up something more interesting than my own absence.

Tools named by the engines: July vs September
Count
Named in July12
Named in September19
Named in both (persisted)9
New in September10
Dropped since July3

75% of July’s tools survived to September — a stable core: Profound, Otterly, Peec, Semrush, Writesonic, Bluefish, AthenaHQ, Scrunch and Geoptie.

But more than half of September’s roster was not there in July. Ten names entered in eight weeks: HubSpot’s AEO Grader, Gauge, Brandi AI, Goodie AI, Ahrefs Brand Radar, Rankscale, Trakkr.ai, SE Ranking, Brand24 and BrandViz.AI. Three dropped out: BrightEdge, RankScope and BotRank.

A precision note: July recorded RankScope and September recorded Rankscale. Those are different products, and I have counted them as one leaving and one entering rather than quietly merging them. If the July capture was a transcription slip on my part, the churn is one lower in each direction — I would rather flag the ambiguity than hide it inside a cleaner statistic.

Two things follow:

The category is not settled, and the engines know it. Half the roster turning over in eight weeks means active re-retrieval, not a cached consensus. The door is not closed.

But the entrants are not independent operators. HubSpot, Ahrefs, SE Ranking, Brand24 — established brands shipping an AI-visibility feature — plus funded point tools with their own marketing sites. Nobody entered by writing a good blog post. They entered by being a tool that roundup sites list.

So what actually gets you cited?

From two dated runs plus the branded result:

  1. Entity work pays off, and it is the easy half. A coherent About page, Person schema and consistent naming got me cited first on my own name. If you have not done that, do it — it is cheap and it works.
  2. Category citation is a different game. The engines cite primary sources and third-party roundups. If you are neither, you are competing for a slot that may not exist.
  3. Getting into the roundups is the lever. Every persistent name across both runs is on multiple third-party listicles. That is an outreach and product problem, not a content-formatting problem.
  4. On-site GEO work is table stakes. It stops you being excluded. It does not get you included.
  5. The timescale is months, not weeks. I unblocked the crawlers inside this window, so this run may simply be too early. Which is why the series continues.

If you want the structural work done properly anyway — and you should, because it is the part you control — the get-cited playbook covers it, and the free citability scorer will tell you where a page falls short.

What happens next

I will run this again, tagging each answer to its engine as I capture it so the per-engine detail matches the first run’s. I will add branded queries on more engines and in English, since one Italian Perplexity answer is a data point, not a pattern.

The whole dataset — both runs, every tool named, every source cited, the branded result — is at the AI citation tracker, free to cite under CC BY 4.0. If you are running the same test on your own site, the prompt set builder generates a comparable set for your category with a CSV tracking sheet.

A single “we’re not cited” post is a complaint. Two dated runs that disagree with each other in an interesting way is data.

See the tools the engines actually name

Frequently asked questions

Do AI engines cite marcodiversi.com?

Yes on branded queries — Perplexity cited it first of six sources for "who is Marco Diversi" and built its description from the site’s own About page. No on category queries — across 18 answers over two runs covering GEO tools, AI visibility and how to get cited by ChatGPT, the site was not cited once. Those are different systems, and the distinction matters more than either number alone.

What is the difference between entity recognition and source citation?

Entity recognition answers "who or what is X" and needs a canonical description, so your own site usually wins because it is the obvious authority on yourself. Source citation answers "what is the best X" or "how do I do Y", where the engine is choosing evidence from among everyone writing on the topic. Being a recognised entity buys you nothing in the second case, which is why schema and an About page can be working perfectly while category visibility stays at zero.

Did the GEO work fail?

No, but it did not do what was hoped. The branded result proves the site is crawlable, readable and trusted enough to be cited first on a question about itself, which rules out the technical explanations for the category zero. What it shows is that on-site structural work is necessary but not sufficient for category citation. Crawler unblocking also happened inside this window, and citation lags crawling by months rather than weeks.

What do ChatGPT, Gemini and Perplexity actually cite for tool questions?

In both runs, two things: third-party roundup pages whose purpose is to list and rank tools in the category, and vendor or official documentation. For "how to get cited by ChatGPT" the engines went to OpenAI and Google developer docs directly. Independent operator write-ups did not appear in either run.

How much does the list of tools change between runs?

A lot. Between 2026-07-12 and 2026-09-03, 9 of the 12 tools named in July were named again in September, but 10 of the 19 named in September were new. Roughly half the roster turned over in eight weeks, which suggests the engines are actively re-retrieving rather than serving a settled consensus.

Why is this run’s per-engine detail weaker than the first?

Because the category answers were not tagged engine-by-engine as they were captured. All three prompts were run on all three engines, so the union of tools and sources is accurate, but no individual category claim can be attributed to a specific engine. The tracker therefore still shows the July run for per-engine charts and says so, rather than blending the two.

Published Last updated

← All posts

navigate openesc close