Three situations, only one of which concerns you
When someone asks a question, the system does not always behave the same way.
It answers from memory. On a general, stable question — a definition, a principle — the model draws on its training. No source is cited, and nothing you publish today changes that in the short term.
It runs a search. As soon as the question calls for current, local, priced or named facts — "a provider in Tunis", "how much does it cost", "who offers" — the system goes out to the web, retrieves pages, extracts passages and cites its sources. That is the only situation that can be worked on, and it is also where the commercially valuable questions sit.
It is given an address. The person pastes a URL and asks for an analysis. The content is then fetched directly. Worth knowing: a page the crawlers cannot reach will fail in this case too.
The first condition, and the one most often failed
Separate crawlers are involved. GPTBot crawls the web for training, OAI-SearchBot feeds the search, and ChatGPT-User fetches a page on demand. They are configured separately in robots.txt, and allowing Googlebot does not allow them.
Blocking is very common and almost never deliberate: a security plugin, a web application firewall, a host setting, a line copied from a tutorial. The result is final — no content retrieved, therefore no citation possible, whatever the quality of the site. It is the first thing our free audit checks, in a few seconds.
What makes a passage get picked
Once the page is retrieved, the system does not keep the page: it keeps a fragment. Four properties separate a usable fragment from an ignored one.
It answers the question, early
A paragraph that gives the answer in its first three sentences is quotable. A paragraph that sets the scene for eight lines before getting to the point is not — the extracted fragment would contain nothing useful.
It makes sense out of context
This is the most common writing mistake. "This solution reduces lead times" means nothing once it leaves the page. "An accessibility audit halves the time to conformance" can be quoted as it stands. Name the subject in the sentence, every time.
It is attributable
A model more readily cites a source it can name. Organisation name, author, date, contact details: these have to be declared, not merely visible. That is what structured data is for.
It is corroborated
A claim found elsewhere — press, professional directories, business listings, public documents — is safer to reuse. Which is why a citation strategy is never played out on your own site alone.
Testing where you stand, in fifteen minutes
- Write six questions as a customer would ask them, never naming your brand. For example: "who builds bespoke web platforms for public bodies", "how much does an accessibility audit cost", "which agency for an institutional website in North Africa".
- Ask them in a fresh session, with no history and, if possible, no account signed in. History steers the answers and skews the reading.
- Note who gets cited, with the exact link. Not "my sector comes up", but which specific sources appear.
- Do it again three days later. An isolated answer proves nothing; recurrence does.
- Sort the result into three cases: you are cited, a direct competitor is cited, or no source from your sector appears at all. The third case is the most interesting — the place is open.
What does not work
Repeating keywords produces nothing: the system does not count occurrences, it judges whether a passage answers. Hiding instructions in a page to try to influence a model — white text on white, disguised directives — is manipulation, it is detected, and it exposes you to lasting exclusion. And publishing thirty thin pages on the same subject dilutes rather than reinforces: one complete page kept up to date beats ten approximations.
Where to start
Check crawler access first, then rewrite the three pages carrying your most commercial questions, applying the four properties above. The rest comes afterwards. The audit tells you in a minute where you stand on those points, and the pillar page puts the whole thing back into its wider mechanism.
