AI Search and Builder Sites: What Is Known and What Is Guessed
Assistants now answer questions your pages used to answer. Almost everything written about influencing them is vendor research or speculation — so here is the short list of things that are actually established, and what a builder site can do about them.
A new category of advice has appeared in the last two years telling site owners how to get cited by AI assistants, and it has a structural problem: the people publishing it cannot see inside the systems they are describing. The search engines do not document citation selection. The assistant vendors do not either. What exists instead is correlation studies run by companies selling optimisation services, which is not worthless and is not the same as knowing.
So this note separates the two piles. The short pile of things that are established. The larger pile that is inference. And then the narrower question this site is actually about: which of it a hosted website builder gives you any means to act on.
Established: the markup you were told to add does not do this
Structured data is the first thing people reach for, on the theory that machine-readable markup must be what machines read. Two facts complicate that.
Google withdrew HowTo rich results in 2023 and restricted FAQ rich results to well-known authoritative government and health sites. Whatever schema is doing now, it is not producing those features for an ordinary site.
And the crawlers behind most assistant products are not documented as consuming JSON-LD in the way a search engine does. They fetch pages. The visible text is what they have to work with.
The practical conclusion is unglamorous: if information matters, put it in the text of the page where a human can read it. Markup it as well, by all means — it is cheap and other consumers use it. But information that exists only in JSON-LD is information you have hidden from the readers that matter here. That holds on every platform in the seven-builder control audit, including the ones with the best schema panels.
Established: blocking the crawlers has a cost
Assistant crawlers identify themselves and can be excluded in robots.txt, and there is a genuine debate about whether publishers should. For a site whose business model is people finding it, the arithmetic is simple enough: a page a system cannot fetch is a page it cannot cite, and the citation is the only distribution on offer from that system.
On most hosted builders the question is moot, because you do not control robots.txt. Squarespace, Hostinger's builder and GoDaddy's builder all generate it for you. Where you do control it — Webflow, Wix, Duda, Shopify — the default is to allow, and this desk's view is that leaving it alone is the correct call for a site that wants to be found. If you have a reason to exclude, exclude deliberately; do not do it because a checklist said to.
Established: at least one builder has started shipping llms.txt
A convention has emerged of publishing an llms.txt file — a plain markdown index of a site's significant pages, intended to make a site's structure legible to language models. It is not a standard anybody has ratified, and no assistant vendor has committed to reading it.
Hostinger documents its builder generating one automatically alongside metadata, sitemap.xml and robots.txt once you publish on a custom domain. Whether it does anything is unknown. It costs nothing, it is the only platform-level move in this whole area any builder in this comparison has made, and it is worth knowing about if only because it is the clearest signal that the platforms consider this a real channel.
Guessed: the citation mechanics everybody quotes
Several figures circulate constantly. That the overwhelming majority of AI Overview citations come from pages already ranking in the organic top ten. That ChatGPT's citations overlap heavily with Bing's results. That Perplexity weights recency and pulls a large share of its citations from a handful of community sites.
These come from vendor research — companies selling SEO tooling, analysing samples of answers. The samples are real and the analyses are often competent. They are also unaudited, unreplicated, measured at one point in time on systems that change monthly, and produced by parties with an interest in the conclusion. Treat them as weather reports rather than physics.
The one inference that seems robust across all of them, and which also happens to be the least actionable-sounding, is that classic search visibility gates AI visibility. If that holds, there is no shortcut: the work that gets a page cited by an assistant is the work that gets it ranked, and this whole field is a new distribution channel for an old discipline rather than a new discipline.
The rule that is not optional
Google made manipulating generative AI responses in Search an explicit spam policy in 2026. Writing text on your page addressed to an AI system — instructions, hidden prompts, phrasing designed to make a model recommend you — is not a clever tactic, it is a named violation.
It is also, separately, terrible writing. A paragraph constructed to be quoted by a machine reads exactly like a paragraph constructed to be quoted by a machine, and the humans who arrive from the citation notice.
What a builder site can actually do
Everything on this list is available on every platform here, which is the useful finding: this is not a capability question.
Answer the question in the first two sentences under each heading. Not as a trick for extraction, but because it is better writing and it makes a page quotable. Elaboration after the answer, not before it.
Attach dates, numbers and named sources to claims. A sentence containing a verifiable specific is a sentence something can be built on. A sentence of general advice is not, for a reader or for a machine.
Keep the visible text complete. Do not put the substance in an image, a video, or a collapsed element that requires interaction to reveal. Builders make all three easy.
Publish a real about page. Who runs the site, what the method is, how corrections are handled. Every framework for assessing content quality that anybody has published — including Google's own — comes back to this, and it is the one thing on the list a competitor cannot copy.
The honest summary
Nobody outside these companies knows how citations are selected, and anybody who tells you they do is selling something. What can be said is that the systems read visible text, that they appear to draw from pages already performing in conventional search, and that the tactics people are being sold for influencing them range from harmless to explicitly against policy.
For a builder site, that reduces to a familiar instruction: write pages that answer something specific, make them legible, do not hide the substance, and do not address the machines. That is the same advice as before the assistants existed, which is either reassuring or disappointing depending on what you were hoping for. The capability differences that do matter are the boring documented ones — the redirect managers, the canonical controls, the robots access — and those are checkable today.
Revisit this in six months
Everything above has a short shelf life, and that is a statement about the subject rather than about this page. The products change monthly, the research is re-run on new samples, and conventions like llms.txt either get adopted by somebody who matters or quietly stop being mentioned.
What will not change quickly is the underlying shape: visible text is what gets read, conventional visibility appears to gate the AI kind, and instructing machines is against policy. Build on those, treat the rest as provisional, and check the date on anything you read about this — including this.
Where this fits
Every note on this site rolls up into one inventory: the best website builder for SEO, which scores seven platforms on how much of the technical surface they hand over.