Skip to the audit
Controls re-read 19 Sept 2026
Menu

Six controls, seven builders, one re-check date on every page.

Field note

Sitemaps, robots.txt and noindex on Hosted Builders

Three controls that get confused with each other constantly, on platforms that expose them unevenly. What each one actually does, which builders let you touch which, and the combination that quietly stops working.

Published 21 August 2026Re-checked 19 September 2026

These three controls get treated as interchangeable ways of saying keep this page out of Google, and they are not. They operate at different stages, they answer different questions, and one common combination of them produces exactly the opposite of what the person intended. On a hosted builder, where you may only be given one of the three, understanding which one you have is the difference between a plan that works and a plan that fails quietly.

Here is the whole distinction in three sentences. A sitemap is a suggestion about what exists. robots.txt is an instruction about what may be fetched. noindex is an instruction about what may be listed. They are not substitutes, and no amount of one compensates for the absence of another.

Three controls, three different questions

Exists

Sitemap

A suggestion about what exists.

Fetched

robots.txt

An instruction about what may be fetched.

Listed

noindex

An instruction about what may be listed.

They are not substitutes, and no amount of one compensates for the absence of another.

Sitemaps: a suggestion, not a promise

Every builder in this comparison generates a sitemap automatically and keeps it current, which is one of the genuine benefits of a hosted platform — it is a job that gets forgotten on hand-built sites for years at a time.

What a sitemap does is tell a crawler which addresses exist and when they last changed. What it does not do is guarantee that any of them get indexed. Submitting a sitemap is not a request for indexing and a page listed in one is not thereby promised a place in the index. A large gap between pages submitted and pages indexed in Search Console is normal on a new site and is not a sitemap fault; it is usually a site-age problem, sometimes a quality problem, and essentially never a sitemap problem.

The control you have over the sitemap on a hosted builder is close to nil, and that is fine. Duda documents the option of replacing its generated sitemap with your own file, which is unusual and useful at agency scale. Everybody else generates one and hands you the URL.

robots.txt: who is allowed to fetch what

robots.txt is a file at the root of a domain listing paths that crawlers should not request. It is a crawling instruction. It says nothing about indexing, which is the source of nearly every misunderstanding in this area.

Access varies more than on any other control here:

  • Webflow — editable in project settings.
  • Wix — editable through a built-in editor.
  • Duda — auto-generated, with a documented option to substitute your own file.
  • Shopify — editable through robots.txt.liquid, available since June 2021, and shipped with an unusually stern warning: Shopify's own help centre describes it as an unsupported customisation, says Support cannot help with edits, and states that incorrect use can result in loss of all traffic.
  • Squarespace — not editable. One shared file across all sites. The only control is a checkbox that blocks crawlers site-wide.
  • Hostinger's builder and GoDaddy's builder — auto-generated, not editable.

Most site owners should never touch this file even where they can. The legitimate uses are narrow: excluding an effectively infinite URL space, a staging path, or a parameter pattern generating thousands of near-identical addresses. If you cannot name which of those applies, the correct edit is none.

noindex: who is allowed to list a page

A noindex directive — usually a meta robots tag in the page head — tells a search engine it may fetch the page but must not list it in results. This is the right tool for the overwhelming majority of keep this out of Google requests: the thank-you page, the internal-use page, the campaign landing page that should not compete with the real one.

Builder support is uneven. Squarespace exposes a "hide page from search results" toggle in each page's SEO tab — with a specific and awkward gap: it is absent on the homepage and on individual collection items, which means blog posts, products and events, the bulk of what most Squarespace sites publish. Wix and Webflow expose per-page index controls. Shopify's is a theme-code edit rather than a setting. Hostinger's and GoDaddy's builders do not document one.

The combination that fails

Here is the mistake, and it is the single most useful thing in this article. Somebody wants a page out of the index, so they block it in robots.txt and add noindex to it.

Those two instructions are in conflict. If a crawler is forbidden from fetching the page, it never reads the page, so it never sees the noindex. If the address is known from elsewhere — a link, a sitemap, a mention — the page can end up listed anyway, with no description, because the engine has an address it is not allowed to look at and no instruction it is permitted to read.

Blocked and noindexed at the same time

The combination that fails

robots.txt forbids the fetch

The crawler never reads the page.

The noindex is never seen

It sits in a page nobody is allowed to open.

Listed anyway

Known from a link or a sitemap, the address can appear with no description.

Allow the crawl, refuse the listing

The crawl is allowed

The page can be fetched.

The noindex is read

The instruction reaches the search engine.

The page drops out

Then block it, if you still need to.

To remove a page from results, allow it to be crawled and tell it not to be indexed. Blocking comes later, if at all.

The rule is: to remove a page from results, allow it to be crawled and tell it not to be indexed. Blocking comes later, if at all, once the page has dropped out. On Squarespace this is the reason the toggle gap hurts — with no per-item noindex and no robots.txt access, there is no clean way to de-list a collection item at all.

The crawl-budget question, answered honestly

People reach for robots.txt because they have read about crawl budget. For a site under a few thousand URLs — which is every site any of these builders will ever host for a typical customer — crawl budget is not your constraint and managing it is not a productive use of an afternoon. Google has been consistent that it matters for very large or very frequently changing sites.

What does matter at small scale is internal linking, which every platform gives you free. A page three clicks from the homepage gets found and re-checked more readily than one reachable only from a sitemap. If a page matters, link to it from somewhere people go.

A four-step check for your own site

Open yourdomain.com/robots.txt and read it. It takes thirty seconds and it is the only way to know what your platform put there.

Open yourdomain.com/sitemap.xml and see whether the pages you care about are in it.

Pick one page you expect to be excluded and check its source for a robots meta tag, to confirm the toggle you flipped did something.

Then check Search Console's page report for pages excluded by an instruction you did not intend to give — the category that catches accidental noindex is the most useful single screen in the product.

Which of these three controls your platform actually hands you is part of the seven-builder control audit. The short version: only Webflow, Wix and Duda give you all three in a form a non-developer can safely use.

Questions people actually ask

Can I edit robots.txt on Squarespace?
No. Every Squarespace site is served the same robots.txt and customers cannot access or edit it. The only related control is a checkbox that blocks search engine crawlers across the whole site, which is appropriate for a site under construction and nothing else.
Should I block a page in robots.txt or use noindex?
Use noindex. Blocking in robots.txt stops the page being fetched, which means the noindex is never read, and a page known from a link elsewhere can end up listed anyway with no description. Allow the crawl, refuse the listing.
Why are my pages in the sitemap but not indexed?
That is normal and it is not a sitemap fault. A sitemap tells a crawler an address exists; it is not a request for indexing and carries no promise of it. New sites routinely wait weeks, and a persistent gap on an established site is a quality or authority signal rather than a technical one.
Do I need to worry about crawl budget on a builder site?
Almost certainly not. Crawl budget is a concern for very large or very frequently changing sites, not for the few hundred URLs a typical builder site holds. Internal linking is the lever that actually matters at that scale, and every platform gives it to you free.
How do I stop a blog post showing in search on Squarespace?
There is no clean way. The per-page 'hide from search results' toggle is absent on individual collection items, which includes blog posts, and robots.txt is not editable either. The practical options are to unpublish the post or to accept that it stays listed.

Where this fits

Every note on this site rolls up into one inventory: the best website builder for SEO, which scores seven platforms on how much of the technical surface they hand over.