Sitemaps, robots.txt and noindex on Hosted Builders
Three controls that get confused with each other constantly, on platforms that expose them unevenly. What each one actually does, which builders let you touch which, and the combination that quietly stops working.
These three controls get treated as interchangeable ways of saying keep this page out of Google, and they are not. They operate at different stages, they answer different questions, and one common combination of them produces exactly the opposite of what the person intended. On a hosted builder, where you may only be given one of the three, understanding which one you have is the difference between a plan that works and a plan that fails quietly.
Here is the whole distinction in three sentences. A sitemap is a suggestion about what exists. robots.txt is an instruction about what may be fetched. noindex is an instruction about what may be listed. They are not substitutes, and no amount of one compensates for the absence of another.
Three controls, three different questions
Exists
Sitemap
A suggestion about what exists.
Fetched
robots.txt
An instruction about what may be fetched.
Listed
noindex
An instruction about what may be listed.
Sitemaps: a suggestion, not a promise
Every builder in this comparison generates a sitemap automatically and keeps it current, which is one of the genuine benefits of a hosted platform — it is a job that gets forgotten on hand-built sites for years at a time.
What a sitemap does is tell a crawler which addresses exist and when they last changed. What it does not do is guarantee that any of them get indexed. Submitting a sitemap is not a request for indexing and a page listed in one is not thereby promised a place in the index. A large gap between pages submitted and pages indexed in Search Console is normal on a new site and is not a sitemap fault; it is usually a site-age problem, sometimes a quality problem, and essentially never a sitemap problem.
The control you have over the sitemap on a hosted builder is close to nil, and that is fine. Duda documents the option of replacing its generated sitemap with your own file, which is unusual and useful at agency scale. Everybody else generates one and hands you the URL.
robots.txt: who is allowed to fetch what
robots.txt is a file at the root of a domain listing paths that crawlers should not request. It is a crawling instruction. It says nothing about indexing, which is the source of nearly every misunderstanding in this area.
Access varies more than on any other control here:
- Webflow — editable in project settings.
- Wix — editable through a built-in editor.
- Duda — auto-generated, with a documented option to substitute your own file.
- Shopify — editable through
robots.txt.liquid, available since June 2021, and shipped with an unusually stern warning: Shopify's own help centre describes it as an unsupported customisation, says Support cannot help with edits, and states that incorrect use can result in loss of all traffic. - Squarespace — not editable. One shared file across all sites. The only control is a checkbox that blocks crawlers site-wide.
- Hostinger's builder and GoDaddy's builder — auto-generated, not editable.
Most site owners should never touch this file even where they can. The legitimate uses are narrow: excluding an effectively infinite URL space, a staging path, or a parameter pattern generating thousands of near-identical addresses. If you cannot name which of those applies, the correct edit is none.
noindex: who is allowed to list a page
A noindex directive — usually a meta robots tag in the page head — tells a search engine it may fetch the page but must not list it in results. This is the right tool for the overwhelming majority of keep this out of Google requests: the thank-you page, the internal-use page, the campaign landing page that should not compete with the real one.
Builder support is uneven. Squarespace exposes a "hide page from search results" toggle in each page's SEO tab — with a specific and awkward gap: it is absent on the homepage and on individual collection items, which means blog posts, products and events, the bulk of what most Squarespace sites publish. Wix and Webflow expose per-page index controls. Shopify's is a theme-code edit rather than a setting. Hostinger's and GoDaddy's builders do not document one.
The combination that fails
Here is the mistake, and it is the single most useful thing in this article. Somebody wants a page out of the index, so they block it in robots.txt and add noindex to it.
Those two instructions are in conflict. If a crawler is forbidden from fetching the page, it never reads the page, so it never sees the noindex. If the address is known from elsewhere — a link, a sitemap, a mention — the page can end up listed anyway, with no description, because the engine has an address it is not allowed to look at and no instruction it is permitted to read.
Blocked and noindexed at the same time
The combination that fails
robots.txt forbids the fetch
The crawler never reads the page.
The noindex is never seen
It sits in a page nobody is allowed to open.
Listed anyway
Known from a link or a sitemap, the address can appear with no description.
Allow the crawl, refuse the listing
The crawl is allowed
The page can be fetched.
The noindex is read
The instruction reaches the search engine.
The page drops out
Then block it, if you still need to.
The rule is: to remove a page from results, allow it to be crawled and tell it not to be indexed. Blocking comes later, if at all, once the page has dropped out. On Squarespace this is the reason the toggle gap hurts — with no per-item noindex and no robots.txt access, there is no clean way to de-list a collection item at all.
The crawl-budget question, answered honestly
People reach for robots.txt because they have read about crawl budget. For a site under a few thousand URLs — which is every site any of these builders will ever host for a typical customer — crawl budget is not your constraint and managing it is not a productive use of an afternoon. Google has been consistent that it matters for very large or very frequently changing sites.
What does matter at small scale is internal linking, which every platform gives you free. A page three clicks from the homepage gets found and re-checked more readily than one reachable only from a sitemap. If a page matters, link to it from somewhere people go.
A four-step check for your own site
Open yourdomain.com/robots.txt and read it. It takes thirty seconds and it is the only way to know what your platform put there.
Open yourdomain.com/sitemap.xml and see whether the pages you care about are in it.
Pick one page you expect to be excluded and check its source for a robots meta tag, to confirm the toggle you flipped did something.
Then check Search Console's page report for pages excluded by an instruction you did not intend to give — the category that catches accidental noindex is the most useful single screen in the product.
Which of these three controls your platform actually hands you is part of the seven-builder control audit. The short version: only Webflow, Wix and Duda give you all three in a form a non-developer can safely use.
Questions people actually ask
Can I edit robots.txt on Squarespace?
Should I block a page in robots.txt or use noindex?
Why are my pages in the sitemap but not indexed?
Do I need to worry about crawl budget on a builder site?
How do I stop a blog post showing in search on Squarespace?
Where this fits
Every note on this site rolls up into one inventory: the best website builder for SEO, which scores seven platforms on how much of the technical surface they hand over.