Most business owners never open their own robots.txt file. There is no reason to, day to day. But this one small text file, sitting quietly at yoursite.com/robots.txt, can stop Google from showing your entire website in search results, and most people who have this problem have no idea it exists until enquiries dry up and nobody can explain why. It is one of the most common issues we find during a first technical audit, and one of the fastest to fix once found.
This guide explains what robots.txt and XML sitemaps actually do, in plain English, how to check your own site in five minutes, and the specific mistakes that quietly keep Melbourne business websites out of Google.
What robots.txt actually does
Robots.txt is a small text file that tells search engines which parts of your site they are allowed to look at. Think of it as a sign on the front gate, not a lock on the door. It is a request, not a security measure. A page blocked in robots.txt can still appear in Google if enough other sites link to it, so it is not a way to hide sensitive information.
Most sites only need a few lines in this file. Problems start when a developer, an old plugin, or a staging site setup accidentally blocks the whole website, which happens more often than most business owners would expect. A single misplaced line can tell Google to stay away from every page you have.
What an XML sitemap actually does
If robots.txt is the sign on the gate, the XML sitemap is the map of everything worth seeing once you are inside. It is a file, usually at yoursite.com/sitemap.xml, that lists every page you want search engines to know about and roughly how important each one is.
A sitemap does not force Google to include your pages. It is a strong hint, not a guarantee. But without one, Google has to discover your pages purely by following links, which takes longer and can miss pages that are not well linked internally.
| robots.txt | XML sitemap | |
|---|---|---|
| What it does | Tells search engines where they can and cannot go | Lists the pages you want found and indexed |
| File location | yoursite.com/robots.txt | yoursite.com/sitemap.xml |
| Effect if missing | Search engines assume everything is allowed | Google finds pages more slowly, through links only |
| Effect if wrong | Can block your entire site from Google | Can list pages that do not exist, or omit new ones |
Not sure whether your site has either of these set up correctly? Our free 30-minute audit will tell you, even if you never become a client. Book a free consultation.
How to check your own site in five minutes
You do not need any special software to do a basic check. Type your website address followed by /robots.txt into a browser and press enter. If you see a line that reads Disallow: / with nothing else qualifying it, your entire site is likely blocked from search engines. If the file does not load at all, that is usually fine, most search engines assume everything is allowed when no file exists.
Next, type your website address followed by /sitemap.xml. You should see a list of page addresses, often with dates showing when each was last updated. If the dates are all more than a year old, or the list is missing pages you know exist on the site, the sitemap has likely gone stale.
If your site runs on WordPress, plugins such as Yoast or Rank Math generate both files automatically, which is convenient but not foolproof. Automatic generation still inherits any blocking rules set elsewhere in the plugin, so a page marked noindex in the plugin settings can still show up in the sitemap by mistake, or vice versa. Automation removes the manual work, not the need to check the result.
Neither check tells you everything a proper audit would, but both take under a minute and can catch the two most damaging mistakes on this list.
How the two files work together
These two files do different jobs but they need to agree with each other. Robots.txt should list the location of your sitemap so search engines find it immediately, rather than stumbling across it eventually. And critically, a page listed in your sitemap should never be a page that robots.txt blocks. That contradiction, telling Google to both index a page and stay away from it, sends a confused signal that can leave the page in limbo, neither properly indexed nor properly excluded.
Googlebot, the program Google uses to read websites, is a heavy user of the internet in its own right. Googlebot alone accounted for 4.5% of all HTML requests across Cloudflare’s global network in 2025, more than every AI crawler combined, according to Cloudflare’s own year in review. That scale is exactly why these two small files matter: with that much automated traffic reading the web, a site that makes it easy for Googlebot to find and understand its pages has a real advantage over one that leaves it guessing.
We cover this alongside the rest of your site’s technical health in our practical guide to technical SEO for Melbourne businesses, which is worth reading if this is the first time you have looked under the bonnet of your own website.
Common mistakes Melbourne businesses make
The most common and most damaging mistake is a leftover line from a staging or development site: Disallow: /. This single line tells every search engine to stay away from the entire website. It is meant to stop Google from indexing a test version of a site before launch, and it gets left in place after the real site goes live far more often than developers would like to admit. We have seen this exact mistake cost a Melbourne business months of invisibility on Google after what should have been a routine website redesign.
The second common mistake is a sitemap that never gets updated. A business adds a dozen new pages over a year and the sitemap still lists the same twelve pages from when the site launched. New pages take longer to be found and can be missed altogether if nothing links to them clearly.
The third is blocking useful pages by accident, such as blocking an entire folder to hide old, unfinished content, without checking whether any live, valuable pages sit inside that same folder. This is especially common on sites that have been through more than one redesign, where old folder structures are blocked out of caution and never revisited.
If a technical audit sounds like exactly what your site needs right now, talk to our Melbourne team. Fixing a broken robots.txt file is usually a same-day job. Talk to our Melbourne team.
What this costs, how long it takes, and what goes wrong
Checking and fixing robots.txt and sitemap issues is one of the cheaper technical SEO jobs there is, because it is configuration work rather than content work. A straightforward fix can often be completed within a day once identified, though the audit to find every issue on a larger site can take a week or two.
The larger cost is what happens while the problem goes unnoticed. Australians spent a record $82.6 billion online in 2025, up 14% year on year, with 24% of all retail spend now happening online, according to the Australia Post eCommerce Report 2026. A Melbourne business with a blocked robots.txt file is invisible for the entire slice of that spending that starts with a Google search, for as long as the mistake goes uncaught.
What typically goes wrong is not the technical fix itself, it is finding the problem in the first place. Most business owners only discover a robots.txt issue when someone mentions their site does not come up on Google, months after the mistake was made, often after a website migration or a change of web developer where nobody checked what changed under the surface.
Frequently asked questions
Do I need both a robots.txt file and a sitemap? Yes. They serve different purposes and most SEO plugins or website platforms will create basic versions of both automatically, but they still need to be checked, not assumed to be correct. See our full FAQ for more.
How do I check if my site has a robots.txt problem? Google’s own Crawl Stats report in Search Console shows exactly how Googlebot is treating your site, including blocked requests, alongside the quick manual check described above.
Should I submit my sitemap to Google directly? Yes, submitting it through Google Search Console speeds up discovery, though it is not a substitute for having the file itself correctly built and linked from robots.txt.
What happens if I fix a blocked robots.txt file today? Google does not re-crawl a site instantly, so expect pages to start reappearing in search results over the following one to three weeks, not overnight, and full recovery can take a little longer for a large site.
A website that reads perfectly to a human but blocks search engines is invisible for the searches that matter most. Book your free consultation. No lock-in contracts, reply within 24 hours, Melbourne-based team. Get started.
For more guides like this one, visit our SEO blog.
Sources: Google Search Central guidance on robots.txt and building a sitemap; Google Search Console Help, Crawl Stats report; Cloudflare Radar 2025 Year in Review; Australia Post eCommerce Report 2026.
Leave a Reply