Hacked robots.txt Files and What They Hide

By · Updated · 4 min read

It is four lines of plain text and almost nobody reads it after the day it was written. That is exactly why it is worth reading now. A modified robots.txt does not break anything visibly, does not trigger a scanner, and can quietly reshape what search engines do with your site — sometimes hiding a spam campaign, sometimes destroying your rankings outright.

What the file actually controls

robots.txt sits in your web root and tells well-behaved crawlers which parts of your site they may request. It can allow or disallow paths, name your sitemap locations, and set rules per crawler. It is a request rather than an enforcement mechanism — it does not prevent access, it asks politely — but the major search engines honour it.

That makes it powerful in one specific way: it shapes what gets crawled and therefore what gets discovered. Change it and you change what search engines look at, without changing a single page of your site.

How it gets abused after a hack

The most common change is an added sitemap line pointing at a spam sitemap, which is a direct submission channel for thousands of injected URLs. My guide on sitemap spam covers where those lead.

The second is selective hiding. Disallow rules can be added to stop crawlers requesting the paths you would use to check on things, or to keep your own security tooling away from a spam directory, while leaving the spam pages themselves perfectly crawlable. And occasionally the change is simple sabotage — a blanket disallow that tells every crawler to stay away from the entire site, which will remove you from search results over the following weeks.

Reading yours properly

Fetch it directly in a browser at your domain followed by /robots.txt and read every line. That is what crawlers see. Do not check the file on disk alone, because WordPress can generate a virtual robots.txt dynamically when no physical file exists, and a plugin or injected code can alter what that generated version contains.

Account for every rule. A normal WordPress robots.txt is short: a disallow for wp-admin with an allow for admin-ajax, perhaps a few paths you deliberately excluded, and your sitemap. Anything else needs an explanation — particularly disallow rules for paths you do not recognize, allow rules for directories you did not create, and sitemap lines pointing anywhere other than your real sitemap.

The blanket disallow problem

If you find a rule disallowing everything for all crawlers, fix it immediately, because every day it stays live is a day of deindexing. The damage here is not subtle and it compounds — search engines drop pages they have been told not to crawl, and getting them back takes considerably longer than losing them did.

Once corrected, use Search Console to confirm the file is being read as you expect, and request recrawling of your important pages to speed recovery. My guide on ranking recovery after a hack covers realistic timelines for coming back from this.

Fixing it and looking further

Restore a correct robots.txt, remove any spam sitemap declarations, and confirm the served version matches what you intended. Then keep going, because a modified robots.txt is evidence that someone could write to your web root, and that is a much bigger finding than the file itself.

Check for added files in the root, look for backdoors and web shells, check your .htaccess files, and audit administrator accounts. Then do the standard cleanup and rotate credentials.

Keeping an eye on it

This file should essentially never change on its own. That makes it an excellent tripwire: if you have file integrity monitoring, make sure it is covered, and if you do not, a quarterly look at it costs nothing. Search Console will also flag crawl anomalies that often trace back to it.

It is a small file that punches well above its weight, and it is one of the first things I read on any site I am assessing — it frequently tells me more about what happened than the pages do. If you want a full sweep rather than a file-by-file check, that is my malware removal service.

Common questions

Where is my robots.txt file?

At your domain followed by /robots.txt. It may exist as a physical file in your web root, or WordPress may generate it dynamically when no file is present. Always check the served version in a browser rather than only looking on disk, since the generated version can be altered by a plugin or by injected code.

Can a modified robots.txt hide a hack from me?

Indirectly, yes. It can stop crawlers and some scanning tools requesting particular paths, which keeps a spam directory out of the reports you might be reading while leaving it perfectly accessible to anyone with the URL. It hides the evidence rather than the infection.

What should a normal WordPress robots.txt contain?

Very little. Typically a disallow for /wp-admin/ with an allow for admin-ajax.php, any paths you deliberately excluded, and a line pointing at your sitemap. If yours contains rules you cannot explain, especially sitemap declarations for files you did not create, something has changed it.

My robots.txt blocks everything. How much damage is done?

It depends entirely on how long it has been live. A few days is usually recoverable with little lasting harm. Several weeks means significant deindexing, and recovery takes longer than the damage did. Fix it immediately, then request recrawling of your key pages in Search Console.

Is robots.txt a security control?

No, and it is important not to treat it as one. It is a request that well-behaved crawlers honour. It does not restrict access, and listing a sensitive path in it actually advertises that path to anyone who reads the file. Use real access controls for anything that genuinely needs protecting.