AhrefsBot and SemrushBot: What They Are and How to Control Them
By Seth Black · Published
If you’ve looked at your server logs, you’ve seen AhrefsBot and SemrushBot. On a lot of small sites they send more requests than Googlebot. They’re not search engines in the usual sense, blocking them won’t affect your Google rankings, and whether you should block them depends on what you want competitors (and yourself) to be able to see.
What AhrefsBot and SemrushBot are
Both are crawlers run by SEO software companies. They crawl the web to build the link indexes and site data sold in their tools: backlink reports, top pages, keyword rankings and competitor research. When someone looks up your domain in Ahrefs or Semrush, much of what they see came from these bots.
They also run the site audit features of those tools, which crawl a specific site when a customer asks.
Ahrefs bots
Ahrefs documents two crawlers at ahrefs.com/robot:
| Bot | User agent contains | Used for |
|---|---|---|
| AhrefsBot | AhrefsBot/7.0; +http://ahrefs.com/robot/ |
Ahrefs’ link and web index, and the Yep search engine |
| AhrefsSiteAudit | AhrefsSiteAudit/6.1; +http://ahrefs.com/robot/site-audit |
The Site Audit tool, crawling sites on request |
Ahrefs says both bots respect Disallow and Allow rules and Crawl-delay. The one exception: a verified owner of a site can let AhrefsSiteAudit ignore robots.txt on their own site, to audit sections that are normally blocked. By default Site Audit crawls at most 30 URLs per minute.
To confirm a request really comes from Ahrefs, check the IP against the ranges Ahrefs publishes, or do a reverse DNS lookup: the hostname should end in ahrefs.com or ahrefs.net.
Semrush bots
Semrush lists its crawlers at semrush.com/bot. The ones you’re most likely to see:
| Bot | Used for |
|---|---|
| SemrushBot | Semrush’s main crawler for backlink analytics and other tools |
| SiteAuditBot | The Site Audit tool |
| SemrushBot-BA | Backlink Audit |
| SemrushBot-SI | On Page SEO Checker and similar tools |
| SemrushBot-SWA | SEO Writing Assistant, checking URL accessibility |
| SplitSignalBot | SEO A/B tests in SplitSignal |
| SemrushBot-OCOB | Content Toolkit reports |
| SemrushBot-FT | Plagiarism Checker |
Semrush says its bots follow robots.txt, and SemrushBot supports Crawl-delay with intervals of up to 10 seconds. Semrush also says not to block its bots by IP, because it doesn’t use consecutive IP ranges. Use robots.txt instead.
Should you block them?
Blocking them does not affect Google or Bing. Search engines don’t use Ahrefs or Semrush data to rank sites.
What blocking does change:
- Your site’s outbound links disappear from their indexes. If your site links to other sites, those links won’t appear in Ahrefs or Semrush backlink reports. That matters to the sites you link to, not to you.
- Competitors see less of your site. Top pages, content and site structure data about your domain become thinner in those tools. Your backlinks from other sites are still visible, because those links are crawled on the other sites.
- Your own audits may stop working. If you or your agency use Ahrefs Site Audit or Semrush Site Audit on your site, blocking the audit bots breaks them. Block the index crawlers and allow the audit bots if that’s the goal.
- Server load goes down. On small or slow hosting, this is the real reason most people block them.
My take: if the crawl rate isn’t hurting your server, leave them alone or slow them down. If it is, slow them first and block only if that doesn’t work.
How to slow them down
Add a crawl delay in robots.txt. Both honor it (Google does not, which is fine; Google ignores Crawl-delay entirely):
User-agent: AhrefsBot
Crawl-delay: 10
User-agent: SemrushBot
Crawl-delay: 10
Ahrefs also says its bot backs off automatically when your server returns 4xx or 5xx errors, for example during maintenance.
How to block them
Block the index crawlers but keep your own audits working:
User-agent: AhrefsBot
Disallow: /
User-agent: SemrushBot
Disallow: /
Block every Ahrefs and Semrush crawler, including site audits:
User-agent: AhrefsBot
User-agent: AhrefsSiteAudit
User-agent: SemrushBot
User-agent: SiteAuditBot
User-agent: SemrushBot-BA
User-agent: SemrushBot-SI
User-agent: SemrushBot-SWA
User-agent: SplitSignalBot
User-agent: SemrushBot-OCOB
User-agent: SemrushBot-FT
Disallow: /
A group with several User-agent lines applies its rules to each of them. Make sure these groups don’t accidentally match Googlebot: rules for a specific bot only apply to that bot, and Googlebot follows its own group or the * group.
Changes can take a while to take effect, because bots cache robots.txt. Test the result for each user agent with our free robots.txt tester.
Blocking at the firewall instead
robots.txt is a request, not an enforcement mechanism. Legitimate bots follow it; scrapers pretending to be AhrefsBot don’t. If you see “AhrefsBot” requests that ignore robots.txt, verify them before blaming Ahrefs:
# reverse DNS on an IP from your logs
host 203.0.113.10
A real Ahrefs IP resolves to a hostname ending in ahrefs.com or ahrefs.net. Anything else is an impostor, and a firewall or CDN bot rule is the right tool. Cloudflare lists both Ahrefs bots as verified bots, which makes rules like “block unverified bots claiming to be AhrefsBot” straightforward.
How to find them in your logs
On a standard combined log format:
# requests per bot
grep -Eio 'AhrefsBot|AhrefsSiteAudit|SemrushBot[-A-Z]*|SiteAuditBot|SplitSignalBot' access.log | sort | uniq -c | sort -rn
# the paths AhrefsBot requests most
grep AhrefsBot access.log | awk '{print $7}' | sort | uniq -c | sort -rn | head -20
If most of their requests go to parameter URLs, filters or calendar pages, that’s a crawl trap Googlebot is probably hitting too. Fix the trap and the bot traffic drops for everyone. How to audit an ecommerce site for crawl waste covers the usual culprits.
What about AI crawlers?
GPTBot, ClaudeBot, PerplexityBot and CCBot are a separate decision with different trade-offs. The AI crawler robots.txt guide covers them.
Auditing your site without a third-party bot
If you block audit bots but still want an audit, run the crawler yourself. bseoa crawls from your own machine, respects robots.txt by default, and has --ignore-robots for crawling sections of your own site that robots.txt disallows. Its user agent is configurable with --user-agent, and --rate-limit and --concurrent-requests control how hard it hits your server. It’s $279 once, with a 14-day free trial.
-Sethers