Stop Blocking AI Agents: Manage AI Crawlers Without Blanket Rules
Somebody on your team wants to block AI crawlers. It's an understandable instinct.
Some of that traffic is a training scrape you never agreed to. Some of it is a prospective customer asking ChatGPT, Claude, or Grok whether your product can do the job, and the assistant fetching your pricing page to find out. Block all of it and you've kept a scraper off your content... and your product out of the answer.
The block doesn't even work very well. The polite crawlers announce themselves and respect robots.txt, so they were never the problem. The rest don't announce themselves at all. One independent test found Grok's live traffic arriving as ordinary Chrome from datacenter proxies, never under the crawler names listed for it. A user-agent rule can't find the assistants, let alone the scripts that borrow their names.
A user agent tells you what a client claims to be. It doesn't tell you what the client is about to do. The policy that works lets fetchers fetch and saves its judgment for whatever signs up. That's the job of a judgment layer, whether it's Dregs or another one.
Why Blanket Rules Fail
You are right to want the registration farm gone. Scripts still invent names, rotate inboxes, and open accounts overnight. That is signup bot abuse, and it belongs off the form. A one-line rule that blocks every automated client does not stop that farm. It blocks the traffic you meant to keep.
The rule that stops "AI bots" tends to stop everything else that isn't a person: the Pingdom
check on /health, the preview bot that builds your card when someone shares a
link on LinkedIn, the research crawler reading your public API docs, and the assistant a
customer sent to read your pricing page. Those are false positives, and they land on your
customers and your search presence, not on the pest you meant to stop. A puzzle on the same
door makes the same mistake: it taxes the assistant, and a solver farm still walks through.
CAPTCHA alternatives covers what to use instead.
Meanwhile the more ambitious of these pests have noticed the panic. Copying an assistant's user agent is cheaper than a residential proxy, and it looks like a helpful client in the logs. A session that calls itself an assistant and then burns a trial is not an assistant. It is a signup bot using a borrowed name. Knowing that a request is automated is the easy part. Deciding whether it is a problem is the part that matters.
How to Identify AI Crawlers
"AI crawler" is not one act. The large operators split their traffic by job and publish a user-agent token for each, which is the most useful thing to know when you are reading logs.
| Job | Examples | What it is doing | Sensible default |
|---|---|---|---|
| Training crawlers | GPTBot, ClaudeBot, CCBot, Bytespider, Meta-ExternalAgent | Collecting pages to train models | Your call, made in robots.txt. Not an abuse signal either way. |
| AI search indexers | OAI-SearchBot, Claude-SearchBot, PerplexityBot | Indexing pages so an AI search product can cite them | Usually allow. This is how you show up in answers. |
| User-initiated fetchers | ChatGPT-User, Claude-User, Perplexity-User | Fetching a page because a person just asked a question | Allow. There is a customer on the other end. |
| Agent browsers | ChatGPT's Cloud browser and other browser-based agents | Driving a full browser to complete a task for a person | Allow the browsing. Judge the account if it signs up or signs in. |
Then there is the traffic that borrows any of those names: signup bots, credential stuffing, scrapers hoarding content or inventory, and bot farms manufacturing fake accounts. Stop these, preferably before they become "users" in your table, whatever header they are wearing. The wider taxonomy, including the good bots that have nothing to do with AI, is in good bots vs bad bots.
Treat the names as a starting point, not a credential. Copying a user agent costs nothing, and not every useful assistant announces itself. Crawler directories list Grok user agents such as GrokBot, but an independent test found Grok's live traffic arriving with generic Chrome, Safari, and Go HTTP client strings instead. An allowlist of names would never have matched it, and neither would a blocklist.
Training is the one row where blocking is a reasonable choice, and it is a content decision rather than a security one. You may want to allow retrieval and refuse training, or the reverse. Either way, it is not a reason to treat every automated client as suspicious activity.
Can AI Crawlers Execute JavaScript?
Mostly not, which surprises people. In Vercel's analysis of crawler traffic on its network, GPTBot, ChatGPT-User, ClaudeBot, PerplexityBot, and the Meta and ByteDance crawlers fetched HTML, and sometimes downloaded JavaScript files, without executing any of it. The exceptions were Gemini, which renders through Googlebot's infrastructure, and Applebot. Agent browsers are the other exception: ChatGPT's Cloud browser and its peers drive a full browser, so they run everything.
Two practical consequences follow. If your pricing or docs only appear after client-side rendering, most AI crawlers can't see them, and that is a content problem, not a bot problem. And JavaScript execution makes a poor on/off switch for policy. A rule that blocks clients that skip JavaScript hits the polite fetchers. A rule that blocks automated browsers hits agents acting for real customers. The well-built signup bot gets through both, because it renders the page the way any modern form expects.
There is a third consequence for account scoring. The Dregs tracker runs in the browser, and an identity only exists in Dregs once your application identifies a user at signup or sign-in. A crawler that fetches HTML and leaves never becomes an account, so there is nothing to judge and nothing to get wrong. Dregs doesn't police the fetch. It judges what signs up. When a client does run script and register, the evidence gets richer: a registration form filled in 400 milliseconds, followed by the next account from a sibling device, is not a fetcher. Dregs scores that from device rendering, automation signatures, and timing, with confidence attached, so a thin first hit is not treated like a completed campaign.
Allowlists Rot; Scoring Scales
Vendor allowlists and "AI bot" blocklists ask whether a request matches a name you already wrote down. New assistants appear faster than those lists are updated, names change, and the afternoon a list ships, a signup script copies the header.
Verification is getting better, and you should use it. Published IP ranges and reverse DNS have been around for years, and signed requests under Web Bot Auth let an agent prove where it came from cryptographically. ChatGPT's Cloud browser signs its requests, and edge providers such as Cloudflare can check the signature for you. But verification tells you who a client is. It doesn't tell you whether the account it just created is a problem.
| Control | Question it answers | Failure mode |
|---|---|---|
| Allowlist of user agents | Is this a crawler we already know? | Lists rot; new assistants have no entry; headers are spoofed |
| Blocklist of "AI crawlers" | Does this look like a published AI client? | Blocks assistants and AI search along with the farm; misses clients that never announce themselves |
| Verified identity (IP ranges, reverse DNS, Web Bot Auth) | Is this client who it says it is? | Proves who, not whether; many agents cannot be verified yet |
| Humanity + Behavior policy | Is this automation a problem for the product? | Needs identity history; not a substitute for volumetric edge defense |
Edge bot management still belongs in front of the app for noisy scrapers and for labeling known crawlers. It won't tell you that the "assistant" which just registered used a throwaway inbox, is the fourth account on the same hardware this week, and went straight for the free credits. That is an account question, and identity scoring is how you answer it.
A Humanity and Behavior Policy
Encode the split in custom rules and lists rather than in a global "block bots" toggle. The useful pattern is:
- Automation that never signs up. Fetchers, indexers, monitors, and link previews. Allow them at the edge, set your training preference in robots.txt, and move on. They never become accounts, so there is nothing to score.
- Low Humanity plus registrations, logins, or trial cycling. A claimed assistant that then signs up. Treat it as fake signups, free trial abuse, or credential stuffing, depending on the path. Stop it before it settles into the user table.
- Low Humanity on a paying customer's account, ordinary everything else. An agent doing a job for someone who pays you: pulling a report, filing a ticket, finishing a setup step. Expect more of this every quarter. Behavior, Authenticity, and Uniqueness are what tell you the account is fine, and they should outvote Humanity here.
- High Humanity, low Behavior. A person driving a routine: the trial cycler, the list tester. Humanity won't save you here. Behavior will.
- Automation you operate. QA bots, partner integrations, invited agents. Mark them disregarded so they stop affecting everyone else's analysis. Don't teach the rest of the product that "low Humanity" means "ban."
Dregs scores every account on Humanity, Behavior, Authenticity, and Uniqueness, updated moments after new activity, and each score opens into its observations so you can see exactly why before you act. Real users stay clear of false positives, including the ones whose assistant is doing the clicking, and there is less manual review to do.
Keep Humanity and Behavior as separate knobs. Badge likely bots, escalate only when several signals agree, and let webhooks refuse to provision, throttle, or simply watch. Dregs doesn't replace an edge WAF or a verified-bot catalog. It covers the part they can't see, which is what an account does after the request was allowed.
llms.txt and Crawler Norms
Crawler norms are useful and incomplete. robots.txt is how well-behaved
clients ask what they may fetch, and it is also where the training decision lives: GPTBot,
ClaudeBot, and CCBot are the crawler names to address, while Google-Extended and
Applebot-Extended are tokens that opt your content out of AI training without touching
search. Some sites also publish llms.txt at the root, a Markdown summary so
assistants can find a clean description of the product.
Both are courtesy. Neither authenticates the client. Neither notices that the same user agent later opened twenty accounts. Honor them at the edge for clients that honor them, and don't confuse a file on disk with bot management for AI crawlers. Abusive scripts ignore the file. Judgment still sits on Humanity, Behavior, and the identity in front of you.
Frequently Asked Questions
Q: What are AI crawlers?
A: AI crawlers are automated clients that fetch web pages for AI companies: to train models, to index content for AI search, or to retrieve a page because a person just asked an assistant a question. Examples include OpenAI's GPTBot, OAI-SearchBot, and ChatGPT-User, Anthropic's ClaudeBot and Claude-User, and PerplexityBot. Agent browsers such as ChatGPT's Cloud browser go further and drive a full browser on a person's behalf. Abusive scripts borrow the same names, so identify AI crawlers by what they do, not by the user agent alone.
Q: Should I block AI crawlers?
A: Not as a category. A blanket block also stops assistants fetching your pages to answer a customer's question, AI search indexers that decide whether you get cited, and agents acting for real users, and it usually catches uptime monitors and link preview bots on the way. Those are false positives: missing citations, red status checks, and customers whose assistant cannot read your docs. Make the training decision separately in robots.txt, let fetchers fetch, and judge the clients that register accounts, cycle trials, or scrape at volume.
Q: How do I identify AI crawlers?
A: Start with the published identity: user-agent tokens such as GPTBot, ChatGPT-User, ClaudeBot, Claude-User, and PerplexityBot, the IP ranges their operators publish, and signed requests (Web Bot Auth) where the agent supports them. Then watch behavior, because a user agent is a self-reported name. Some assistants do not announce themselves at all: an independent test in 2026 found Grok's live traffic arriving as ordinary Chrome from datacenter proxies rather than under the crawler names listed for it.
Q: Can AI crawlers execute JavaScript?
A: Mostly not. In Vercel's analysis of crawler traffic on its network, GPTBot, ChatGPT-User, ClaudeBot, PerplexityBot, and the Meta and ByteDance crawlers fetched HTML, and sometimes downloaded JavaScript files, without executing them. Gemini (through Googlebot's rendering) and Applebot do render JavaScript, and agent browsers such as ChatGPT's Cloud browser run a full browser. Content that only appears after client-side rendering is invisible to most AI crawlers, and a rule that blocks clients that skip JavaScript hits polite fetchers while missing the headless signup bots that render the page.
Q: What is llms.txt, and does it stop AI crawlers?
A: llms.txt is an optional file some sites publish at /llms.txt with a Markdown summary that assistants can read. It is a courtesy, in the same family as crawler norms such as robots.txt, not an enforcement control. Compliant crawlers may honor robots.txt; abusive scripts will not. Neither file authenticates the client or tells you whether an account that later appears in your product is a problem. Use them as secondary signals. Put judgment in Humanity and Behavior.
Q: How is bot management for AI crawlers different from an allowlist?
A: An allowlist answers whether a request matches a crawler you already wrote down. Verification (published IP ranges, reverse DNS, signed requests) makes that answer more trustworthy, but it still only tells you who the client is. Bot management for AI crawlers also asks whether the automation is a problem for the product. New assistants appear before lists catch up, and names are copied the afternoon a list ships. Behavior and identity scoring keep working as the roster changes, and every score opens into its observations so you can see exactly why before you act.
Further Reading
Stop abusive automation without blocking useful AI crawlers on your SaaS.
Dregs helps you judge whether automated traffic is actually a problem. It scores Humanity, Behavior, Authenticity, and Uniqueness on every account so you can allow useful assistants and stop the ones farming your product.
Start Free Trial