AI Traffic Control Compared 2026: How Cloudflare, Akamai and Vercel Classify Bots Differently

Blocking the wrong category of AI bot costs you different things. Block training crawlers and your catalog is less present in the models people ask for recommendations. Block the real-time fetchers and a shopper who clicks your link inside a chat assistant gets a page that will not load. Cloudflare covered both with a single switch until July 2026. In 2026 all three major options split it apart, and they split it the same way while calling the parts different names.

Three vendors, one split, three vocabularies

ProviderCategory namesDefault postureGranularityHard block possibleWhere the setting lives
CloudflareSearch / Agent / TrainingFrom September 15, 2026, domains newly onboarding get Training and Agent blocked on pages displaying ads, Search allowedThree independently managed categoriesYesBot management settings in the dashboard
AkamaiTraining crawlers / search crawlers / fetchersPolicy-driven, no single defaultCategory plus individual bot signaturesYesBot Manager policy
VercelTraining / search / user-generated fetchesRuleset must be enabled deliberatelyCategories within the managed rulesetYesAI bots managed ruleset in the firewall
robots.txtWhatever User-agent lines you writeNone, declarative onlyPer user agent and per pathNoSite root

The convergence on three buckets is not coincidence. Each bucket is a different business transaction wearing the same HTTP request. A training crawl pays out months later, when a model answers a question and your brand is in the answer. A search crawl pays out in days, when your page becomes a citation. A real-time fetch pays out immediately, because a person is waiting. The three can arrive from the same IP ranges and sometimes the same company, and still have nothing in common commercially.

Cloudflare made the change in July 2026, replacing the single “block AI bots” switch with three categories you manage independently. Agent covers automated activity acting in real time on a person’s behalf, which in practice means the fetch a chat assistant performs while a user waits for an answer.

The September 15 default applies to domains newly onboarding to Cloudflare, and only to pages that display ads: Training and Agent blocked, Search allowed. Existing domains keep whatever they already have, which is exactly why it is worth opening the settings and reading the current state of all three categories rather than assuming. A page-by-page audit of that change is covered separately and not repeated here.

Cloudflare’s own wording is that multi-purpose crawlers combining Search and Training are affected. Read that as a statement about mixed-identity crawlers, and do not attach it to any one search engine’s crawler by name.

The three categories are not worth the same to a commerce site, and the ranking does not change between vendors. Agent traffic is worth the most because a real shopper is waiting on the other end of it, comparing you against two other brands. Search is next, determining whether you surface in AI answers on a horizon of days to weeks. Training matters on a horizon of months, deciding whether a model recalls your brand unprompted. Protect Agent first, Search second, and take your time on Training.

Akamai: same three buckets, enterprise controls

Akamai sorts AI traffic into training crawlers, search crawlers, and fetchers. Line that up against the table and it maps cleanly onto Cloudflare’s Training, Search, and Agent, with fetchers playing the real-time role.

Akamai reported AI bot traffic on its own network rising more than 300% between 2025 and early 2026. Every rule you write now touches far more requests than the same rule would have a year ago, so a misconfiguration is proportionally more expensive.

That growth has a cost dimension worth measuring before you write any rule. Pull the cache hit ratio on AI bot requests specifically. A low ratio means those requests are landing on your origin, so both the bill and origin load scale with the traffic, and improving cacheability returns more than a block does while costing you nothing in visibility. A high ratio means the cost pressure is limited and you can go straight to the category policy.

Akamai’s security team has written publicly that blocking AI agents outright means opting out of how purchase decisions are increasingly being made. A commerce site pays for that directly. When a shopper asks an assistant which trekking pole is lightest and the assistant fetches product pages to answer, a blanket block removes you from the shortlist before the comparison starts.

Granularity is where Akamai leads. Beyond the three categories you can write policy against individual bot signatures, which is the only way to treat two crawlers in the same category differently. The cost of that control is the commercial model: Bot Manager is an enterprise product with enterprise contracting, and a small store will not get to it.

Signature-level control earns its keep when two bots in one category behave nothing alike. Both may be fetchers, but one arrives a few hundred times a day and leaves, while the other sweeps every product URL you have in a twenty-minute burst. The first belongs on the allow list, the second belongs behind a rate limit. A category toggle cannot tell them apart, so you either accept both or lose both.

Vercel: rules that ship with the deployment

Vercel’s AI bots managed ruleset controls traffic from AI bots crawling for training data, for search, or for user-generated fetches. Being a managed ruleset, Vercel maintains the bot list and you decide the action per category.

Vercel also ships BotID, which answers a different question. The ruleset decides whether a known AI crawler gets through. BotID decides whether an unidentified request is automated at all. Most teams running commerce on Vercel want both on, because the second one protects checkout and form endpoints from scripted abuse that has nothing to do with AI search.

The operational advantage is that these rules live with the project rather than in a separate ops console, so the front-end team that owns the deployment can change them. For a headless store where the front-end team is the ops team, that removes a handoff.

The limit is coverage. Only traffic served through Vercel is subject to the ruleset. If product pages run on Vercel while images, reviews, or a search API sit on another origin, that traffic needs its own controls.

Choosing the action per category needs a moment of thought too. Between allow and block there is usually a challenge option, and for AI crawlers a challenge is functionally a block, because they do not solve human verification. If your intent is to slow a crawler rather than exclude it, a challenge does not express that. Rate limiting does.

robots.txt is the free baseline and nothing more

Write robots.txt correctly regardless of which provider you use. It costs nothing, takes effect on the next crawl, and gives you per-user-agent and per-path precision no category toggle can match.

It also has no teeth. robots.txt is a declaration, and compliance is voluntary. Well-behaved crawlers read it and honor it. Crawlers that ignore it are exactly the ones you most wanted to stop. Use it to state intent, never as a control.

The working setup is two layers that agree with each other. robots.txt states the policy, the edge enforces it. When the two disagree, typically robots.txt allowing a crawler the edge is quietly blocking, you lose days later trying to work out why your brand stopped appearing in AI answers.

A commerce robots.txt also needs path rules that have nothing to do with AI. Faceted listing URLs, internal search result pages, cart and checkout paths carry no value for any crawler, and leaving them open spends crawl budget on tens of thousands of near-identical URLs. Product pages, category pages, and editorial content are what you want fetched, by search crawlers and AI fetchers alike.

Worth tracking as direction: the IETF is developing the Bot Service Index (BSI), a discovery infrastructure that lets agents cryptographically prove their identity. Every rule described above rests on user agent strings and IP ranges, both forgeable. BSI would make “this request genuinely came from that assistant” verifiable for the first time. Nothing to implement today, but it belongs in any traffic policy you expect to still be running a year from now.

Which setup fits which seller

Shopify store on Cloudflare’s free tier. Your controls are the three category toggles plus robots.txt. Allow Search. Allow Agent. Decide Training on your own view of model training, keeping in mind that letting your product copy into training data is how a smaller brand gets recalled by name later. The one setting to verify rather than assume is Agent: leaving that blocked drops live pre-purchase traffic, which is the most valuable AI traffic you receive.

Headless store on Vercel. Enable the AI bots managed ruleset rather than leaving it at its unconfigured state, then inventory your domains and handle anything not served through Vercel separately. Turn on BotID at the same time; it addresses scripted abuse of checkout and forms, a distinct problem with the same urgency.

Enterprise seller on Akamai. Use the granularity you are paying for. Set a category-level baseline, then open specific high-value AI entry points and tighten specific suspect signatures individually. Before configuring anything, pull your own AI bot traffic composition rather than working from an industry default, because the mix varies sharply between product categories.

A fourth configuration turns up often enough to name: a site deployed on Vercel with Cloudflare in front of it. Both rule sets apply and Cloudflare evaluates first, so debug from the outside in rather than hunting through Vercel rules for a request that never arrived. If both layers have AI rules configured, set one of them to allow everything and keep the decision in a single place. Two layers of partial policy is how a team ends up unable to say which rule produced the outcome.

All three should do the same follow-up. A week after any change, ask a few AI assistants about your brand terms and lead products and compare citations against what you saw before. A citation that was there last week and is gone this week points at the category you just changed, and no dashboard on any of these three platforms will report that for you.

Related Articles

Cloudflare Changed Its Defaults on September 15: The AI Agents Shopping for Your Customers May Be Locked Out

In July, Cloudflare split its single AI-bot switch into three separately managed categories: Search, Agent, and Training. From September 15, domains newly onboarding get a default that blocks Training and Agent bots on pages that display ads while Search stays allowed. Here is what each category covers, who the new default actually touches, and the dashboard audit that takes half an hour.