Anthropic built its entire brand on being the responsible one. Founded by researchers who left OpenAI specifically to build a safer alternative, the company has spent years positioning itself as the AI lab that takes ethics seriously. So it's a little jarring to look at Cloudflare's traffic data and see Anthropic sitting at the very bottom of the pack when it comes to giving back to the web it depends on.
Cloudflare sits in front of roughly a fifth of all internet traffic, which makes it one of the best vantage points anyone has for measuring how AI companies actually behave online. Its numbers tell a consistent story: Anthropic's bots crawl the web harder, relative to the traffic they send back, than almost anyone else building a major AI product.
What the Crawl-to-Refer Ratio Actually Measures
Cloudflare tracks something it calls the crawl-to-refer ratio. It's a simple idea. For every page a company's bots crawl, how many actual human visitors does that company send back to the site the content came from?
A healthy ratio looks like traditional search. Google crawls a page, indexes it, and when someone searches for that topic, Google sends a real visitor to the site. That's roughly a 5:1 exchange right now, according to Cloudflare Radar. You let the crawler in, and it pays you back in clicks.
An AI answer engine works differently. It reads a page, absorbs what it needs, and then answers the user's question directly inside the chat window. The user gets what they came for and never clicks through. The site that supplied the answer gets a server bill and nothing else.
Anthropic's Numbers, and Why They Keep Moving
For the week of July 1 to 7, 2026, Cloudflare's data puts Anthropic's crawl-to-refer ratio at roughly 2,800:1. That means for every visitor Claude sends back to a publisher, its bots, ClaudeBot for training and Claude-SearchBot for live search, have already crawled around 2,800 pages.
That 2,800:1 figure looks almost reassuring next to where Anthropic has been recently. In early April 2026, Cloudflare pegged the ratio at about 8,800:1. By the first week of May, it had spiked to roughly 24,700:1. Other trackers using slightly different methodology and time windows have put Anthropic's ratio anywhere from around 2,900:1 to over 11,000:1 in various weeks this spring and summer. The number swings wildly depending on the exact week you measure, largely because Anthropic's training crawler and its newer search crawler behave very differently, and the mix between them shifts constantly.
A year ago, in January 2025, Anthropic's ratio reportedly sat above 280,000:1. It has come down by more than an order of magnitude since then, mostly because Claude's web search feature now adds clickable citations that occasionally send someone back to the source. But even on its best recent weeks, Anthropic remains the most extractive of the major AI labs by a wide margin.
Here's the full picture, pulled directly from Cloudflare's AI Insights dashboard, for both a single week and a full month:
Anthropic isn't just worse than the field, it's in a different category entirely. Its closest competitor on this list, OpenAI, still refers traffic roughly eight to ten times more often than Anthropic does, and OpenAI itself sits well above everyone else beneath it. Perplexity, despite building its entire business around live web retrieval, refers at a rate over ten times better than Anthropic in both windows.
Microsoft, Yandex, and Baidu all sit in a fairly stable middle band, month to month, in the 10:1 to 35:1 range. These are companies running traditional search products (Bing, Yandex Search, Baidu Search) alongside AI features, and their ratios barely moved between June and July. That stability is itself informative: it suggests these ratios are not inherently volatile by nature. They move when the underlying business model is AI-answer-first rather than search-first, which is exactly what's driving Anthropic's swings.
Google and DuckDuckGo anchor the bottom of the table, both holding close to the traditional search bargain: crawl a page, and reliably send a visitor back for it.
Why the "Ethical" Company Crawls the Hardest
Anthropic rolled out real-time web search for Claude to compete with ChatGPT and Perplexity. Building that feature without leaning entirely on a licensed index like Bing's means Claude's bots have to fetch fresh pages constantly to answer questions accurately. That drives crawl volume up sharply. At the same time, more than half of Anthropic's crawling, and AI crawling industry-wide, is still dedicated to training rather than live search. Training crawls never send anyone anywhere. They just feed the model.
Put those two things together and you get exactly the pattern Cloudflare is measuring: heavy, constant crawling with a comparatively small trickle of referral traffic flowing back out.
Two Very Different Definitions of "Ethical"
When Anthropic describes itself as a safety-focused, ethical company, it means something specific by that. In AI alignment circles, ethics is about the model's behavior: not giving out bioweapon instructions, not generating hateful content, not confidently making up dangerous medical advice, and building in safeguards against a future system going rogue.
But that framework has almost nothing to say about how the model got built in the first place. To a publisher, a freelance writer, or a small site owner, ethics means something closer to consent, credit, and some form of compensation. Those two definitions of "ethical AI" don't just fail to overlap much. In Anthropic's case, they can point in opposite directions at once. A company can build genuinely careful safety systems into its model while running one of the most extractive crawlers on the internet, and both of those things can be true simultaneously.
The Old Deal Is Broken, and Everyone Knows It
For three decades, the web ran on an implicit trade. You let search engines crawl your content, and in exchange, they sent you readers you could monetize through ads, subscriptions, or sales. It wasn't perfect, but it worked well enough to fund an entire industry of independent publishing.
Cloudflare's own research has found that over half of all AI crawler requests are purely for training, meaning the content gets absorbed permanently with no expectation of ever sending a visitor back. That's a fundamentally different relationship than the one search engines built with the web. The publisher still pays the hosting costs for every one of those crawl requests. They just don't get anything in return.
Anthropic's usual defenses here are the same ones the rest of the industry leans on: training a model is comparable to a person reading widely and building expertise, the company respects robots.txt and doesn't hide its bots, and high-quality training data is genuinely necessary to keep the model from being useless or dangerous. Those arguments have some legal and technical merit, and courts are still actively working through how much of it holds up. But none of them really answer the publisher's actual complaint, which isn't about legality. It's about whether taking someone's work at scale, without payment or meaningful attribution, and then never sending them a reader again, counts as fair dealing.
The Double Standard Gets Sharper With the Distillation Fight
Anthropic has spent much of 2026 publicly accusing Chinese AI labs of stealing its work. In February, the company said it had caught DeepSeek, Moonshot AI, and MiniMax running coordinated campaigns, collectively generating more than 16 million exchanges with Claude through roughly 24,000 fake accounts, in an apparent effort to distill Claude's capabilities into cheaper competing models. In June, Anthropic escalated further, telling the Senate Banking Committee that operators linked to Alibaba's Qwen lab had run nearly 28.8 million exchanges through about 25,000 fraudulent accounts over a six-week period, using proxy services to get around geographic restrictions. Anthropic called it the largest distillation campaign it had ever detected, larger than the three prior campaigns combined, and framed it not just as IP theft but as a national security issue, warning that it lets foreign competitors strip away the safety guardrails Anthropic spent heavily to build.
These are Anthropic's own figures and Anthropic's own attribution. Alibaba has not addressed the specifics publicly, and outside parties have not independently verified the numbers or confirmed the accounts are actually linked to Alibaba rather than some other operator. Distillation is genuinely hard to prove from outside a company's own logs, since the evidence is a pattern of API usage rather than anything more concrete.
Even granting Anthropic's version of events at face value, the logic underneath it is nearly identical to the complaint publishers have about Anthropic itself. Both are, in essence, the same sentence: you used my expensive, hard-won output to build something that competes with me. A publisher spends years building an audience and a body of work, and ClaudeBot crawls it at scale to train a product that answers questions without ever sending a reader back. A Chinese lab allegedly runs millions of queries against Claude to build a cheaper model that competes with Claude. Anthropic calls the second one industrial-scale theft and takes it to the Senate. It calls the first one fair use.
The legal distinction Anthropic draws is that its consumer terms of service explicitly prohibit using Claude's outputs to train competing models, while most publishers never had a contract with Anthropic at all, just an open website and, at best, a robots.txt file. That's a real legal difference. But it's a thin one to lean on when the underlying grievance, in both directions, is about who gets to profit from work they didn't do the hard part of creating.
I don't think Anthropic is uniquely villainous here. Every major AI lab is extracting value from the open web faster than it's returning any, and the industry-wide numbers make that plain. Google, the one company that has historically paid the web back through referral traffic, is itself under pressure as AI Overviews start keeping searchers on Google's own results pages instead of sending them onward.
But I do think Anthropic's specific position deserves more scrutiny than it usually gets, precisely because the company has built so much of its public identity around being the safety-conscious, values-driven alternative to everyone else. Safety and fairness are not the same thing. A model can be carefully aligned against generating harmful content and still be built on a crawling strategy that most publishers would recognize as taking without asking. Anthropic gets to define which of those problems counts as "ethics" when it talks about its own product, and it has consistently chosen the framing that doesn't require paying anyone.
The company has pushed back on Cloudflare's methodology, saying it can't independently verify how the ratios are calculated and that its newer search features are already sending more traffic back than the numbers suggest. That's a fair point to raise, and it's worth watching whether Anthropic's ratio keeps trending down as Claude's search product matures. Worth watching, but not yet resolved. As of the most recent data, Anthropic is still taking far more than it gives back, and it's still the loudest voice in Washington arguing that taking from Claude is a national security threat.


Comments