Tech Souls, Connected.

Cloudflare Claims Perplexity Circumvents Robots.txt, Masks Bot Identity

Cloudflare claims Perplexity is evading anti-scraping measures across thousands of websites—fueling renewed debate over AI ethics and internet publishing rights.


Cloudflare Says Perplexity Is Scraping Despite Blocks

AI startup Perplexity is under fire for allegedly scraping websites that have explicitly opted out of AI crawling, according to a new report by Cloudflare, one of the world’s largest internet infrastructure providers.

  • Cloudflare says it detected Perplexity bots disguising their identity and bypassing standard anti-scraping tools like robots.txt.
  • The scraping activity was reportedly widespread, observed across tens of thousands of domains and millions of requests per day.

“We observed that Perplexity uses not only their declared user-agent, but also a generic browser intended to impersonate Google Chrome on macOS,” Cloudflare wrote.


How Perplexity Allegedly Circumvented Scraping Protections

According to Cloudflare, Perplexity’s bots used cloaked user agents and rotated autonomous system numbers (ASNs) — identifiers that can help websites trace where internet traffic is coming from — to mask their origin.

  • This effectively allows bots to slip past site blocks that would otherwise reject known scrapers.
  • Cloudflare claims to have fingerprinted the traffic patterns using a mix of machine learning and network-level analysis.

This behavior contradicts the established web standard that robots.txt files are meant to offer — a mutual understanding between website owners and web crawlers on what can or can’t be accessed.


Perplexity Responds, Denies Allegations

In response, Perplexity spokesperson Jesse Dwyer called the report a “sales pitch” and said the bots named in Cloudflare’s blog aren’t affiliated with the company.

“The screenshots in the post show that no content was accessed,” Dwyer said via email.

However, Cloudflare says it confirmed Perplexity’s involvement through customer reports, independent testing, and repeated crawler behavior consistent with the startup’s known infrastructure.

Cloudflare has since removed Perplexity from its list of verified bots and implemented new detection tools to block their activity.


A Pattern of Scraping Controversy

This isn’t the first time Perplexity has been accused of unauthorized content harvesting.

  • In 2023, Wired and other news outlets claimed Perplexity was plagiarizing their content without proper attribution.
  • During a live interview at TechCrunch Disrupt 2024, CEO Aravind Srinivas notably struggled to define the company’s stance on plagiarism, raising further questions about ethical boundaries.

These incidents underscore growing industry tensions around AI training data, copyright, and the value of original web content in the age of AI aggregation.


Cloudflare Steps Up Its Anti-AI Push

Cloudflare’s report is part of a broader push to defend publishers and creators against unauthorized scraping:

  • In July, it launched a new marketplace allowing websites to charge AI scrapers for access to their content.
  • CEO Matthew Prince has publicly warned that AI is undermining the internet’s business model, especially for journalism and independent publishing.
  • Last year, the company also released a free tool to block AI scrapers, aligning itself with a growing resistance to data harvesting without consent.

The accusations against Perplexity add to an intensifying debate about how AI companies source their training and real-time data:

  • Are robots.txt files still respected, or are they being quietly ignored?
  • Can startups justify scraping content in the name of AI innovation?
  • Who should decide how public web data is used—the site owners, or the model developers?

As legal battles over AI scraping heat up — including lawsuits from The New York Times, Getty Images, and other rights holders — cases like this may shape the future of data governance and AI development norms.

Share this article
Shareable URL
Prev Post

OpenAI’s ChatGPT Nears 700M Users, Dominates AI Engagement Charts

Next Post

No Shutdown, But Major Changes: Amazon Overhauls Wondery’s Structure

Read next