有人正在进行大规模漏洞扫描,并伪装成 ClaudeBot 等人工智能机器人。
Someone is running mass vulnerability scans, spoofing AI bots like ClaudeBot

原始链接: https://knownagents.com/insights

“代理网络指数”(Agentic Web Index)追踪了超过 5000 个网站的生态指标,报告显示过去 90 天内与人工智能相关的机器人流量增加了 11%。该数据将网络代理分为多种类型,包括 AI 爬虫、数据提供商、搜索爬取器和 AI 助手。 **关键发现:** * **主要活动:** AI 数据爬虫(如 ClaudeBot、meta-externalagent 和 Amazonbot)在爬取总量中领先,而 ChatGPT-User 则是 AI 助手流量的主要驱动力。 * **搜索索引:** 以 PetalBot 和 Amzn-SearchBot 为首的 AI 搜索爬取器,正在为 AI 驱动的搜索结果大量索引内容。 * **安全风险:** 目前存在活跃的“机器人欺诈”活动,恶意行为者冒充合法 AI 代理,扫描网站以获取敏感配置文件和凭据。 * **Robots.txt 的有效性:** 尽管一些代理遵守“禁止访问”(disallow)规则,但其他代理经常忽略这些规则。拦截效果因代理而异。 * **AI 推荐:** 该指数追踪了 AI 平台如何引用并将人类流量引导至网站,“计算机与电子产品”成为引文和推荐流量的主要类别。 这些数据为希望管理机器人流量、优化 AI 驱动的发现以及保护数字基础设施的网站所有者提供了关键见解。

Hacker News 最近的一场讨论指出,针对成千上万个网站的大规模漏洞扫描出现了明显的激增。尽管自动化扫描在互联网上是常态,但安全专家观察到了一种新的复杂趋势:扫描程序正越来越多地伪装成“ClaudeBot”等合法的 AI 机器人用户代理(user-agent)。 讨论参与者指出,这些请求往往无法通过 IP 验证和网页机器人认证,证实了它们是冒充搜索引擎或 AI 爬虫的恶意行为者。这些攻击活动似乎专门针对新型 AI 编程工具中的漏洞。 虽然一些评论者认为这只是典型的互联网“噪音”,可以通过屏蔽来自常见 VPS 提供商的流量来缓解,但原帖作者强调,近期这种特定伪装模式在统计学上的显著增长值得关注。社区认为,这种激增很可能意味着有人正在协同行动,以利用某个新发现的漏洞。
相关文章

原文

Key ecosystem metrics across 5,000+ websites using Agent Analytics and AI Chat Referral Tracking.

Agentrification

↑ 11%

Compared to the previous 90 days

The percentage of bot traffic that's AI-related

Traffic by Agent Type

AI Agent

AI Agent

Uses an actual web browser to autonomously complete complex tasks on behalf of a human user

AI Assistant

AI Assistant

Fetches website content in response to a user prompt, to include in an AI-generated answer

AI Coding Agent

AI Coding Agent

Fetches documentation and other resources to help build software

AI Data Provider

AI Data Provider

Crawls websites to supply structured content to AI systems as a third-party service

AI Data Scraper

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

AI Search Crawler

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

Archiver

Archiver

Captures and stores historical website snapshots for long-term digital preservation

Automated Agent

Automated Agent

Automates browser interactions programmatically without direct human supervision

Developer Helper

Developer Helper

Assists with testing, debugging, and ensuring website functionality

Fetcher

Fetcher

Retrieves web page metadata to power app features like link previews or feeds

Intelligence Gatherer

Intelligence Gatherer

Analyzes web content for brand safety, competitive insights, and ad targeting

Scraper

Scraper

Extracts large amounts of web data, often without explicit website permission

Search Engine Crawler

Search Engine Crawler

Systematically scans and indexes web pages to include in search results

Security Scanner

Security Scanner

Scans websites for security vulnerabilities, threats, and configuration weaknesses

SEO Crawler

SEO Crawler

Analyzes website structure and content to identify SEO improvement opportunities

Uncategorized

Uncategorized

Not yet assigned a type

Undocumented AI Agent

Undocumented AI Agent

Crawls websites without disclosing its purpose, collecting data for an unknown AI use case

Hover over each agent type for more information about what they do

Top Visiting Agents

bingbot
SRCH

Search Engine Crawler

Systematically scans and indexes web pages to include in search results

8.2%

Googlebot
SRCH

Search Engine Crawler

Systematically scans and indexes web pages to include in search results

7.9%

AhrefsBot
SEO

SEO Crawler

Analyzes website structure and content to identify SEO improvement opportunities

6.2%

Known Agent
DEV

Developer Helper

Assists with testing, debugging, and ensuring website functionality

5.5%

ChatGPT-User
ASST

AI Assistant

Fetches website content in response to a user prompt, to include in an AI-generated answer

3.3%

ClaudeBot
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

3.2%

PetalBot
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

3.2%

SemrushBot
SEO

SEO Crawler

Analyzes website structure and content to identify SEO improvement opportunities

3.0%

facebookexternalhit
FTCH

Fetcher

Retrieves web page metadata to power app features like link previews or feeds

2.6%

meta-externalagent
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

2.3%

Amazonbot
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

2.3%

MJ12bot
SEO

SEO Crawler

Analyzes website structure and content to identify SEO improvement opportunities

2.2%

Amzn-SearchBot
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

2.2%

DotBot
SEO

SEO Crawler

Analyzes website structure and content to identify SEO improvement opportunities

2.1%

Applebot
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

2.1%

Agents with the most activity

These bots scrape website content to train AI models. Some belong to AI companies, while others belong to third-party services that resell the data. Automatic Robots.txt can block unwanted scraping. Included agent types include AI Data Providers and AI Data Scrapers.

AI Data Provider

AI Data Provider

Crawls websites to supply structured content to AI systems as a third-party service

AI Data Scraper

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

AI scraping activity by agent type over time

Top Visited Website Categories

Computers and Electronics

5.3%

Business and Industrial

5.0%

Internet and Telecom

4.5%

Website categories with most activity

Top Agents

ClaudeBot
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

27.0%

meta-externalagent
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

19.5%

Amazonbot
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

19.2%

GPTBot
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

9.4%

Bytespider
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

7.3%

ShapBot
PVDR

AI Data Provider

Crawls websites to supply structured content to AI systems as a third-party service

5.6%

GoogleOther
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

3.9%

CCBot
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

2.3%

Timpibot
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

1.8%

YouBot
PVDR

AI Data Provider

Crawls websites to supply structured content to AI systems as a third-party service

1.4%

VelenPublicWebCrawler
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

0.8%

Diffbot
PVDR

AI Data Provider

Crawls websites to supply structured content to AI systems as a third-party service

0.5%

DeepSeekBot
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

0.4%

TerraCotta
PVDR

AI Data Provider

Crawls websites to supply structured content to AI systems as a third-party service

0.3%

AIWebIndex
PVDR

AI Data Provider

Crawls websites to supply structured content to AI systems as a third-party service

0.2%

Agents doing the most AI scraping

These bots fetch website content in real time to power AI assistants, coding agents, and other retrieval-augmented generation (RAG) tasks. Pages inform responses on the spot, such as when an assistant summarizes an article or a coding agent references documentation. Included agent types include AI Assistants and AI Coding Agents.

AI Assistant

AI Assistant

Fetches website content in response to a user prompt, to include in an AI-generated answer

AI Coding Agent

AI Coding Agent

Fetches documentation and other resources to help build software

AI fetching activity by agent type over time

Top Visited Website Categories

Computers and Electronics

2.3%

Travel and Transportation

1.9%

Business and Industrial

1.8%

Internet and Telecom

1.7%

Website categories with most activity

Top Agents

ChatGPT-User
ASST

AI Assistant

Fetches website content in response to a user prompt, to include in an AI-generated answer

86.4%

DuckAssistBot
ASST

AI Assistant

Fetches website content in response to a user prompt, to include in an AI-generated answer

3.3%

Perplexity-User
ASST

AI Assistant

Fetches website content in response to a user prompt, to include in an AI-generated answer

2.7%

Claude-User
ASST

AI Assistant

Fetches website content in response to a user prompt, to include in an AI-generated answer

2.5%

MistralAI-User
ASST

AI Assistant

Fetches website content in response to a user prompt, to include in an AI-generated answer

2.1%

Claude-Code
CODE

AI Coding Agent

Fetches documentation and other resources to help build software

1.1%

Shap-User
ASST

AI Assistant

Fetches website content in response to a user prompt, to include in an AI-generated answer

1.0%

Google-NotebookLM
ASST

AI Assistant

Fetches website content in response to a user prompt, to include in an AI-generated answer

0.4%

Cursor
CODE

AI Coding Agent

Fetches documentation and other resources to help build software

0.4%

Gemini-Deep-Research
ASST

AI Assistant

Fetches website content in response to a user prompt, to include in an AI-generated answer

0.0%

meta-externalfetcher
ASST

AI Assistant

Fetches website content in response to a user prompt, to include in an AI-generated answer

0.0%

GoogleAgent-URLContext
ASST

AI Assistant

Fetches website content in response to a user prompt, to include in an AI-generated answer

0.0%

Code
CODE

AI Coding Agent

Fetches documentation and other resources to help build software

0.0%

kagi-fetcher
ASST

AI Assistant

Fetches website content in response to a user prompt, to include in an AI-generated answer

0.0%

QualifiedBot
ASST

AI Assistant

Fetches website content in response to a user prompt, to include in an AI-generated answer

0.0%

Agents doing the most AI fetching

These bots crawl website content so it can be surfaced in AI search engines and AI-generated answers. Those answers often include citations or links back to the source pages. Included agent types include AI Search Crawlers.

AI Search Crawler

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

AI search indexing activity by agent type over time

Top Visited Website Categories

Travel and Transportation

6.8%

Books and Literature

4.9%

Computers and Electronics

4.6%

Arts and Entertainment

4.5%

Website categories with most activity

Top Agents

PetalBot
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

25.4%

Amzn-SearchBot
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

17.3%

Applebot
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

16.7%

OAI-SearchBot
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

11.9%

meta-webindexer
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

10.8%

Claude-SearchBot
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

10.4%

PerplexityBot
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

3.4%

LinkupBot
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

3.4%

Google-CloudVertexBot
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

0.3%

xAI-SearchBot
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

0.2%

AzureAI-SearchBot
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

0.2%

AddSearchBot
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

0.0%

Anomura
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

0.0%

MistralAI-Index
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

0.0%

amazon-kendra
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

0.0%

Agents doing the most AI search indexing

These bots use browsers to autonomously navigate websites, click through pages, and make decisions to complete tasks for people. Agentic UX best practices and Google PageSpeed Insights help evaluate how well websites support them. Included agent types include AI Agents.

AI Agent

AI Agent

Uses an actual web browser to autonomously complete complex tasks on behalf of a human user

AI browsing activity by agent type over time

Top Visited Website Categories

Computers and Electronics

0.0%

Internet and Telecom

0.0%

Travel and Transportation

0.0%

Business and Industrial

0.0%

Books and Literature

0.0%

Arts and Entertainment

0.0%

Website categories with most activity

Top Agents

Google-Agent
AGNT

AI Agent

Uses an actual web browser to autonomously complete complex tasks on behalf of a human user

43.0%

Manus-User
AGNT

AI Agent

Uses an actual web browser to autonomously complete complex tasks on behalf of a human user

40.1%

ChatGPT Agent
AGNT

AI Agent

Uses an actual web browser to autonomously complete complex tasks on behalf of a human user

16.9%

NovaAct
AGNT

AI Agent

Uses an actual web browser to autonomously complete complex tasks on behalf of a human user

0.0%

GoogleAgent-Mariner
AGNT

AI Agent

Uses an actual web browser to autonomously complete complex tasks on behalf of a human user

0.0%

AmazonBuyForMe
AGNT

AI Agent

Uses an actual web browser to autonomously complete complex tasks on behalf of a human user

0.0%

TwinAgent
AGNT

AI Agent

Uses an actual web browser to autonomously complete complex tasks on behalf of a human user

0.0%

Agents doing the most AI browsing

See which robots.txt rules are set across the web and how well agents follow them. An agent's Robots.txt Effectiveness measures the effectiveness of a disallow rule for it by estimating how much the agent reduces its traffic after it's blocked.

Top Rule-Following Agents

LinkupBot
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

100.0%

serpstatbot
SEO

SEO Crawler

Analyzes website structure and content to identify SEO improvement opportunities

100.0%

proximic
INT

Intelligence Gatherer

Analyzes web content for brand safety, competitive insights, and ad targeting

100.0%

ClarityBot
SEO

SEO Crawler

Analyzes website structure and content to identify SEO improvement opportunities

100.0%

IAS crawler
INT

Intelligence Gatherer

Analyzes web content for brand safety, competitive insights, and ad targeting

100.0%

Barkrowler
SEO

SEO Crawler

Analyzes website structure and content to identify SEO improvement opportunities

100.0%

um-IC
INT

Intelligence Gatherer

Analyzes web content for brand safety, competitive insights, and ad targeting

100.0%

SEOkicks
SEO

SEO Crawler

Analyzes website structure and content to identify SEO improvement opportunities

100.0%

Dragonfly
INT

Intelligence Gatherer

Analyzes web content for brand safety, competitive insights, and ad targeting

100.0%

AffsignalCrawler
INT

Intelligence Gatherer

Analyzes web content for brand safety, competitive insights, and ad targeting

100.0%

AdsBot-Google-Mobile
INT

Intelligence Gatherer

Analyzes web content for brand safety, competitive insights, and ad targeting

100.0%

meta-externalads
INT

Intelligence Gatherer

Analyzes web content for brand safety, competitive insights, and ad targeting

100.0%

UptimeRobot
DEV

Developer Helper

Assists with testing, debugging, and ensuring website functionality

100.0%

Leikibot
INT

Intelligence Gatherer

Analyzes web content for brand safety, competitive insights, and ad targeting

100.0%

TTD-Content
INT

Intelligence Gatherer

Analyzes web content for brand safety, competitive insights, and ad targeting

100.0%

Agents with the best Robots.txt Effectiveness percentages

Top Rule-Breaking Agents

Baiduspider
SRCH

Search Engine Crawler

Systematically scans and indexes web pages to include in search results

82.6%

SirdataBot
UNC

Uncategorized

Not yet assigned a type

84.5%

ShapBot
PVDR

AI Data Provider

Crawls websites to supply structured content to AI systems as a third-party service

90.4%

YandexBot
SRCH

Search Engine Crawler

Systematically scans and indexes web pages to include in search results

91.2%

Sogou web spider
SRCH

Search Engine Crawler

Systematically scans and indexes web pages to include in search results

92.7%

Mediapartners-Google
INT

Intelligence Gatherer

Analyzes web content for brand safety, competitive insights, and ad targeting

93.2%

Scrapy
SCRP

Scraper

Extracts large amounts of web data, often without explicit website permission

93.8%

Amazonbot
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

94.7%

HeadlessChrome
AUTO

Automated Agent

Automates browser interactions programmatically without direct human supervision

95.4%

linkfluence
INT

Intelligence Gatherer

Analyzes web content for brand safety, competitive insights, and ad targeting

95.9%

DotBot
SEO

SEO Crawler

Analyzes website structure and content to identify SEO improvement opportunities

96.1%

OAI-SearchBot
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

96.3%

ChatGPT-User
ASST

AI Assistant

Fetches website content in response to a user prompt, to include in an AI-generated answer

96.8%

archive.org_bot
ARCH

Archiver

Captures and stores historical website snapshots for long-term digital preservation

96.9%

Agents with the worst Robots.txt Effectiveness percentages

Top Blocked Agents

GPTBot
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

24.6%

CCBot
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

22.6%

ClaudeBot
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

21.7%

Bytespider
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

18.9%

PerplexityBot
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

17.9%

ChatGPT-User
ASST

AI Assistant

Fetches website content in response to a user prompt, to include in an AI-generated answer

16.1%

anthropic-ai
UND

Undocumented AI Agent

Crawls websites without disclosing its purpose, collecting data for an unknown AI use case

15.8%

meta-externalagent
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

15.6%

Amazonbot
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

15.6%

Diffbot
PVDR

AI Data Provider

Crawls websites to supply structured content to AI systems as a third-party service

14.0%

Claude-Web
UND

Undocumented AI Agent

Crawls websites without disclosing its purpose, collecting data for an unknown AI use case

14.0%

cohere-ai
UND

Undocumented AI Agent

Crawls websites without disclosing its purpose, collecting data for an unknown AI use case

14.0%

omgili
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

13.5%

Agents blocked by the most top websites

See which agents are most frequently impersonated, and how spoofing activity changes over time. A visit is considered spoofed when it claims a recognized agent identity but fails that agent's supported authentication method, such as verified IP or Web Bot Auth.

Active Threat: AI Bot Spoofing Campaign

We are observing a widespread campaign impersonating AI bots to scan websites for vulnerabilities. The attacker appears to be targeting credential and configuration paths used by AI coding tools. Contact us for more information, or check your own website.

Top Spoofed Agent Identities

Googlebot
SRCH

Search Engine Crawler

Systematically scans and indexes web pages to include in search results

0.5%

ChatGPT-User
ASST

AI Assistant

Fetches website content in response to a user prompt, to include in an AI-generated answer

0.1%

GPTBot
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

0.1%

OAI-SearchBot
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

0.1%

PerplexityBot
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

0.1%

ClaudeBot
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

0.0%

Applebot
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

0.0%

bingbot
SRCH

Search Engine Crawler

Systematically scans and indexes web pages to include in search results

0.0%

Perplexity-User
ASST

AI Assistant

Fetches website content in response to a user prompt, to include in an AI-generated answer

0.0%

MistralAI-User
ASST

AI Assistant

Fetches website content in response to a user prompt, to include in an AI-generated answer

0.0%

GoogleOther
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

0.0%

AhrefsBot
SEO

SEO Crawler

Analyzes website structure and content to identify SEO improvement opportunities

0.0%

Claude-User
ASST

AI Assistant

Fetches website content in response to a user prompt, to include in an AI-generated answer

0.0%

Claude-SearchBot
SRCH

AI Search Crawler

Indexes website content to possibly include as citations in AI-powered search results

0.0%

CCBot
SCRP

AI Data Scraper

Downloads website content to include in datasets used for training AI models such as LLMs

0.0%

The most impersonated agent identities

Recent Top Targeted Paths

/.config/anthropic/credentials/default.json

/.claude/settings.json

/.claude.json

/.hermes/.env

/.openclaw/.env

/.codex/config.toml

/.continue/config.json

/.aider.conf.yml

/service-account.json

/serviceaccountkey.json

/service_account.json

/firebase-adminsdk.json

/firebase-service-account.json

/.aws/credentials

/.aws/config

/.s3cfg

/.boto

/.npmrc

/.env.example

/.env.local

/.env.production

/.env.backup

/.env.old

/backend/.env

/api/.env

/admin/.env

/dockerfile

/docker-compose.yaml

/.docker/config.json

/terraform.tfstate

/credentials.json

/secrets.json

/secrets.yml

/key.json

/rclone.conf

Examples of recent top targeted request paths

See which AI platforms like ChatGPT, Perplexity, and Gemini cite websites and send them human referral traffic. Citations are estimated. Google's guide explains how websites can optimize their content to be more visible in AI chat responses (GEO).

Traffic by AI Chat Platform

ChatGPT

Claude

Copilot

DeepSeek

Gemini

Meta AI

Mistral

Perplexity

Referral activity by AI chat platform over time

Top Mentioned (Cited) Website Categories

Computers and Electronics

2.3%

Travel and Transportation

1.9%

Business and Industrial

1.8%

Internet and Telecom

1.6%

Website categories most frequently cited in AI chat responses

Top Clicked (Referred) Website Categories

Travel and Transportation

0.3%

Computers and Electronics

0.1%

Business and Industrial

0.1%

Internet and Telecom

0.0%

Website categories receiving the most referrals from AI chat

Data Scope

The Index is updated daily with completed days of traffic, security, and referral data from more than 5,000 websites using Agent Analytics and AI Chat Referral Tracking. The current partial day is excluded. Percentage-change tags compare the current period with the preceding period of the same duration. Agent names, operators, and classifications come from the Agent Directory, which is updated as new agents are discovered or existing agents change. Website categories follow the taxonomy used by Google AdSense. Participating websites are not a random sample of the entire web, and the qualifying set can change as websites connect, disconnect, or cross activity thresholds. Results characterize the observed network and broader directional trends; they should not be interpreted as a precise census of global web traffic.

Qualification & Aggregation

Only websites meeting minimum activity and data-quality requirements are included. Internal, test, incomplete, or anomalous data is excluded. Bot traffic percentages use total server traffic as their denominator. AI chat referral percentages use estimated human traffic, calculated by excluding identified bot visits from total server traffic. Rates are calculated for each qualifying website first, then averaged across websites and completed days. This gives each website equal weight regardless of traffic volume and prevents a small number of high-traffic websites from dominating the results. Daily charts are not smoothed, allowing normal seasonality to remain visible.

Measuring Robots.txt Effectiveness

An agent's Robots.txt Effectiveness estimates the reduction in its request rate associated with a full disallow rule. For each completed day, Known Agents establishes an agent-specific baseline from qualifying websites where that agent is allowed, adjusts the baseline for the overall traffic of each website where the agent is disallowed, and compares the expected activity with the activity actually observed. Only website-day observations with sufficient site traffic, agent activity, cross-site coverage, and expected volume qualify. Scores also require repeated observations across multiple websites and days. When an agent publishes a supported authentication method, only verified traffic is attributed to it.

Each qualifying website-day contributes equally. Scores range from 0%, meaning no measurable reduction, to 100%, meaning no qualifying requests were observed where the agent was disallowed. The headline Robots.txt Effectiveness metric gives each qualifying agent equal weight. Because this is an observational estimate rather than a controlled experiment, it measures an association with robots.txt rules but does not claim that robots.txt caused every observed difference. Top Blocked Bots is calculated separately using daily robots.txt scans of Similarweb's top 1,000 websites.

Identifying Spoofed Bots

Spoofing statistics measure traffic from visits that claim the identity of a known agent but fail a supported authentication method, such as published IP verification or HTTP message signatures. Each agent's daily rate is calculated against total server traffic for every qualifying website, then averaged across websites. A failed check indicates that the visit was likely impersonating the named agent; it does not identify the software or operator that actually made the request. Agents without a supported authentication method are not included in these measurements.

AI Chat Citations & Referrals

AI chat referral statistics count directly observed human visits carrying a recognized AI platform in the referring URL or campaign source. Visits without usable referral information cannot be attributed to an AI platform. Citation statistics are estimates based on requests from agents known to retrieve content for AI platforms. Those requests indicate that content may have informed a response, but they do not confirm that a source appeared as a citation to a user. Because AI platforms do not provide a complete public record of their sources, citation results should be interpreted as directional patterns rather than exact citation counts.

Can journalists and media organizations use this data?

Absolutely. You may cite The Agentic Web Index with attribution and a link to this page. For interviews, fact-checking, background context, or a more specific breakdown for a story, contact us and include your deadline.


Do you work with researchers?

Absolutely. We welcome thoughtful research into how agents and bots are changing the web. Tell us about your research question, timeframe, and intended use. Depending on the scope and data constraints, we may be able to provide additional context, compare approaches, or explore a joint analysis.


Can I request a specific analysis?

Yes. If you need a breakdown by agent, operator, activity type, website category, or time period that is not shown here, contact us. When the underlying data supports it, we can examine the question and provide a focused analysis.


联系我们 contact @ memedata.com