logo
|
Blog
    BotManager

    Should You Block All AI Bots? Why Understanding and Distinguishing Their Purposes Matters

    Should AI search crawlers, training crawlers, and agents all be managed the same way? Learn why access policies should differ by content, function, and bot purpose.
    Lo
    Louise
    Oct 01, 2026
    Should You Block All AI Bots? Why Understanding and Distinguishing Their Purposes Matters
    Contents
    SummaryWhat Should You Look at First When Classifying AI Bots?Search: Is This Content Intended to Be Discoverable?Training: Does Public Content Also Need to Be Available for Training?Agent: Reading Content and Taking Actions Should Be Treated DifferentlySetting AI Bot Traffic Policies for Your ServiceWhat Matters in AI Bot Management Is the Policy Applied After ClassificationFAQ

    Summary

    • Allowing or blocking all AI bots indiscriminately can create a conflict between opportunities for search visibility and the need to protect data.

    • Search crawlers, model training crawlers, and AI tools that access websites on behalf of users have different purposes and should be granted different levels of access.

    • Different access policies should be applied to public content, private data, and functions such as login and purchasing.

    • Rather than relying solely on a bot’s name or User-Agent, organizations should evaluate identity, access paths, request frequency, and actual behavior together.


    AI bots visit websites for many different reasons. Some look for information to use in search responses, while others collect content for model training. Still others access websites on behalf of users to compare products or carry out tasks such as making reservations.

    If businesses simply group all of these automated visitors together as “AI bots” and block them indiscriminately, they may miss opportunities that AI-driven traffic can provide.

    In July 2026, Cloudflare announced options for managing AI traffic by dividing it into Search, Agent, and Training categories. For newly registered domains, Cloudflare stated that Search traffic would be allowed by default, while Training and Agent traffic would be blocked on pages displaying ads. Akamai also refined its existing AI Bots category in September, dividing it into AI training crawlers, AI search crawlers, and real-time fetchers and agents.

    Although the two companies do not use exactly the same classification system, they share the same overall direction: access policies should be designed according to the purpose of AI bot traffic.


    What Should You Look at First When Classifying AI Bots?

    Type

    Purpose

    AI search crawlers

    Collect and index pages so that information can be discovered and used in AI-powered search

    AI training crawlers

    Collect content for model training and improvement

    Real-time fetchers & AI agents

    Retrieve pages or perform tasks on the web in response to user requests

    AI bots operate with a variety of purposes. However, even a request claiming to come from a search crawler should be evaluated separately if its identity cannot be verified or if it operates outside the permitted scope.

    Such traffic may actually be malicious automation and could violate service policies through unauthorized data collection, account abuse, excessive requests, or other activities.

    At the same time, not every automated request should automatically be considered malicious. This is why organizations need to determine which types of AI bot activity should be permitted within their services.

    Search: Is This Content Intended to Be Discoverable?

    For service pages designed to be publicly accessible—such as product descriptions, help documentation, and blog posts—it is worth checking whether AI search crawlers can access them.

    More people are now using AI-powered tools to search for information. As visibility and citations in AI-generated answers can create additional opportunities for customers to discover a brand, the access granted to AI search crawlers may affect whether content can be discovered or used by AI search services.

    Training: Does Public Content Also Need to Be Available for Training?

    Making information publicly available and allowing that information to be used for model training are two separate decisions.

    Organizations should determine how their content, pricing information, and documentation may be used according to their internal policies, and then decide how requests for training-related data collection should be handled.

    Some AI bots may perform both search and training functions. Therefore, classifying a bot solely by its name may result in policies that behave differently from what was originally intended.

    Agent: Reading Content and Taking Actions Should Be Treated Differently

    There is a meaningful difference between an AI tool retrieving a product page and summarizing it for a user, and that same tool logging in, claiming a coupon, or submitting an order.

    A fetcher retrieves content required for a specific request, while an agent performs multiple steps on behalf of a user. Although the two may sometimes be managed together, it is useful to distinguish between simply accessing content and performing actions that change the state of a service.

    For example, a service may allow an AI tool to view publicly available hotel listings while requiring separate authentication, authorization, request limits, and other controls when the tool attempts to create an account, reserve inventory, or confirm a booking.

    Authorization to act on behalf of a user does not automatically mean that every transaction is authorized. This distinction should therefore be considered carefully.


    Setting AI Bot Traffic Policies for Your Service

    The table below is not intended as a default configuration. Instead, it provides examples of criteria that can be used when distinguishing AI bot traffic. Policies should be adjusted according to each service’s public access scope, usage policies, and authentication methods.

    Access Target

    Search Crawlers

    Training Crawlers

    Real-Time Fetchers & Agents

    Malicious or Unverified Automation

    Public blogs & product descriptions

    Consider allowing access based on search visibility goals

    Allow or restrict according to content usage policies

    Consider read access; evaluate abnormal high-volume requests separately

    Verify identity and request patterns, then restrict or block as necessary

    Members-only information & personal data

    Exclude from public search

    Deny access

    Determine permitted access after verifying user authentication and authorization

    Block access and investigate

    Login, sign-up, coupon & purchase APIs

    Generally not intended for search indexing

    Deny access

    Verify authentication, action-level permissions, request frequency, and transaction flows

    Restrict or block based on abusive behavior

    When applying these criteria, organizations need to consider both ‘who is making the request?’ and ‘what are they doing?’

    User-Agent and IP information can provide useful signals, but neither independently proves who the requester is or whether the traffic is being used for an authorized purpose.

    Request paths, repeated calls within short periods, behavior after login, unusually high collection volumes, and error rates should also be analyzed together.

    In addition, robots.txt is a way to communicate access preferences to cooperative crawlers. Protecting private information or preventing unauthorized requests requires additional measures such as authentication and access controls.

    ① Define which pages should be public and which functions should be protected
    → ② Verify the purpose and identity of known crawlers
    → ③ Define permitted access for each page and API
    → ④ Adjust policies based on logs and performance

    Organizations should monitor both indexing and traffic changes on pages that remain accessible to search crawlers, as well as automated requests and false positives involving sensitive functions. Looking at only one side of these metrics can lead to incorrect decisions.

    What Matters in AI Bot Management Is the Policy Applied After Classification

    Depending on the service, both being discoverable through search and AI-generated answers and protecting content, accounts, and transaction functions can be important.

    Even within the same website, these requirements may differ depending on the page or action involved.

    Therefore, instead of asking:

    ‘Should we allow or block AI bots?’

    organizations should ask:

    ‘What types of automation should be allowed to perform which actions on which resources?’

    STCLab BotManager analyzes access environments and behavioral patterns to detect malicious bots and macros, while using filters and policies that combine signals such as IP addresses, User-Agents, and request attributes to manage access.

    These capabilities can be used to identify abnormal automation according to the access criteria defined by each service. Organizations can first distinguish between content they want to make publicly accessible and functions they need to protect, and then review the filters and policies that can be applied to actual traffic.

    Learn more about BotManager · Learn more about BotManager filters and policies

    The starting point for AI bot management is to stop treating every automated request as a single category of bot traffic.

    For content that organizations want users and AI services to discover, access paths should be reviewed accordingly. For data and transaction functions that need protection, policies should instead be defined according to purpose and authorization.


    FAQ

    Q. If we allow AI search crawlers, will our brand definitely appear in AI-generated answers?

    No. Allowing access is only one of the conditions that may enable a service to crawl or index your content. Whether your brand is actually displayed or cited depends on the criteria used by each search or AI service.

    Q. If we block AI training crawlers, does that mean we will no longer appear in AI search?

    Not necessarily. Some services use separate crawlers for search and model training. However, because a single crawler may serve multiple purposes, organizations should review both the service provider’s documentation and the policies being applied.

    Q. If a tool identifies itself as an AI agent, can we allow it to make purchases or reservations?

    The agent’s name alone is not sufficient. Organizations should verify the account owner’s scope of authorization, the agent’s actual permissions, and the potential for abuse at each stage of the transaction before determining what actions should be permitted under their service policies.

    Share article
    Contents
    SummaryWhat Should You Look at First When Classifying AI Bots?Search: Is This Content Intended to Be Discoverable?Training: Does Public Content Also Need to Be Available for Training?Agent: Reading Content and Taking Actions Should Be Treated DifferentlySetting AI Bot Traffic Policies for Your ServiceWhat Matters in AI Bot Management Is the Policy Applied After ClassificationFAQ

    STCLab Inc.

    RSS·Powered by Inblog