Skip to content

Bringing you stories that vibe with your fashion

Latest News

ChatGPT-User bypassed robots.txt on 54% of restricted pages

OpenAI agent ChatGPT-User accessed restricted web pages in 54 percent of cases during early 2026, according to traffic platform TollBit.

ChatGPT-User bypassed robots.txt on 54% of restricted pages

The OpenAI web agent ChatGPT-User accessed pages forbidden by robots.txt files in 54 percent of recorded cases during the first half of 2026, according to data released by web monitoring platform TollBit.

Among European websites in the monitored sample, ChatGPT-User recorded the highest rate of unauthorized page requests among all tracked artificial intelligence bots. By comparison, Bytespider recorded a 48 percent bypass rate, while PerplexityBot reached 42 percent. Across the complete sample of European and North American websites, artificial intelligence crawlers directed about 15 percent of total requests to pages with explicit robots.txt prohibitions.

A robots.txt file is a standard web protocol used by domain administrators to instruct automated web crawlers on which pages should not be visited. TollBit is a digital media technology platform that monitors web traffic and content access for online publishers.

Image generated: Nano Banana

Publisher Traffic and Bot Identification

TollBit clarified that its statistics reflect web traffic observed strictly within its own publisher network rather than across the entire public internet. The company defined a successful bypass as any successful server response to a URL request that a publisher had explicitly designated as off-limits for that specific crawler.

The analytics platform also noted that while the figures document actual requests made to restricted pages, the data does not establish the underlying reason for each individual query. Furthermore, TollBit stated that the figures cannot rule out the possibility of user-agent spoofing by third-party requests.

OpenAI Documentation and Access Controls

The high bypass frequency is linked to the core operational purpose of ChatGPT-User. According to OpenAI documentation, the agent accesses specific web pages when a user submits a prompt in ChatGPT or a custom GPT that requires live web content to answer.

OpenAI specifies that ChatGPT-User is not an automated web crawler and that because its requests are triggered directly by human end users, standard robots.txt restrictions may not apply to its activity. This distinguishes ChatGPT-User from OAI-SearchBot, which determines content eligibility for ChatGPT search answers, and GPTBot, which crawls website content to train base artificial intelligence models.

Because robots.txt files communicate publisher preferences rather than technical barriers, web servers continue to serve content to requesting agents unless strict security protocols are enforced. Technical experts note that preventing external agents from accessing sensitive resources requires access control mechanisms such as user authentication, authorization rules, network policies, or server-level blocking. Web management teams must therefore actively monitor site visibility and access controls to safeguard digital content.

Related

Leave a comment

Your email address will not be published. Required fields are marked *