Article 6DS1V Sites scramble to block ChatGPT web crawler after instructions emerge

Sites scramble to block ChatGPT web crawler after instructions emerge

by
Benj Edwards
from Ars Technica - All content on (#6DS1V)
hiding_hero_1-800x450.jpg

Enlarge (credit: Getty Images)

Without announcement, OpenAI recently added details about its web crawler, GPTBot, to its online documentation site. GPTBot is the name of the user agent that the company uses to retrieve webpages to train the AI models behind ChatGPT, such as GPT-4. Earlier this week, some sites quickly announced their intention to block GPTBot's access to their content.

In the new documentation, OpenAI says that webpages crawled with GPTBot "may potentially be used to improve future models," and that allowing GPTBot to access your site "can help AI models become more accurate and improve their general capabilities and safety."

OpenAI claims it has implemented filters ensuring that sources behind paywalls, those collecting personally identifiable information, or any content violating OpenAI's policies will not be accessed by GPTBot.

Read 12 remaining paragraphs | Comments

External Content
Source RSS or Atom Feed
Feed Location http://feeds.arstechnica.com/arstechnica/index
Feed Title Ars Technica - All content
Feed Link https://arstechnica.com/
Reply 0 comments