It might be the summer break for publishers in the Northern Hemisphere, but one segment of the audience never goes on holiday: AI crawlers.
For years, robots.txt files were how publishers allowed web spiders to crawl their website, and signalled to search engines their permission to index articles. Now it can also signal publishers’ approach to access by AI models and services.
Robots.txt is often compared to a ‘do not walk on the lawn’ sign: a polite request that visitors are free to ignore. However, with licensing deals between AI and content companies negotiated behind closed doors, we decided to look at the robots.txt files of 70 major publishers to analyse their differing approaches and identify some best practices.
As you’ll read, some are prioritising visibility in AI search and responses. Others are taking a more defensive position, protecting content that they believe has value through subscriptions and licensing. Most — 60 of the 70 we looked at — restrict at least one AI crawler, most often those associated with model training.
Your AI crawler policy will depend on the importance of content discovery, the maturity of your subscription business and whether there is a market for content licensing. Robots.txt might be just an initial layer of protection, but research from Known Agents shows that 96% of bots respect its specifications, so it's worth checking whether your file is up to date.
Whatever route you take, who you let crawl your content and how they do it are commercial - and possibly legal - decisions that need careful consideration.
Our eight pieces of housekeeping advice are a good place to start, while our recent two-part series on resilient content businesses in the age of AI (part one and part two) provide further food for thought.
Get in touch if you would like to discuss the analysis — or the trade-offs for your publication or outlet — in more detail.