According to iAfrica.com, two-thirds of South African news sites take no action against AI scraping. The report highlights a split between larger publishers that can afford technical blocks and smaller outlets that let automated tools copy their articles unchecked.
AI scraping is the practice of using software to automatically copy website content for training large-language-models (LLMs), the type of artificial intelligence that generates text. When a bot visits a news page, it can download the headline, body text and images in seconds, then feed that data into an AI system that later reproduces similar articles without credit or payment to the original publisher.
The cost of defending against such bots can be significant. Effective measures include deploying bot-detection firewalls, adding CAPTCHAs, or negotiating licensing agreements with AI providers. Those tools require specialised staff or third-party services that many small and medium-sized publishers cannot justify in their budgets.
Regulatory backdrop
South Africa’s current copyright law protects original works, but it does not specifically address the bulk extraction of text for AI training. The Protection of Personal Information Act (POPIA) governs the handling of personal data, yet it does not cover non-personal news content. Without clear statutory guidance, publishers rely on general copyright infringement claims, which are costly to pursue and often result in lengthy court battles.
For advertisers and small businesses that depend on local news for brand exposure, the trend raises practical concerns. If AI systems repurpose news stories, the original outlet may see reduced traffic and lower ad revenue, limiting the reach of local marketing campaigns. Smaller publishers, which already compete for limited ad spend, could lose a vital platform for reaching regional customers.
Policy makers are watching the issue closely. Some observers suggest a licensing framework where AI developers pay a fee per article used for training, similar to models being debated in Europe. Others argue for a voluntary code of conduct that encourages AI firms to seek permission before harvesting large volumes of content. Until a clear approach emerges, the burden remains on individual news sites to decide whether to invest in protection or accept the risk.
The divide highlighted by iAfrica.com points to a broader question about the sustainability of South Africa’s news ecosystem. If the majority of outlets continue without safeguards, the long-term impact could be a shrinking pool of local voices, which in turn affects the diversity of information available to businesses and consumers alike.
The tension this report describes is playing out well beyond South Africa. Major international publishers and news groups have in recent years pursued a mix of strategies against AI scraping: some have signed direct licensing deals with AI developers for the right to train on their archives, others have pursued litigation over unauthorised use, and a growing number have adopted technical blocking measures such as the robots.txt standard or dedicated bot-detection services. What is comparatively unusual about the South African picture this report describes is not the existence of unprotected content, which is common globally among smaller outlets everywhere, but the scale: two-thirds is a notably high share for a national news market, and it suggests the gap here is less about awareness of the risk than about the cost of the available defences relative to the advertising revenue a smaller South African outlet actually earns.
The absence of a South Africa-specific legal answer also has a knock-on effect on smaller publishers’ own commercial options. Without clear domestic case law establishing whether AI training on scraped news content infringes copyright, a small outlet weighing whether to pursue a licensing negotiation with an AI developer has little leverage: the developer has no clear legal reason to pay for something it is not yet clearly required to pay for, leaving voluntary technical blocking as the only immediately available defence for publishers unable to litigate the question themselves.



