CNN outlines how AI changed its reporting on the Epstein files

8 August 2026

CNN journalists used artificial intelligence, automated data collection and document-search tools to examine the large volume of material released by the US Department of Justice in connection with Jeffrey Epstein, while retaining human verification for editorial decisions, according to a presentation at AI4 2026 in Las Vegas.

The session, titled How Journalists Used AI to Cover the Epstein Files, featured Ryan Struyk, CNN Director of AI Innovation, senior reporter Marshall Cohen and senior data editor Matt Stiles. AI4’s current speaker programme confirms Struyk’s role as CNN’s Director of AI Innovation.

The presentation offered a practical example of how AI is beginning to change investigative journalism without necessarily replacing the traditional work of reporters.

Cohen compared the Epstein documents with previous large government releases, including Hillary Clinton’s State Department emails, the Mueller report and transcripts from the congressional investigation into the January 6 attack on the US Capitol.

In those cases, newsrooms often divided hundreds or thousands of pages between teams of journalists and reviewed them manually. The scale of the Epstein material made that approach increasingly difficult.

CNN therefore used Google Pinpoint as one of its principal tools for organising and searching the collection. Pinpoint can use optical character recognition and other technologies to search scanned documents, images, handwritten material, audio and other records while automatically identifying frequently mentioned people, organisations and locations. Google has subsequently added generative AI functions allowing journalists to query documents while linking responses back to the underlying source material.

For CNN, one of the immediate advantages was being able to place large volumes of material within a searchable environment rather than relying on individual journalists to keep separate records of names, references and relationships appearing across the documents.

Cohen said the technology helped reporters identify communications between individuals and begin mapping relationships contained in the files. Journalists then carried out the conventional reporting process of checking the original records, establishing context, contacting those mentioned and requesting responses before publication.

That distinction was central to the presentation.

CNN said it does not treat AI-generated findings as finished journalism. Instead, the systems are used to narrow large datasets, locate potentially relevant material and accelerate technical work that would otherwise absorb significant newsroom resources.

Building a system to capture millions of files

Stiles described a separate engineering challenge: acquiring the material itself.

The Justice Department did not provide the files as a single straightforward download. According to the CNN presentation, records were spread across large numbers of individual links and pages, requiring the newsroom to construct an automated system capable of detecting releases, downloading files, recording what had been collected and identifying gaps.

CNN first created a monitoring process using Python that checked the relevant Justice Department pages at regular intervals. When the structure or content of the page changed, the system sent an alert to the newsroom through Slack.

Once new material appeared, CNN used a combination of web crawling and systematic requests to identify and retrieve the underlying files.

Stiles said AI coding assistants helped the data team build and modify parts of this infrastructure quickly, including dealing with changes to website behaviour and maintaining logs of downloaded material.

The logging system was particularly important because journalists needed to know which records CNN had obtained and which remained outstanding.

It also allowed the newsroom to identify material that had subsequently disappeared from the Justice Department website while retaining a record of files already downloaded.

Rather than allowing AI to determine the conclusions of the investigation, CNN used it largely as an engineering and research assistant.

Stiles described the approach as similar to having additional junior developers available to help write scrapers, test systems and consider possible technical problems, while journalists and engineers remained responsible for the architecture and verification.

AI finds redaction problems

One of the more significant reporting outcomes came from material that traditional document-search tools could not easily examine.

The Justice Department release included photographs and video as well as conventional documents. CNN worked with AI visual-search company Visual Layer to examine more than 100,000 images contained within the released material.

That analysis identified images that appeared to contain information that should have been redacted, including images of children and documents containing personal identifying information.

CNN subsequently reported that more than a dozen photographs remained publicly accessible without appropriate redactions for nearly a month. After CNN contacted the Justice Department, the affected files were removed and replaced with redacted versions, according to CNN’s reporting.

CNN’s February reporting said the material included unredacted images of minors as well as passports and driving licences containing information such as identification numbers, addresses and dates of birth. The Justice Department told CNN at the time that it was continuing to address victim concerns, personally identifiable information and files requiring further redaction.

The example illustrates a potentially important application of AI in journalism: using machine-search capabilities to find accountability stories within datasets too large for practical manual examination.

Human experience still determines what is news

The CNN team was also clear about where the technology failed.

Cohen said one of the most basic questions in journalism remained difficult for AI to answer: what is actually new?

The Epstein case has generated court proceedings, criminal investigations, civil cases and extensive reporting over almost two decades. Millions of newly released records can therefore include a mixture of previously undisclosed material and documents that have already been public for years.

An AI system may locate a name, email or document, but it cannot necessarily determine whether the information represents a genuine new development.

That required reporters with long-standing knowledge of the case, including journalists who had followed previous Epstein proceedings and the trial of Ghislaine Maxwell.

The presentation also referred to the wider reporting community. Journalists could identify potentially important material surfaced by other reporters, search CNN’s own document collection to locate the original record and then independently examine and verify it.

This combination of machine search and human institutional knowledge was presented as more useful than asking AI simply to summarise an enormous collection.

Accuracy remains the limit

The session repeatedly returned to the risks of using AI in reporting where errors can have serious consequences.

The Epstein material is particularly sensitive because merely appearing in an email, address book, photograph or government record does not itself demonstrate wrongdoing.

CNN therefore said human review remained necessary throughout the process.

The newsroom tracked where files originated, logged downloads, tested collection systems and reviewed the source material behind potential stories.

That approach is consistent with the design of tools such as Pinpoint, where AI-generated responses can be traced back to the original documents rather than being treated as independent evidence.

AI expands what data teams can attempt

The practical effect, according to Stiles, was not eliminating reporters or data journalists but expanding the scale of projects a newsroom could reasonably attempt.

Systems that previously might have required days of development could sometimes be constructed or modified within hours. When websites changed or downloads failed, AI coding tools could help engineers develop and test alternative approaches more quickly.

The underlying journalistic standards, however, remained unchanged: reporters still needed to establish what a document meant, whether information was previously known, whether allegations were supported and what responses should be included.

Recent work by Stiles also shows that his role remains centred on conventional data journalism. In July, he was among the CNN journalists producing mapping and data visualisations following the Kumamoto earthquake in Japan, using geographic and seismic data to explain the scale and location of the event.

The Epstein project therefore offers a useful indication of where AI may have its strongest immediate impact in large newsrooms.

Rather than automatically producing articles, the technology is increasingly being used further upstream: monitoring government sources, acquiring records, processing scans, recognising entities, searching images, writing data-processing code and finding connections that reporters can subsequently investigate.

For journalism, that may prove more consequential than automated writing. The technology reduces the amount of time spent locating information, but the judgement over whether that information is new, accurate and newsworthy remains with journalists.

© 2026 cij.world

front page info
LATEST NEWS