ChatGPT Outperforms Claude and NotebookLM in Long Document Tests

A new comparative test reveals significant differences in how major AI tools handle complex, lengthy reports, with one model demonstrating superior depth in data retrieval and synthesis.
While artificial intelligence tools are often marketed for complex coding or workflow automation, their most practical utility for the average user remains the ability to digest large amounts of text. Recent testing by XDA Developers compared three leading platforms—ChatGPT, Claude, and NotebookLM—by feeding them an identical 113-page report on a railway collision in England. The goal was to determine which tool could accurately extract specific details, synthesize scattered information, and interpret visual data without hallucinating or oversimplifying the content.
The test results challenged recent assumptions about the state of these models. Contrary to the narrative that ChatGPT has fallen behind its competitors, the tool emerged as the clear winner in this specific scenario. The evaluation focused on eight distinct questions ranging from simple fact retrieval to complex chronological reconstruction and chart interpretation. The findings suggest that for document-heavy tasks, the choice of AI assistant matters significantly, and the most popular option is not always the most accurate.
Precise data retrieval defines the winner
The first indicator of ChatGPT’s superiority was its attention to minute details. When asked for the exact time of the collision, the tool provided the specific timestamp of 6:42:57 PM, whereas the other two models rounded the time to approximately 6:43 PM. While this may seem like a trivial difference, it highlighted a deeper engagement with the source material. ChatGPT did not settle for the most obvious summary-level answer but instead dug into the specific data points within the report.
This advantage became more pronounced as the questions increased in complexity. When tasked with reconstructing the events leading up to the accident, ChatGPT provided the most detailed and logically ordered account. The other tools tended to generalize the sequence of events, whereas ChatGPT maintained a strict chronological flow that aligned with the technical evidence presented in the document. This capability is crucial for users who need to understand the 'how' and 'why' of complex events rather than just the 'what'.
Synthesis capabilities reveal deeper reasoning
The true test of an AI’s document understanding lies in its ability to connect disparate pieces of information. One question asked which pieces of evidence supported the report’s conclusion regarding why the train failed to stop. ChatGPT successfully aggregated multiple data streams, including the train’s recorder data, wheel-slide protection activity, and physical testing results of the contaminated rails. It also referenced post-accident braking tests and the investigation’s findings on when the driver initiated braking.
In contrast, the other models struggled to synthesize these different types of evidence into a coherent argument. They often isolated facts without clearly explaining how they interlinked to form the basis of the report's conclusion. This gap in reasoning ability is significant for professionals who rely on AI to summarize technical reports or legal documents, where missing a subtle connection between data points can lead to a fundamentally flawed understanding of the situation.
Interpreting visual data remains challenging
A key component of the test involved interpreting Figure 41 from the report, which required the AI to determine the position of a specific work bank relative to a mean-average line on a graph. This task tested the model’s ability to process visual data and correlate it with textual descriptions. The results indicated that while all three tools made an attempt, the accuracy varied. This highlights a remaining trade-off in current AI technology: while text processing has improved dramatically, the integration of visual analysis with textual reasoning is still a developing area.
For users relying on these tools for daily document management, the takeaway is clear. If your primary need is to extract specific, precise data points and synthesize complex arguments from lengthy reports, ChatGPT currently holds the edge. However, users should remain cautious, as no tool is infallible. The test underscores that while AI can streamline the process of understanding large documents, human oversight remains essential to verify the accuracy of the synthesized information, especially when dealing with critical or technical content.






