NewsTradingSentimentEventsCommunityBriefing
Tech

AI Models Lose Focus on Long Tasks, Risking Legal Compliance

By Tech Desk · · 2 min read
A stack of paper documents with a single sheet highlighted in the middle

Research shows AI performance drops 39% in long conversations, causing critical instruction failures.

Key points

  • Chroma's 2025 study found 18 leading AI models degrade in performance as input length increases, a phenomenon termed context rot.
  • Microsoft Research and Salesforce data shows a 39% average performance drop when tasks are spread across multiple conversational turns.
  • Legal professionals face compliance risks because AI models often fail to maintain safety instructions and accuracy in long documents.

A growing body of research reveals that leading AI models suffer from significant performance degradation as input length increases. This phenomenon, often described as context rot, means that even the most advanced systems struggle to maintain accuracy and instruction adherence when handling lengthy documents or extended conversations.

The issue is not merely a matter of memory capacity but a structural weakness in how models process information. As inputs grow, the ability to retain initial instructions and identify critical details buried in the middle of a text declines sharply, posing serious risks for professional applications like legal analysis.

Performance Degrades in Long Contexts

Researchers at Chroma tested eighteen major AI models in 2025 and found that all exhibited declining performance as the size of the input grew. The failures began well before the input filled the model's maximum context window, indicating that the problem is inherent to the architecture rather than a simple limit of storage. This effect was independently confirmed by earlier studies from Stanford and UC Berkeley, which showed that models are least accurate when the key information is located in the middle of a document stack.

The implications for professional workflows are direct. When a lawyer asks an AI to extract specific clauses from a 300-page agreement, the model is forced to attend to the middle of a long input, the precise position where its reliability is lowest. This creates a blind spot in tasks that require meticulous attention to detail across large volumes of text.

Multi-Turn Conversations Reduce Accuracy

Beyond static documents, the way users interact with AI in real-time chats further exacerbates the problem. A 2025 study by Microsoft Research and Salesforce found that performance fell by an average of 39% when tasks were spread across multiple conversational turns compared to single-shot prompts. Models that make an early error in a long dialogue tend to build upon that mistake, failing to recover as the conversation continues.

This decay in control means that safety guardrails and style guidelines provided at the start of a session lose their effectiveness over time. A rule that is strictly followed in a two-paragraph answer may be ignored in a twenty-page output, creating a gap between intended behavior and actual results.

Compliance Risks for Legal Practice

The American Bar Association’s Formal Opinion 512 requires lawyers to have a reasonable understanding of AI tools' limitations. The documented decay in model performance over long inputs represents a critical limitation that must be accounted for in professional practice. Relying on these tools for extensive diligence without recognizing their tendency to drift or lose focus may violate standards of competence.

As reported by JD Supra, the consensus among researchers is that current models are not yet reliable for unsupervised, long-horizon tasks. Professionals must adopt strategies such as breaking down complex requests into smaller segments or frequently restarting conversations to mitigate the risk of instruction drift and ensure that critical constraints remain enforced.

Based on reporting by JD Supra, compiled by the Tradingbird desk.

Read next

More in Tech

More from the Tech desk

All desk stories