Regulation vs. reality in AI deployment
Companies are increasing the autonomy of their AI models but are not fully confident in the safeguards meant to keep them in check. In the UK and Germany, regulators are pushing for more rigorous validation before deployment. However, according to VentureBeat Pulse Research, over half of the 157 surveyed enterprise teams reported incidents where their AI systems passed internal tests but then caused visible problems with customers. This disconnect is widening as companies accelerate deployment without enough confidence in their tools.
Automated trust is a false sense of security
Despite growing use of AI in customer service, finance, and logistics, trust in automated evaluation systems remains low. Only 5% of enterprises say they fully trust their automated systems today. More than a quarter of those surveyed say their main problem is that the tests don’t predict real-world performance.
Many organizations rely on built-in evaluations from their model providers or have no evaluation tools at all. This lack of robust testing means that even when agents appear to behave well during internal checks, they can misfire in real-world scenarios. The risk is clear, and yet many companies are moving forward with limited oversight.
Autonomy outpaces evaluation maturity
Two-thirds of the companies surveyed either allow automated deployment of low-risk agents without human oversight or are building systems to do so within a year. This rapid expansion of autonomy, however, is not matched by a mature evaluation landscape. Only 25% of organizations run real-time quality checks on their models in production.
This growing trust in automation is leading to what the study calls an “evaluation gap.” The more independence enterprises give to AI agents, the more they are relying on evaluations that cannot fully predict success. The risk is not just technical — it’s reputational, as customer-facing failures can erode user confidence.

