According to a new report by VentureBeat, many large organizations in the AI sector have given their automated agents more autonomy, but still do not fully trust automated evaluations to guarantee the quality and performance of these agents. Among the 157 companies surveyed, half had deployed at least one AI agent in the past year that, despite passing internal evaluation, ultimately failed with customers. Only 5 percent of organizations stated they fully trust automated evaluation, with the most commonly cited weakness being that evaluations often do not align with actual results.
While agent autonomy is advancing faster than the development of evaluation infrastructure, two-thirds of organizations either currently allow agents to operate without human intervention or are designing mechanisms to enable such operations within a year. Most companies still rely on basic tools or native evaluations provided by model vendors, and only a quarter of organizations implement real-time quality control on agent outputs. Meanwhile, future investments are increasingly directed towards human supervision and monitoring of production. This evaluation gap is considered a fundamental challenge for the AI industry.

