Aidanix
Aidanix
From learning to earning, with AI
Startfree
AI News & UpdatesBusiness

Evaluation gap: Low trust in automated assessments in AI organizations

Aidanix Team3 minJuly 17, 2026
Evaluation gap: Low trust in automated assessments in AI organizations

According to a new report by VentureBeat, many large organizations in the AI sector have given their automated agents more autonomy, but still do not fully trust automated evaluations to guarantee the quality and performance of these agents. Among the 157 companies surveyed, half had deployed at least one AI agent in the past year that, despite passing internal evaluation, ultimately failed with customers. Only 5 percent of organizations stated they fully trust automated evaluation, with the most commonly cited weakness being that evaluations often do not align with actual results.

While agent autonomy is advancing faster than the development of evaluation infrastructure, two-thirds of organizations either currently allow agents to operate without human intervention or are designing mechanisms to enable such operations within a year. Most companies still rely on basic tools or native evaluations provided by model vendors, and only a quarter of organizations implement real-time quality control on agent outputs. Meanwhile, future investments are increasingly directed towards human supervision and monitoring of production. This evaluation gap is considered a fundamental challenge for the AI industry.

ارزیابی عامل هوش مصنوعیاعتماد به ارزیابی خودکارخطاهای عملیاتیخودکارسازیپایش تولید
نظرات

هنوز نظری ثبت نشده — اولین نفر باش.

Meta's AI Glasses: Increased Risk of Covert Recording and Bans in Public PlacesGoogle purchased Spirit Airlines employee data; flight attendants are concerned about privacy.Meta ran an ad for an AI app that generates porn of American female politicians.
Get motivated