Enterprises burned by AI agents still race to remove human approval

VentureBeat's latest Pulse data show an inversion at the center of enterprise agent rollout: companies that have already watched an AI feature pass internal testing only to disappoint customers are moving most aggressively to remove humans from the release decision. Among 108 enterprises surveyed in July, 49% reported at least one customer-visible incident over the previous year after a feature cleared internal checks, essentially unchanged from 50% in June. Trust in automated evaluation rose from 5% to 13% over the same window. Among the burned respondents, 85% said they were already letting agents push code or change systems without approval in limited cases, or were building toward that practice. Among those with no comparable incident, that figure was 61%. The source describes these splits, drawn from groups of 41 to 53 respondents, as directional rather than a market census.

The most revealing cross-tab is the one the survey itself flags. Among the 53 enterprises that had experienced an AI feature pass testing only to disappoint a customer, 4% placed complete faith in automated checks. Among the 41 enterprises with no comparable incident, 24% said the same, a sixfold difference in full confidence between those with the most and the least direct evidence that the release gate can fail. Ten of those 41 unburned respondents placed full faith in the gate. The burned group fell to 2 of 53. The source describes that gap as unsurprising, and it is. The interesting question is what those respondents then do with their diminished trust.

The counterintuitive finding sits one question later. Among the burned group, 85% were pursuing end-to-end deployment automation for at least some low-risk workflows. Among the unburned group, the figure was 61%. Only 11% of burned respondents said they would reject no-approval deployment automation over the coming years, against 24% of unburned respondents. The overall number did not shift: 67% of the 108 enterprises either already permit agents to push code in limited cases or are modifying pipelines to allow it, with 37% in the first bucket and 30% in the second. The burned respondents are not driving that average up; they are pulling further away from it than the unburned cohort pulls back. A passing evaluation score, for them, sits inside a deployment pipeline they are still scaling.

Two plausible readings sit behind that gap. The first is recklessness: burned respondents have decided to trust the very automated checks they have already seen fail. The second is deployment maturity: organizations running more agents at higher volume and across more consequential workflows are more likely both to encounter customer-visible incidents and to have the engineering infrastructure that makes automated deployment operationally tractable. The source does not establish which reading dominates. What it does establish is that a customer-visible incident has not slowed the move toward autonomy in the group with the most evidence that internal testing can miss defects.

Volume is the bridge to the next-order consequence. Per-deployment failure rates were not measured in either wave of the survey. If that rate stays constant while the proportion of enterprises using no-approval deployment rises, and while each of those enterprises ships more decisions through the gate, absolute incident counts could grow even when the share of affected companies does not. The July wave holds the customer-visible-incident proportion steady at roughly half. Production-side incidents could still be rising inside that flat headline. The data does not show that, but it does not foreclose it.

The production monitoring side tells the rest of the story. Among 106 valid responses, 26% said inline quality assertions, automated judges or guardrails that score live outputs, were their primary way of watching deployed agents. Another 26% centered on transaction traces; 24% emphasized gateway metrics such as latency, error and cost. Trace and gateway data reveal outages, slowdowns and broken requests. A fluent, fast and confidently wrong answer can pass all three. Roughly half of respondents monitored whether the agent was functioning, and just over a quarter automatically checked whether its output was correct in production.

The gap is sharpest among the 40 respondents already permitting no-approval deployment in at least limited cases. Only 28% of that group reported automatically checking the meaning and correctness of live answers. In other words, most enterprises that have already removed a person from at least some release decisions have not installed semantic-quality monitoring as a production backstop. The source does not characterize how those enterprises handle bad outputs after release, but the asymmetry is direct: a release gate has been automated faster than the production check that would catch the gate's failures.

The vendor data offers a different signal. OpenAI's native evals and traces narrowly led as the primary platform at 18%, followed by Confident AI's DeepEval at 17% and Braintrust at 15%. Anthropic's Claude Console and Workbench tied with organizations reporting no dedicated evaluation platform, at 12%. Braintrust's primary-platform share rose from 8% in June to 15% in July, the report's largest single-vendor move and the only one it flags as statistically significant. DeepEval climbed from 12% to 17%. Use of no purpose-built platform declined five points to 12%, though that smaller shift does not by itself confirm a trend. Specialist tools are gaining share faster than the generic ones.

Broader footprint figures are larger because most enterprises run more than one tool. OpenAI's native evaluation appeared somewhere in 31% of stacks, DeepEval in 27%, Braintrust in 22% and Anthropic's native tooling in 20%. Custom internal tooling reached 14%, with Weights & Biases Weave and open-source Langfuse at 11% each. Those figures describe adoption, not product performance, and the source does not establish that any vendor produces more reliable agents than the others. What the data does support is a separate observation about buying criteria. Integration ease replaced cost as the leading factor, rising 12 points to 39%; cost fell from 28% to 23%; evaluation accuracy ranked second at 28%. Buyers want a tool that drops into existing development and monitoring pipelines. Their leading success metric remains evaluation consistency at 38%, followed by fewer failures and regressions at 20%.

Budget intentions bracket the contradiction. People-centered review workflows edged ahead of production observability as the most frequently cited area for increased investment, 31% to 30%. Automated evaluation pipelines ranked third at 19%, followed by testing for safety and policy compliance at 16%. Only 6% said their reliability and evaluation budget was not increasing. Among the burned respondents, 38% said people-centered review would see the fastest investment growth; 24% of unburned respondents said the same. The strategy that emerges is automation with a human backstop: remove the person from the release checkpoint, then spend more on people downstream to catch the misses that automated checks let through.

The source points at one implication and stops there: the model may not scale. Agent deployments and automated checks can grow in step with software volume. Reviewer hours do not fall at the same rate. Enterprises may be replacing a human approval step with a larger downstream review function rather than eliminating human oversight entirely, a tradeoff the survey was not designed to measure. Raindrop's CTO, quoted in the source, framed the shift as a movement away from enumeration toward anomaly detection, describing enterprise eval sets as being reduced and deprioritized as agent systems grow more complex. That is an attributed vendor view, not an independent finding.

The July wave shows confidence rising before measured failure incidence improved, and a release gate being automated while the production check that would catch its failures remains incompletely installed. Whether the downstream human review function, growing fastest among the most-burned respondents, can absorb an increasing share of automated-deployment volume without becoming the new bottleneck is the operational question the survey itself does not attempt to answer.

Subscribe to AI Enthusiast Log

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe