AI
Aug 25, 2026


As AI moves into production, the teams responsible for keeping enterprise systems reliable are being asked to monitor behavior that traditional infrastructure tools were never designed to see.
[For more news, click here]
An AI system can remain online, respond to users, and pass the usual infrastructure checks while its answers become less accurate, its costs rise, or an agent starts behaving differently from what its developers expected. For the engineers responsible for keeping enterprise technology running, that creates a reliability problem that cannot be understood through uptime and latency alone. SRE teams have traditionally worked with signals such as uptime, latency, errors, and service level objectives, but those measures tell only part of the story when an application depends on an AI model. A model can return a response without an infrastructure failure while producing an inaccurate result, drifting from expected behavior, or consuming more resources than anticipated. Platform teams face a related challenge as they give developers access to AI tools while maintaining the security, standards, and controls around them. That shift is increasingly showing up in the work of site reliability engineering and platform teams.
New research from Dynatrace surveyed 919 senior leaders, decision makers, managers, and supervisors involved in SRE, platform engineering, or IT operations at enterprises with annual revenues of at least $500 million. The findings suggest these teams are now having to adapt established SRE and platform engineering practices as AI becomes part of production infrastructure. Monitoring AI systems for performance and accuracy is the most common AI powered capability among SREs, with 58% reporting its use, while half say they are using AI for automated incident response. Yet the growing amount of information available to these teams can create problems of its own. Nearly half of SRE respondents, 49%, say too many data sources make it harder to define and manage service level objectives, while 45% point to too many metrics as a challenge.
AI Is Adding New Problems to an Already Complex Stack
An enterprise AI application can span several layers of technology, with different teams responsible for different parts of the system. When an AI service starts producing poor results, checking whether the underlying infrastructure is healthy may not reveal where the problem began. Among platform engineers, 37% identify integration with existing tools and systems as their biggest challenge, reinforcing the difficulty of bringing these different layers into the same operational view. It is against that backdrop that Dynatrace signed a definitive agreement on August 13 to acquire AI observability company Arize in a cash and stock transaction valued at $915 million. The proposed combination would bring AI native evaluation and observability together, connecting work done during development with what teams see once AI applications reach production.
“SRE and platform engineering laid the groundwork for modern digital reliability, but AI is rewriting the rules. Enterprises need to now move from managing systems to orchestrating them, connecting observability, automation, and agentic AI to operate at the speed these initiatives demand, turning insight into action at scale,” said Steve Tack, Chief Product Officer at Dynatrace. “This research also reflects why we recently announced our intent to acquire Arize. AI engineering teams have been evaluating in one set of tools while operations teams monitor in another, and that gap is no longer sustainable as AI moves deeper into enterprise production.”
Bringing those two sides closer together matters because testing a model before deployment does not tell an operations team everything it needs to know once the system is live. Behavior can change after deployment, and teams may need to determine whether a problem comes from the model, its data, the application, or the infrastructure around it. Arize brings AI observability and evaluation capabilities, while Dynatrace's existing platform covers application, infrastructure, and broader system observability. The company has also been expanding its telemetry capabilities through its planned acquisition of Bindplane earlier this year, as the volume and complexity of operational data continues to grow. AI is generally meeting expectations around system reliability and developer productivity, according to the research, but the reported gains are weaker when it comes to reducing operating costs and shortening incident detection and resolution times. Organizations are already using AI for automated incident response, yet they are continuing to prioritize visibility and human oversight before expanding automation. That caution matters when an AI system can produce an unexpected result without a conventional infrastructure failure.
The deal is intended to give teams a more connected view of those systems. Whether it delivers the level of control enterprises need will become clearer if the acquisition closes, but the operational problem identified in the research is already emerging as companies put more AI into live systems. Models can change, workloads can grow, and agents can behave differently from conventional software, leaving SRE and platform teams responsible for a wider set of signals than uptime and latency alone.
Related Articles
Dynatrace to acquire Bindplane to expand telemetry data capabilities
Bank Muscat Launches Oman First AI Powered Banking Command Center
AI Agent Governance Is Falling Behind Enterprise Adoption, Optro Research Finds
Related Articles