Size: 200
Location: San Francisco, CA
Customer Since: 2020
In February 2023, Google launched its Bard chatbot with a slick promotional demo. In one of the example prompts, Bard incorrectly claimed the James Webb Space Telescope had captured the first images of an exoplanet. Astronomers spotted the error within hours—but by the next trading day, parent company Alphabet’s market cap had dropped by roughly $100 billion.
The cost of a single hallucination was bigger than the GDP of most countries.
This incident captures something every team shipping AI agents now has to reckon with: AI failures don’t look like traditional software failures. They tend to be more subtle, and much harder to catch before they impact operations.
That’s the problem PagerDuty and Arize are solving together.
Arize is an AI and agent engineering platform built around a three-stage operating model: observe, evaluate, improve. It connects those three stages into a continuous loop, giving engineering teams the visibility to catch quality issues in production and the workflows to fix them.
Integrated with PagerDuty’s digital operations and incident management platform, that loop now turns AI quality signals into prioritized incidents—routing the right humans to the right tasks, fast.
Read more about how PagerDuty and Arize built end-to-end observability for AI agents.
Why AI Agents Drift After Launch
In a demo, an AI agent may perform perfectly. But in production, real users introduce behaviors you never tested for.
“One of the key patterns that we see is that these agents may launch successfully…but over time, they start to degrade,” says Richard Young, Director, Partner Solutions Architecture at Arize.
Quality inevitably drifts. Tool calls go wrong. Unlike a server going offline (where every monitor lights up at once), AI failures tend to be more subtle, often surfacing as slow degradations in response quality.
Traditional monitoring covers uptime, system errors, and latency, but legacy platforms weren’t designed to surface AI-specific failure modes like model drift, hallucinations, or incorrect tool calls. To stay ahead of these issues, teams need a system that treats AI failures as high-priority incidents.
Why PagerDuty
Together, Arize and PagerDuty turn the inevitable challenges of AI deployment into rapid learning opportunities. By feeding incident data back into observability workflows, this integration speeds up response times and removes friction for the teams shipping these agents.
“Arize and PagerDuty together turn AI quality into a proactive operational discipline,”
– Richard Young, Director of Technical Partnerships, Arize
Arize Detects, PagerDuty Acts
Arize captures a trace of every AI interaction—every prompt, tool call, and model output, plus latency and token counts. Together, those traces form a system of record for the agent’s behavior. After curating the criteria and datasets, the system runs continuous evaluations on those traces, scoring things like answer relevance, hallucinations, tool choice, etc. Each evaluation then feeds into custom metrics that track behavior over time, like the percentage of responses flagged as hallucinations in a 24-hour window.
When a metric crosses a defined threshold, Arize fires an alert into PagerDuty, which routes it to the right responder with the context they need to triage and act fast. As Rich Young puts it, PagerDuty “helps to operationalize the insights and the alerts that Arize detects in our system.”
Both Arize and PagerDuty know the value of this loop firsthand. Both teams use the integration internally to keep our own AI agents on track.
At Arize, the company dogfoods its platform on Alyx, its own AI engineering agent. When Alyx started skipping an explanation step in a multi-step workflow, Arize evaluations caught the regression, an incident fired into PagerDuty, and the team triaged with context-rich alerts to ship a correction.
At PagerDuty, the Shift Agent (which helps on-call teams manage scheduling and coverage) has a response relevance threshold of 95%. If the chance of irrelevant responses crosses 5%, Arize fires an alert into PagerDuty and the team that owns the agent immediately gets to work.
“PagerDuty is part of that ecosystem that helps to ensure that we keep the quality of your agents up to par,” Young says.
From Reactive Triage to Proactive Quality Management
Together, PagerDuty and Arize deliver measurable impact:
- Early detection: Arize provides the framework to catch quality issues early, lowering mean time to detection.
- Faster resolution: PagerDuty lowers mean time to resolution by routing each AI incident to the right responder with the context they need to triage. Plus, with the new SRE Agent integration, PagerDuty can ingest the same Arize logs that triggered the alert, recommend specific mitigation steps, and reference similar past incidents to accelerate resolution.
- Self-improving agents: Instead of a static system that ships once and slowly degrades, AI agents connected to Arize and PagerDuty get better with use. Every interaction becomes data, closing the loop between what the agent does in the real world and how the team builds the next version.
Resolve Issues Before They Impact Operations
If you’re shipping AI agents, the question isn’t whether they’ll drift in production—it’s whether you’ll catch it before your customers do. Discover how PagerDuty and Arize can help you ship AI agents with confidence. Try PagerDuty free for 14 days.
Watch the full session to see PagerDuty and Arize cover this integration in detail at PagerDuty On Tour, including a look at how the SRE Agent ingests Arize traces to accelerate AI incident response.