- PagerDuty /
- Engineering Blog /
- Swapping a Product Engine Mid-Flight
Engineering Blog
Swapping a Product Engine Mid-Flight
The hardest problem in platform engineering isn’t building something new. It’s replacing something customers already depend on, while they’re still depending on it. We’ve done that five times at PagerDuty, in five different ways. Each stressed a different part of the system. The pattern underneath them is what I want to write down.
Swapping an engine mid-flight is a useful metaphor because the hard part isn’t building a better engine. You have to understand what the plane depends on, train the crew, prepare the aircraft, carry the replacement and tooling before takeoff, engineer the swap without losing altitude, and keep an escape path until you know it is complete.
In that sense, software isn’t very different.
Customers spend years building workflows, integrations, automation, and customizations around a product. Their users build muscle memory on top of all of it. At PagerDuty, those dependencies stretch across the software development lifecycle and become especially visible in the incident lifecycle:
Detect → Triage → Diagnose → Remediate → Improve
We ship more than 600 production changes a week across 15,500+ paid customers. A change can look perfectly reasonable inside one feature and still break an important handoff somewhere else.
The five examples here start out looking unrelated: modernizing Schedules, upgrading the Chat experience responders use during incidents, making Post-Incident Review native after the Jeli acquisition, evolving Pricing and Packaging beyond a per-seat model, and introducing AI Agents into a high-trust operational workflow.

For a platform, the lifecycle is the product.
Not Schedules. Not Chat. Not Post-Incident Review. The chain, and the handoffs between them. Foundational change is a systems problem before it’s a product one.
Five principles that have held up for me
I don’t think these are comprehensive, nor are they a lift-and-shift playbook. They’re the ones I keep coming back to as the product surface, customer context, and dominant risk change underneath us.
1. Customer Journeys > Services
SLOs and change-failure rates tell us whether the underlying systems are healthy. Watchtower asks the complementary question: can a customer still complete a critical workflow end to end?
A service can be green while the person using it is stuck.
For migrations, the implication is simple: progression should be gated by evidence, not just the calendar.
2. Cohorts, Not Individual Customers
If your migration plan assumes customers use the product the way you designed it, it has already missed the hard part. This is what I’ve written about before as Design for Disobedience: customers will use extensible products in ways we never intended, documented, or sometimes even imagined.
Map what customers actually built before designing the migration. Then recognize that the unit of migration at this scale isn’t a customer. It’s a cohort.
Revenue tier tells us who to sequence carefully and how to communicate. It doesn’t tell us how hard someone is to move. Segment on use case, complexity, dependency depth, reversibility, and the actual risk of moving. Then build a path for each cohort instead of a bespoke plan for every customer.
3. Preserve on Purpose
Some boundaries are contracts customers depend on. Others are accidental architecture, artifacts built over time.
Preserve the first kind. Aggressively remove the second.
And be just as deliberate about eventually retiring the first. Compatibility you keep forever stops being a promise and starts becoming a tax.
4. Manage One-Way Doors Appropriately
Sequence changes by how hard they are to undo, not by what’s convenient to build. A UI can usually be restored quickly. Once new APIs have written records against a data model the old system can’t understand, there may be no logically safe way back.
Run both paths in parallel where you can. Take the reversible steps first. Spend the confidence you earn on the irreversible ones.
Easy to reverse isn’t the same as safe to introduce. That’s where timing comes in.
5. Firm Deadline, Their Calendar
Hold a firm internal deadline for retirement. Let customers choose their date inside it. Without an end date, nothing ever retires. You run two systems indefinitely, and the cost shows up years later as the reason you can’t move quickly.
But the date inside that window belongs to the customer, and so does the moment. An Admin evaluating configuration on a Tuesday afternoon and a Responder handling an outage at 2 AM have very different tolerance for surprise. A payments company may not absorb foundational change during the holiday season. A tax platform isn’t touching workflows in April. Much of Europe slows down in August.
Migration timing that ignores this isn’t rigor. It’s just a calendar we made up.
Plan the end of life from day one. Reach it at the end of the migration.
Before takeoff
Those principles only matter if they change what we do before the first customer moves:
- Talk to more customers than feels necessary. Treat every assumption about how they use the old thing as unverified until somebody confirms it.
- Feature-flag and sample from every cohort. Disproportionate confusion means a mis-drawn cohort, not a difficult customer.
- Give customers sandboxes to test. Beta feedback tells us what customers say; sandbox behavior tells us what their users actually do.
- Enable Customer Success before the announcement. A customer’s perception of the migration is set by how Customer Success handles the first call.
- Run a premortem. Assume the migration failed, write down why, and you have your escape paths before you need them.
What happens if we don’t?
We discover what customers depend on at the same time they do, except they’re already frustrated when we find out.
Schedules is where most of this got tested at once.
Schedules: build the cohorts first
Schedules are one of PagerDuty’s oldest and most foundational product surfaces. We invested in rearchitecting it to ensure we don’t break core customer workflows, such as on-call, as these are critical parts of customer trust.
Moving from layer-based schedules to the newer shift-based model is a genuine one-way-door problem. The underlying model changes substantially (such as, shifts as first-class objects and native iCalendar support) to enable more flexible coverage, shifts, and responder configurations. The destination is better, but that doesn’t give us license to disrupt customers on the way there.
We started by mapping what customers had actually built. Terraform pipelines, API automation, overrides, and operational conventions surfaced that an ordinary beta release wouldn’t have uncovered. Some weren’t explicit product contracts. They were still dependencies.
The first move was to route all net-new customers onto the new API only. That doesn’t shrink the work already in front of us, but it stops the problem from growing. An open-ended problem becomes a bounded one.
Then came segmentation.
I’ve come to think of migratability as the right axis, not customer size. Migratability isn’t one number. We score a handful of signals we can read directly from a customer’s configuration and usage, and combine them into a composite that determines which cohort a schedule belongs to.
Every schedule has a lossless equivalent in the new model. That part was never the question. What varies is how many ways there are to express the same pattern. Schedules that map with no ambiguity upgrade automatically. Schedules with more than one valid configuration go through a user-initiated flow where an agent proposes the mapping and a human approves it. Here’s an example of the input signals that led us to cohorting:
| Signal | What it tells us |
|---|---|
| Total layers | Structural complexity of the coverage pattern |
| Type of schedule | Whether the pattern is machine-detectable |
| Anti-pattern overrides | Whether the schedule means what it says |
| Edit surface (UI, API, Terraform) | Where the migration has to land |
| Escalation policies referencing it | Blast radius of getting it wrong |
| Users on the schedule | How much muscle memory is attached |
| Incident volume | How much room there is to be wrong |
| Customer Success input | Customer context that configuration alone can’t tell us |
Cohorting told us which path a customer should take. The sandbox told us whether we’d hypothesized correctly. Configurations surfaced during dry runs that we’d mis-bucketed, and each became a correction to the upgrade tooling instead of a support ticket after cutover.

The distribution turned out to be the point. Schedules aren’t snowflakes; they cluster into a handful of recognizable shapes. 60% map with no ambiguity and upgrade automatically, with no human involved. The other 40% have more than one valid configuration, so the analysis is done for the customer and the only lift is reviewing the proposed mapping and saving it. Nobody has to reconstruct their own coverage from scratch. My team has written more in depth about how we clustered them and handled the migration.
The rearchitecture was a multi-year effort. The phased upgrade is ongoing, and it spans multiple quarters by design. Onboarding a customer onto a fresh tool is trivial next to changing the data model under every customer at once while both keep running. What it has already shown us is that cohorting changes the economics. Doing it this way means evidence precedes moving customers.
Self-serve the low-risk majority so you can afford to white-glove the complex few.
We also ran old and new paths in parallel, kept both APIs available while customers moved, and gated each stage before taking the next less-reversible step.
A changelog entry is documentation; it isn’t a migration plan.
What happens if we don’t?
Without cohorts, there’s effectively one upgrade path, and it fails in one of two directions. Build it for the hardest configurations and everyone moves at the pace of the most complex few. Build it for the typical customer and it breaks on the ones it can’t parse.
The second failure is the expensive one. A path designed for the median can therefore hurt exactly the accounts that have built the most on top.
Chat: migrating muscle memory
Chat is almost the inverse problem. There may be very little data to migrate, but enormous behavioral dependency.
Responders work through Slack and Microsoft Teams using muscle memory. The new Incident cards and dedicated incident-channel experience make it easy for engineers to absorb information and agents to work in context. That didn’t make surprising someone during an outage acceptable.
Our rollout was deliberately boring. We explained the customer problem first, exposed dedicated incident channels through Early Access, and phased the rollout rather than flipping everyone at once. In-product walkthroughs, a preparation checklist, and a safe test incident gave responders a chance to build familiarity before they needed the new behavior under pressure.
We also preserved customer control through per-service configuration and opt-out, and avoided meaningful responder-facing changes during predictably bad operational windows.
Better isn’t the bar; predictability is.
Without predictability, we turn a UX improvement into operational friction. It’s like asking a pilot to learn a new cockpit after the warning lights have already come on.
That dedicated channel turned out to matter for another reason. Scribe Agent can capture bridge transcripts on its own, but Chat is where responders see that context, the summaries, and agent output while the incident is still unfolding.
That changed how I think about Chat in the lifecycle. It isn’t just where humans coordinate. It’s also one of the places where an agent’s work becomes legible to humans.
Post-Incident Review: make the seams disappear
Jeli is where Preserve on Purpose points in the other direction.
We acquired Jeli for its product thinking about how teams learn from incidents: conversation, timeline, responders, and evidence brought together into a structured review. One of my favorite quotes from Nora Jones, founder of Jeli, is “Every incident is an opportunity for learning”. But Jeli was also a standalone product, with its own login, user and incident models, permission system, ingestion path, and vendors. Keeping those seams would have meant paying a translation tax on every capability we built afterward.
Jeli customers are going through a real migration, far fewer of them than Schedules, with the standalone product reaching end of life this December. What they get is a review sitting next to the incident it came from, inside the same security and compliance boundary. What they give up is a separate login.
Permissions show the tradeoff concretely. Rather than carry Jeli’s standalone authorization forward, review access now derives from access to the underlying incident: one less model to synchronize, one less boundary to explain.
Compatibility is valuable when customers depend on the boundary. It becomes technical debt when nobody should depend on it.
Native data also compounds. When an incident closes, a review draft assembles from incident context, the Slack or Teams conversation, transcript, and timeline, so teams refine a narrative instead of reconstructing one. And once reviews are native, they become structured input the rest of the lifecycle can use: key moments, recurring-cause analysis, cross-incident insight, action items tracked in the customer’s own system of record.
Here’s how the pieces fit together, with SRE and Scribe Agents + Chat Experience + AI Post-Incident Reviews completing the Incident Management Lifecycle.

What happens if we don’t?
The acquisition never really integrates. More importantly, the Improve phase remains a document stranded at the edge of the workflow. Making the review native turns yesterday’s incidents into learning that the rest of the lifecycle can apply for the future.
Pricing and Packaging: the product boundary isn’t the system boundary
Pricing changes look like product changes, but in practice several systems move together: billing and entitlements, the feature-to-SKU map, and the sales and renewal motion.
I learned this one the hard way. In one transition, a customer ended up back on a plan we thought we had retired because one renewal path was still open. We found it because we went looking, and closed the path. In another rollout, the packaging was ready, but enabling the capability broadly would have added load to the back-office systems at exactly the wrong time. We held the rollout.
The packaging was ready. The wider system wasn’t.
What I’d carry forward: “Can the customer move from old → new?” and “What paths can put them back into old, or make the new state unsafe?”. A migration isn’t complete until both have good answers. Otherwise, we declare victory when the product ships while back-office, GTM, or infrastructure dependencies quietly recreate the state we thought we’d retired.
There was another consequence that became clearer as we moved toward agents:
Per-seat pricing counts people. An increasing share of the work isn’t being done by people.
Agents: AI-native means rethinking the lifecycle
You can’t bolt the future onto foundations you were unwilling to change.
Transforming a product to be AI-first and AI-native isn’t about bolting AI features onto the existing workflow. It starts with a more fundamental question:
If the SRE Agent is going to operate as a trusted first responder, what would we design differently from first principles?
The first SRE Agent architecture gave us a useful lesson. A single agent with a large context and collection of tools worked well for initial triage. As we pushed it toward deeper investigation, the architecture started showing its limits: context degraded as investigations grew, hypotheses ran sequentially, moderately complex investigations could stretch past ten minutes, and a human couldn’t inject new information while the agent was running.
Those weren’t isolated bugs. They were consequences of the architecture.
That forced a different design: multiple focused agents and a more reactive execution model where hypotheses can run concurrently, results arrive incrementally, and human input can interrupt or redirect the investigation.
That is what I mean by AI-native: the lifecycle itself changes rather than AI sitting on top of the old one. Autonomous operations isn’t a feature that gets shipped. It’s the direction that change points.
It also changes the trust problem. A traditional feature produces reasonably deterministic behavior inside a bounded workflow. An agent reasons, calls tools, acts on enterprise context, and sometimes encounters situations we didn’t anticipate in the original test set.
That’s why our production AI agents treat evaluation, permissions, observability, auditability, and guardrails as parts of the architecture rather than launch gates bolted on at the end. Real-world misses feed back into golden-set evaluations scored in CI, so failures become regression cases instead of anecdotes. And when an agent’s output quality drops below threshold in production, it pages the owning team like any other incident.
A trusted agent needs clean foundational data and routing, a surface where its work is visible to responders, and incident context that survives beyond the event itself. It also needs a commercial model that reflects the value and usage of software doing work alongside humans.
None of those migrations existed for the agent. But together they created a much better foundation for one.
Shared agent memory is one of the concrete mechanisms that connects those phases. Operational history, prior incidents, conversations, diagnostics, and context captured earlier in the lifecycle can carry forward instead of being reconstructed from scratch each time the work moves into another phase.
The interesting change isn’t that we added AI.
It’s that who flies the plane was never a design decision. Now it is.
When we miss something
We will.
At this scale, some dependency or edge case will escape the model. When that happens, acknowledge it quickly, put a human on the problem, and feed the miss back into the next cohort.
Customers can tolerate a mistake. They’re far less tolerant of a vendor that minimizes the disruption or hides behind process.
Closing
The interesting part isn’t the moment of cutover. By then, most of the difficult work should already be behind you: dependencies understood, the replacement onboard, temporary scaffolding in place, teams prepared, and the escape route tested.
Reliability is what customers actually buy, and on a platform it isn’t a property of any single feature. It’s a property of the handoffs. Our foundational architecture is built for zero-planned downtime. Last quarter’s Event-to-Notification pipeline reliability was 99.9981%, even as we swapped these engines.
Five different mechanics. Five different ways to get the migration wrong. The discipline underneath them is the same.
You can’t bolt the future onto foundations you were unwilling to change.