Engineering Blog

How We Rebuilt Post-Incident Reviews Inside PagerDuty

by PagerDuty Engineering September 17, 2026 | 10 min read

A post-incident review is where an incident turns into improvement. It is the difference between a team that gets a little better after every incident and one that keeps relearning the same lessons.

Jeli’s post-incident reviews brought a structured approach to incident analysis: gathering the evidence of an incident, the conversation, the timeline, and the people involved, then turning that context into a review a team can learn from. That product thinking became the foundation for PagerDuty’s new Post-Incident Reviews.

Bringing that experience natively into PagerDuty created a new engineering opportunity: build on the platform’s existing identity, incident, permissions, security, and operational primitives while preserving the workflows and product principles that made Jeli valuable. The result is a review experience that sits directly in the incident lifecycle and creates a stronger foundation for what we can build next. This is the engineering story: the architecture we chose, the decisions behind it, and the tradeoffs we accepted. We wrote the product view of what it unlocks separately.

Designing for the platform

Jeli’s product thinking is the foundation this work stands on. The architecture around it was built for a different job. Jeli was designed to stand alone, and it had the shape of a good standalone product: its own front door and login, its own data model for users and incidents, its own permission system, its own ingestion path, and its own set of vendors. Jeli kept serving customers throughout the work described here; the native capability was designed and built alongside it.

As we brought that experience into PagerDuty, the design question was how to make the most of the platform around it. PagerDuty already has mature primitives for incidents, identity, permissions, security, analytics, and operational scale. Building Post-Incident Reviews directly on those primitives let us create a more connected experience while keeping the core product principles intact.

That led to a simple set of requirements. Reviewing an incident should be the natural next step after resolving one, in the same place the incident already lives, with relevant context (from Slack or Teams) available to the review. Access should follow the incident’s existing controls. The capability should inherit the platform’s operational, security, and compliance foundations. And the architecture should make it easier to extend the experience over time.

Those requirements produced the target architecture: preserve the product thinking, and build the capability on what PagerDuty already does well.

One boundary: the platform’s own primitives, and far fewer external dependencies. We organized the engineering around four outcomes: customer trust, reliability, time to market, and innovation speed.

Customer trust

For many teams, the most important property of a review tool is whether they are allowed to use it. When a capability lives outside the platform’s security and compliance boundary, regulated and international customers often cannot enable it without legal sign-off, and some never get there. We could have taken the standalone stack through certification, vendor by vendor, audit by audit. We chose to build inside the boundary instead, so reviews are covered by default, with no separate opt-in, under the same protections as the rest of a customer’s incident data.

Access comes down to whether you can see the underlying incident, which is itself scoped by PagerDuty’s existing roles and team controls. Every request to the review service is authorized against incident access, and review content lives with the incident as a single source of truth rather than being scattered across chat, documents, and tickets.

Reusing the platform’s existing access model also reduced the amount of authorization logic the new capability needed to own: less code, one model to reason about, fewer synchronization paths, and an access story that follows the incident itself.

Reliability

A retrospective about reliability ought to be reliable itself. Much of what a review depends on is gathered in the background: the incident timeline, the conversation in Slack or Teams, the transcript from the response call. If that gathering is fragile, the review starts on a weak foundation.

The foundations are what you would expect from a platform pipeline: a durable, replayable queue, retries with exponential backoff, and a dead-letter queue where failed imports can be inspected and replayed rather than lost.

The decisions worth writing about are the ones where we departed from the default. We tuned the queue for fairness over raw throughput. It is partitioned across many consumers and keyed per user, so one account importing an enormous channel occupies its own partitions instead of everyone’s, and we bound the number of messages per import so no single task can grow unbounded. The consumer lives in its own Kubernetes deployment, off the interactive API path, so ingestion load cannot slow the review someone else has open. Big imports land as bulk inserts while the read-heavy API leans on the database cluster’s reader instances, keeping the two workloads out of each other’s way.

The most contrarian call was about freshness. The default for a product like this is a fleet of background jobs keeping every review’s conversation data up to date, whether or not anyone is looking. We went the other way and deleted the background refresh entirely: a review’s data refreshes just in time, when someone opens it, with a minimum interval so we are not hammering upstream services and a manual refresh for anyone who wants the latest immediately.

Time to market

The fastest way to build something well is often to build less of it. PagerDuty already has hardened answers for the undifferentiated parts of an application: who the users are, what an incident is, how permissions work, how analytics are reported. By standing the new service on our standard platform patterns and reusing those primitives, we did not have to rebuild the scaffolding. We could spend our time on the review itself, and on a REST API whose query patterns we can see, load-test, and tune rather than discover under load.

That same instinct drove two larger decisions. The first was the data model. A model designed for a standalone world, with its own concepts of users and incidents, would have meant translating between two worlds indefinitely and slowing everything built on top of it. Designing it around PagerDuty’s own primitives from the start gave the native experience a common foundation with the rest of the platform and reduced the number of mappings future capabilities would need to understand. We kept the product thinking that worked while choosing a model designed to extend naturally with PagerDuty over time.

The second was real-time collaborative editing, which teams told us they wanted and which makes a group retrospective work. Letting several people write, comment, and assign follow-ups in the same document at the same time is one of the hard problems in software, and it is not where we add unique value. So we bought that capability as a headless collaboration framework and run it ourselves, as containers inside our own Kubernetes environment on our service mesh, with its supporting pieces on infrastructure we already operate: a small relational store for document metadata, an ephemeral cache for real-time fan-out, and object storage for the document bodies, which keeps rich documents from bloating the primary database. Our engineering effort went into the boundary around it: every request to the editor is proxied through our service and authorized against incident access. Buying the generic piece and securing its perimeter let us ship the differentiated experience far sooner.

Innovation speed

The strongest reason to build reviews natively is what becomes possible when review data can participate directly in the platform. Alongside incidents, services, and responders, it becomes another source of context for the capabilities we want to build on top: drafts and key moments generated automatically, cross-incident insights that surface recurring causes and frequently affected services, and an SRE agent whose recommendations get sharper because it can ground them in how past incidents were reviewed and resolved.

A shared platform foundation makes those connections much more direct. Instead of each new capability having to establish its own path to review context, incident and review data can evolve together as part of the same system.

We are already building on that foundation. The new experience can draft a review for you, and because generating a thoughtful draft takes time, we made the wait feel like progress rather than a frozen screen. Drafts are gated by an evaluation pipeline before we widen access, and a person reviews and edits every draft before it becomes the record.

Generation runs asynchronously through our AI platform, which publishes a completion onto a queue when the draft is ready. The interesting part is the last hop: getting that signal back to a browser, which cannot receive an arbitrary server callback.

We relay it, and we deliberately kept the realtime channel dumb. The socket carries only a small ready signal to a WebSocket room scoped to the incident; the page hears the nudge and refetches the real content from the REST API. We could have streamed the draft itself over the socket, but a dumb channel means the API stays the single source of truth, and because the queue delivers at least once, a missed or duplicate notification can never corrupt the draft.

The machinery behind the nudge stays lightweight for the same reason. A worker alongside the API tracks which clients are subscribed to which incident, polls job state on a short interval, and compares a fingerprint of the latest state so it only broadcasts on real change, backing off exponentially on long-running jobs. The backend-for-frontend layer that holds the live connections runs as many pods, so an event arriving on one instance is broadcast internally to reach sockets attached to another.

What responders see, and what comes next

Several people can already write in the same review at once, with comments and mentions; what we are still building is live updates for the structured pieces around the document, like markers and follow-ups, along with continued refinement of the AI drafting experience. We would rather ship the parts that remove real friction today and keep improving in the open than wait for everything to be perfect.

For responders, all of this engineering surfaces as one new tab on the incident they were already working on: a review that starts itself from the incident’s own context, inside the platform’s boundary, on the platform’s primitives.

For teams coming from Jeli, the continuity is deliberate. The product insight behind structured incident analysis remains, now connected more directly to the incident lifecycle in PagerDuty. Reviews can sit beside the incidents they came from, inherit the same access and security foundations, draw from incident context, and contribute to the cross-incident intelligence we can build on top of the platform.

That is the foundation we wanted: preserve what teams value about the review itself, use the strengths of the platform around it, and make every capability we add next easier to connect to the incident lifecycle.

This is the engineering companion to From Incidents to Insight: Closing the Post-Incident Review Gap. For the product view of what these changes unlock, start there.

About the authors

Girish Shankarraman is a Director of Engineering at PagerDuty, where he leads the Incident Analysis and Chat Experience products. Since joining PagerDuty, he has built and led teams across the product platform. An engineering leader with more than 15 years of experience, he focuses on turning ambiguous customer problems and operational data into platform capabilities.

Jacob Castle is a senior software engineer on PagerDuty’s Incident Analysis team and the tech lead for the Incident Reviews product. Previously, he helped lead projects across several of PagerDuty’s platform teams, including building the notifications platform.