Tickd.ai
← The Tickd Guide

Ethics & Responsible Use

Why You Shouldn't Use LLMs to Write Engineering Post-Mortems (And How to Keep Your Blameless Culture Intact)

It is tempting to dump raw PagerDuty logs into Claude and ask for a polished root cause analysis. Here is why outsourcing your team's post-mortems to generative AI destroys psychological safety and misses the real systemic failures.

Updated 10/5/2026

We have all been there. It is 3:00 AM, the database cluster is screaming, the production environment is a smouldering crater, and your Slack channels are a blur of frantic commands. By 4:15 AM, things are stable. You are exhausted, cranky, and desperate for sleep. But before you can close your laptop, there is a nagging ticket looming over you: the Post-Mortem, or Root Cause Analysis (RCA).

Enter the modern temptation. You open your terminal, copy the messy Slack triage history and raw PagerDuty logs, paste them into Claude or OpenAI's GPT-4o, and write a prompt: "Summarise this incident timeline and write a blameless post-mortem."

Within ten seconds, you have a beautifully formatted, professional-sounding markdown document. It looks perfect. It reads like a seasoned engineering lead wrote it. You copy-paste it into Confluence, tag the stakeholders, and go to bed.

It feels like a massive win. But you have just committed a quiet, insidious architectural error—not in your codebase, but in the cultural machinery that makes your engineering organisation tick.

Here is why outsourcing your post-mortems to an LLM is a terrible idea, and how to use AI during incident reviews without destroying your team's psychological safety.

LLMs Seek Narrative Cohesion over Messy Truth

By their very nature, Large Language Models are designed to predict the most plausible next token. They are masters of narrative coherence. They hate loose ends, contradictory human memories, and unexplained gaps in data.

But real production incidents are messy, confusing, and full of cognitive biases. When an outage occurs, three different engineers might have three completely different mental models of how the system was behaving. The true value of a post-mortem is not the final document; it is the friction-filled debate that happens when those three engineers sit in a room and reconstruct the timeline together.

When you ask an LLM to synthesise an incident, it automatically smoothens the rough edges. It resolves contradictions by halluxinating a logical path of cause and effect that might not have existed. It turns a chaotic, non-linear system failure into a neat, linear story. In doing so, you lose the exact anomalies and weird human workarounds that you need to understand to prevent the next, even bigger failure.

The Sanitisation Trap: Erasing the 'Why'

If you ask an LLM to make a post-mortem "blameless," it will often do so by sanitising the human elements out of the report entirely. It replaces human confusion with clean abstractions.

For example, instead of writing: "Dave ran the migration script because he thought the staging flag was set, but the CLI tool's help menu is incredibly misleading," an LLM will output: "An operator executed the database migration on the incorrect environment due to configuration ambiguity."

This looks clean, but it is useless. By erasing Dave’s actual cognitive context—the misleading CLI tool—you miss the real systemic fix. The fix isn't "train operators to be more careful"; the fix is redesigning the CLI's interface so it is physically impossible to run a production migration without explicit, double-key confirmation.

When we let AI write our post-mortems, we replace deep human empathy and systems-thinking with clinical, corporate boilerplate. We end up addressing the symptoms of the outage rather than the underlying sociotechnical system.

The Ethical Hazard of Automated Accountability

Post-mortems rely entirely on psychological safety. If your team believes that their raw, honest Slack messages during a high-pressure crisis will be fed into an LLM to generate a performance-trackable summary, they will change how they behave during an incident.

They will stop admitting mistakes in public channels. They will stop proposing wild, out-of-the-box hypotheses. They will start performing for the machine, ensuring their written communication is hyper-polished and defensive while production is burning.

Using AI to automate these summaries signals to your team that management values speed and bureaucracy over genuine, painful learning. If you are having trouble setting up your workflows or finding the right balance with your prompts, check out our debugging guides over at the OpenAI hub for tips on structuring raw data, but keep the core narrative human-authored.

How to Actually Use LLMs Safely in an Incident Review

This does not mean AI has no place in your incident response toolkit. You can use LLMs to do the heavy lifting of raw data prep, leaving the human-centric narrative to your team. Here is the ethical, highly effective way to split the labour:

  1. Use LLMs as Timeline Parsers, Not Authors: Feed your raw logs, JSON payloads, and timestamps into the model. Ask it to do one specific job: "Convert these unstructured timestamps and log lines into a chronological table of events." Do not ask it to draw conclusions, assign cause, or write the summary.
  2. Run Human-Led 'Why' Sessions: Use the AI-generated timeline as a starting point, but gather your team to write the analysis. Ask: "What did we think was happening at 03:14? Why did that query seem like the culprit?"
  3. Draft Action Items Manually: Never let an LLM generate your remediation tasks. They will almost always suggest generic, low-value work like "add more logging" or "write more tests." True systemic fixes require deep context that only your senior engineers possess.

By keeping the analysis human-led, you preserve the trust and vulnerability that makes a blameless culture work. The paper trail is cheap; the collective learning is priceless. Don't automate away the only part of an outage that actually makes your team smarter.

ethicsengineeringworkflowsculture

Keep going

Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.