Tickd.ai
← The Tickd Guide

Ethics & Responsible Use

Why You Shouldn't Use LLMs to Write Performance Reviews for Your Engineering Team

Staring at a blank page during review season? Before you paste your developer's commit history into Claude, learn why AI-generated appraisals destroy trust and how to ethically use LLMs instead.

Updated 10/5/2026

We’ve all sat staring at a blank document, trying to summarise twelve months of a developer’s professional life into three neat paragraphs. Review season is a gruelling, emotionally draining exercise. It is easy to treat reviews as a bureaucratic exercise—something to tick off your massive to-do list so you can get back to shipping code.

When you are staring down the barrel of writing ten performance appraisals by Friday, prompting a model like /platforms/claude to "write a balanced, constructive performance review based on these bullet points" feels like a lifesaver.

But it is a trap. Relying on Large Language Models (LLMs) to write performance reviews for your engineering team is a shortcut that carries a devastatingly high trust tax. Here is why you should step away from the prompt window, and how to ethically use AI to support your feedback process without losing your humanity.

The Trust Tax: Why Developers Spot AI-Generated Reviews Instantly

Software engineers are trained to spot patterns, edge cases, and anomalies. If you think your team won't recognise the distinct, sanitised prose of an LLM, you are sorely mistaken.

When a developer reads an appraisal containing paragraphs that "delve into" their achievements, celebrate their "multifaceted approach to problem-solving," or serve as a "testament to their dedication," their internal alarm bells go off. They immediately realise that their manager didn't care enough to write their review themselves.

Performance reviews are a primary currency of career growth, compensation, and recognition. Receiving feedback that has been outsourced to a machine feels deeply alienating. It signals that you value their real-world contributions less than the fifteen minutes of intellectual effort required to write a genuine critique. Once a developer loses trust in the authenticity of your feedback, repairing that relationship is incredibly difficult.

The Erasure of "Glue Work"

LLMs are predictive text engines; they thrive on conventional structures and high-frequency patterns. If you feed an LLM a list of JIRA tickets or commit logs, it will write a review that focuses entirely on the loudest, most quantifiable metrics.

This leads to the systematic erasure of "glue work"—the vital, non-quantifiable tasks that keep engineering teams functional. Glue work includes things like helping junior devs debug a tricky state machine, writing internal documentation, smoothing over team friction, and quietly unblocking cross-functional projects.

An LLM cannot read between the lines of your notes. It will sanitise your feedback, flattening the unique, weird, brilliant quirks of your best developers into a generic template of corporate competency. Your quiet high-performers, the ones who do the heavy lifting in the background, will see their specific impact diluted into generic corporate jargon.

The Hidden Hazard of Unconscious Bias Amplification

It is well-documented that human performance reviews are riddled with unconscious bias. We tend to judge men on their potential and women on their past performance; we use different language to describe assertiveness depending on a developer's gender or background.

If you feed raw, unrefined notes into an LLM, the model will not magically clean up your biases. Instead, it acts as an amplifier, smoothing your half-formed, biased thoughts into highly polished, authoritative-sounding prose. It gives your subjective, potentially flawed assumptions the veneer of objective truth. Because the output sounds so professional, you are less likely to question whether your initial assessment was fair. For deep dives into how LLM behaviours can slide into unexpected patterns, checking troubleshooting guides over on /platforms/openai/articles can show just how much human oversight is required to keep models aligned.

The Ethical Way to Use AI in Your Review Process

This does not mean you must completely banish AI from your workflow. Writing is hard, and LLMs are excellent brainstorming partners. The boundary between ethical assistance and unethical outsourcing lies in who is doing the thinking.

Here is how to ethically partner with an LLM during review season:

1. Use AI to Audit Your Own Bias, Not to Write the Copy Instead of asking the LLM to write the review, write the raw draft yourself. Then, paste your draft into the LLM with a highly specific prompt designed to challenge your perspective.

Try a prompt like this (which you can build out further in our /prompts builder):

> "I have written a draft performance review for a software engineer on my team. I want you to act as an objective, unbiased HR consultant. Analyse my draft specifically for unconscious bias, gendered language, or assumptions that lack concrete evidence. Highlight these areas and explain why they might be problematic, but do not rewrite the review for me."

This keeps you in the driver's seat while using the LLM's analytical capabilities to make you a fairer manager.

2. Use AI to Help You Structure Tough Conversations Delivering constructive criticism is uncomfortable. If you are struggling to frame a difficult piece of feedback, you can use an LLM to brainstorm different structural approaches.

For example, ask the model: "I need to give feedback to a senior engineer who is brilliant technically but constantly derails sprint planning meetings with aggressive debates. What are three different frameworks I can use to structure this conversation during our 1-on-1 so it is collaborative rather than defensive?"

This uses the model to expand your management toolkit, leaving the final delivery and phrasing entirely up to you.

3. Use AI for SPaG (Spelling, Punctuation, and Grammar) Only If you suffer from dyslexia or simply struggle with written English, it is entirely ethical to run your completed, raw drafts through a model to clean up grammar and phrasing. The key is to instruct the model to preserve your exact voice and vocabulary: *"Proofread this performance review draft for spelling and grammar errors. Do not change the vocabulary, do not add corporate jargon, and preserve the casual, direct tone of the original writer."*

Feedback is Human Debt

As a manager, your primary job is to grow and protect your people. When you accept the responsibility of managing a team, you agree to take on the cognitive and emotional labor of evaluating their work.

Using an LLM to generate reviews is a form of emotional debt. It saves you an hour today, but you pay for it in the long run with degraded trust, overlooked contributions, and a team that knows you couldn't be bothered to write their appraisal yourself. Write the review yourself. It might be messy, and it might take time, but your team deserves nothing less.

ethicsmanagementengineeringcareerstrust

Keep going

Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.