Ethics & Responsible Use
Why You Shouldn't Use AI to Screen Candidate CVs (And the Ethical Way to Hire Engineers with LLMs)
Feeding a stack of PDFs to an LLM to find your next senior engineer is tempting, lazy, and fundamentally broken. Here is why automated CV sifting fails ethical and technical tests, and how to use models responsibly instead.
Updated 10/5/2026
We have all been there. You post a remote-friendly role for a Senior Backend Engineer, and within forty-eight hours, your ATS is buckling under the weight of six hundred applications. Your calendar is already full, your coffee is cold, and the temptation to write a quick Python script that feeds those PDFs into the Claude API with a prompt like "Rank these candidates based on technical fit" is incredibly strong.
It feels like a modern solution to a modern headache. But doing this is a shortcut that is both ethically compromised and technically short-sighted.
Let’s be entirely honest: using LLMs to make gatekeeping decisions about human careers is a bad look. It is also an incredibly inefficient way to find genuinely great software engineers. Here is a breakdown of why raw LLM screening is broken, what makes these automated systems tick (and fail), and how you can actually use language models to build a fair, efficient, and ethical technical hiring pipeline.
The Fallacy of the "Standard" Engineer
To understand why LLMs make terrible CV filters, you have to look at what they are trained to do. Large language models are statistical prediction engines. They are trained to identify, replicate, and predict the most likely next word in a sequence based on vast amounts of historical data.
When you ask an LLM to evaluate a resume, you are asking it to compare that resume against its internal model of what a "successful engineer" looks like. And where did it get that model? From the internet. Specifically, from decades of historical resumes, corporate bio pages, and tech blog posts.
This creates two major ethical and practical problems:
- Systemic Bias Amplification: Historical hiring data is not neutral. It is heavily skewed toward specific demographics, universities, and pedigree employers. When you ask OpenAI's GPT-4o to filter resumes, it will unconsciously reward candidates who look like the industry's historical average. It doesn't know why a candidate from an elite university is a safe bet; it just knows that sequence of characters correlates with corporate success in its training set.
- The Death of the Brilliant Oddball: The best hires are rarely the ones who followed a perfectly linear path. They are the self-taught developers, the career-switchers, or the engineers who spent three years building a failed indie game before returning to enterprise SaaS. Because LLMs operate on statistical averages, they actively penalise non-traditional profiles. They seek conformity because conformity is mathematically predictable.
The "Aesthetic" Trap: How LLMs Evaluate Formatting Over Substance
If you have spent any time in our Prompt Generator, you know that LLMs are highly sensitive to style, structure, and readability. This sensitivity becomes a major liability when parsing CVs.
An LLM does not have a concept of lived experience. It cannot read between the lines of a poorly formatted but technically brilliant resume. Instead, it rewards polish. A candidate who used an expensive resume template, optimized their keywords for ATS compatibility, and used clean, structured bullet points will score significantly higher than a brilliant systems engineer who threw their career history onto a plain text document.
By relying on LLMs to rank candidates, you aren’t hiring the best engineers. You are hiring the people who are best at writing LLM-friendly markdown.
The Ethical Alternative: Objective Feature Extraction
So, do we have to ban LLMs from the hiring pipeline entirely? Absolutely not. The key to using AI ethically in recruitment is shifting the model's role from evaluation to synthesis.
Instead of asking an LLM to make a subjective judgment ("Is this candidate good?"), you should use it to extract objective, structured data to save you manual reading time. This keeps the human in control of the actual decision-making process.
Here is how an ethical extraction pipeline works:
- Ask for Facts, Not Opinions: Do not ask the LLM to grade the resume. Instead, use a structured schema (via tools like Pydantic) to extract raw data. For example: "Does this resume mention production experience with Kubernetes? [Yes / No / Unmentioned]".
- Anonymise on the Fly: Before you even look at the candidates, use a lightweight script to strip out names, locations, graduation years, and gendered language from the extracted data. This forces your human review team to focus purely on the technical output.
- Surface Relatable Scale, Not Names: Ask the LLM to categorise the scale of systems the candidate has worked on (e.g., "handled migrations of >1TB databases" or "managed teams of 5-10 people"). This helps you quickly filter for scale alignment without being blinded by prestige company logos.
How to Build a Fair Technical Assessment Pipeline
If you want to design a pipeline that actually respects candidates and surface real talent, stop trying to automate the initial filter. Instead, use LLMs to help you design better, fairer technical assessments further down the funnel.
For example, rather than using generic LeetCode puzzles (which are easily solved by LLMs anyway), use models to generate realistic, domain-specific code review challenges. You can check out our guides on building resilient code analysis workflows over at our Claude Articles Hub to see how to structure these prompts.
When you use AI to help you build better human-centric tests, rather than using it to avoid looking at humans entirely, you build a hiring pipeline that candidates will actually respect. You save time, you avoid bias, and most importantly, you keep your integrity intact.
Keep going
Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.