AI resume screening: what it is good at, and where it goes wrong
Ranking a stack of résumés against a role is a real use for a model, and a bad one for a keyword filter. What to automate, what to keep human, and the failure modes worth watching for.
· 4 min read · by the PhoneScreen.AI team
Resume screening was automated long before anyone called it AI, and the early version gave the whole category a bad name.
Keyword filters do not read. They match. A candidate who wrote “forklift” ranks above one who wrote “powered industrial truck” for the same ten years of experience, and nobody finds out, because the second résumé was never shown to a person.
What changed is that a model can read the thing. That is a genuinely different tool, and it is worth being precise about what it does well.
What it is good at
Reading around vocabulary. A model handles “reach truck,” “cherry picker,” “order picker” and “sit-down” as the same family of experience without anyone maintaining a synonym list. For hourly roles, where the same job has a dozen names across employers, this alone changes the shortlist.
Applying the same standard to everyone. A person reading two hundred résumés is stricter at the start and looser by the end, or the reverse if they are behind. A model reading two hundred applies whatever standard it was given to all of them, at résumé one and résumé two hundred.
Explaining itself. This is the underrated one. A ranking without reasons is just a number you either trust or ignore. A ranking that says why this person scored where they did is something you can check, argue with, and correct.
Working at a volume nobody was working at anyway. Most high-volume roles are not carefully screened, they are screened until someone runs out of time and then filled from whoever is left at the top of the pile. The comparison is not “model versus careful human,” it is “model versus the bottom of the stack never being opened.”
Where it goes wrong
It rewards good writing. A model reading a résumé is reading a document that somebody wrote about themselves. Candidates who write well look better than candidates who do the job well. This is the single biggest problem with the technique and it does not fully go away.
It infers things it should not. Schools, addresses, employment gaps, names. All of these correlate with protected characteristics, and a model reading a whole document is perfectly capable of picking up on them whether or not you asked it to. Explicit instruction helps. Removing the fields entirely helps more.
It is confident about gaps. A two-year gap gets read as a risk. It is often caregiving, illness, immigration, or a layoff in an industry that shed everyone at once. A model has no way to tell, and it will pick a story.
Calibration drifts with the pool. Score a strong pool and the scores cluster high. Score a weak one and the same candidate scores differently. If you set a hard cutoff and forget it, what that cutoff means changes month to month.
The line worth drawing
Use it to rank and explain, not to reject.
The useful output is an ordered list with reasons attached, so a person opens the top of the pile first and can see why anyone is where they are. The harmful output is an automatic no that nobody ever sees.
That distinction sounds like a nicety. It is the whole difference between a tool that widens how many applicants get a fair look and one that narrows it while feeling efficient.
What to check before you trust the ranking
Read the bottom ten. Not the top. The top will look right, which tells you nothing. If the bottom is full of people who obviously could do the job, your criteria are wrong or the model is reading for the wrong thing.
Check two résumés that should tie. Similar experience, different writing quality. If they are far apart, you are measuring prose.
Look at what it says, not just what it scored. If the reasons are vague, the score is vague. “Strong relevant experience” is not a reason. “Six years on reach trucks, current certification, no gaps” is.
Re-read your criteria out loud. Most bad rankings are bad criteria. “Reliable” is not screenable. “No more than two employers in the last three years” is, if that is what you actually mean, and stating it plainly also makes you notice whether you want to.
Where screening stops
A résumé tells you what someone claims and how well they wrote it down. That is a legitimate first filter and a terrible last one.
Everything that actually predicts the hire comes after: how they answer a question in their own words, what someone who managed them says, whether the history checks out. Résumé screening is useful precisely because it is cheap and shallow, and it goes wrong when it is asked to carry weight it cannot carry.
The resume screener ranks a stack against your role and shows the reason for every score. It is free.
Curious what the candidate hears?
Take the three-question sample screen for a paper sales job at Dunder Mifflin. Your scorecard lands in your inbox.