Insights
AI in recruitment – and what still requires people
AI shortens the search and the administration. But the step research identifies as the most predictive in any selection process is still a conversation between people.
Written by Jonas Renander
"Human intelligence" is a term borrowed from the intelligence community: information that comes from people, not systems. In recruitment it means roughly the same thing – what you learn through conversations, networks and your own judgement, and what is not found in any database.
The question boards ask us most often right now is not whether we use AI, but where. This is how we divide the work, and why.
What actually predicts performance
The current point of reference is Sackett, Zhang, Berry and Lievens' recalculation of the validity of selection methods in the Journal of Applied Psychology. It corrects Schmidt and Hunter's 1998 summary, long the standard, and lowers most of the values. The figures below show each method's correlation with actual job performance; 1.0 would be perfect accuracy, 0 pure chance.
| Selection method | Validity |
|---|---|
| Structured interview | 0.42 |
| Job knowledge test | 0.40 |
| Biodata (empirically keyed) | 0.38 |
| Work sample test | 0.33 |
| Assessment centre | 0.33 |
| General mental ability test | 0.31 |
| Integrity test | 0.31 |
| Conscientiousness (personality test) | 0.21 |
| Unstructured interview | 0.19 |
The structured interview, then, is more than twice as accurate as the unstructured one. What separates them is how the conversation is set up: the same questions for every candidate, the same assessment criteria, everything written down.
Three caveats, which we would rather state ourselves than have someone else point out.
The 0.42 figure is an average with considerable spread. The researchers state that 80 per cent of values fall between 0.18 and 0.66. It is the execution that determines where in that range you land.
The values concern job performance broadly, not specifically at CEO and board level, where the evidence base is thinner. The direction, however, is stable.
And not every value fell in the recalculation. Biodata is one of the exceptions that was adjusted upwards instead.
What the machine does better than us
Mapping. Working through an entire segment and sorting out people with a certain profile is faster by machine, provided someone checks the output. That is a search problem, not a judgement problem.
Structure. Compiling notes, keeping a process in order and producing a first draft of a role specification.
Consistency. The same questions for every candidate, the same evaluation template. That is exactly what separates the structured interview from the unstructured one. The machine keeps the template in order; the person conducts the conversation.
This saves real time, and we use it. But note what the three points have in common: none of them is a decision.
The three judgements that require a person
1. Whether the person wants it. A list shows who could take the role. It says nothing about who is prepared to leave a company where they have just been given a new mandate, or who stays out of loyalty to a manager. That emerges in a conversation, often only in the second one.
2. Whether the person fits here, specifically. Two CEO roles with an identical role specification can be entirely different assignments depending on ownership, the board's way of working and where the company finds itself. A candidate who is right for a growth company can be wrong for a company in the middle of a transition. That match cannot be made on titles. How deeply it is done is also a large part of the difference between headhunting and executive search.
3. What the references do not say. The most valuable part of a reference call rarely lies in the answers. It lies in the hesitation before the answer, in the phrasing that is discarded and in the subject that never comes up.
The risk of automating the wrong step
A model learns from what has already happened. Feed it ten years of recruitment decisions and it learns who used to get hired, and suggests more of the same. It sounds neutral, because numbers always sound neutral. In practice, it is history being built into the future.
That history is rarely what the assignment brief describes. When a board describes what it needs in its next management team, the answer is almost never "more of the people we already have". A tool that optimises against history will suggest exactly that, and do so with great confidence in its voice.
The mistake is not the technology. The mistake is asking it for a kind of answer it was not built to give.
There is also a regulatory side. AI tools used to select candidates belong to the more heavily regulated use cases within the EU, and the requirements are being tightened in stages. Exactly what applies, and when, has shifted several times. The direction, however, has been the same throughout: a person must be in control of how the tool is used and accountable for the decision, the selection must be explainable after the fact, and the candidate must know how they are being assessed. Parts of this already follow from data protection rules.
We think those principles are reasonable to work by regardless of what the calendar says. Anyone who builds the process that way from the start never has to think about transitional rules. In assignments requiring security vetting, the habit is already ingrained: the material must stand up to review afterwards, by someone other than the person who produced it.
What the candidates themselves think
Pew Research Center asked 11,004 Americans in December 2022. 71 per cent opposed letting AI make the final hiring decision, against 7 per cent in favour. 66 per cent said they would not want to apply for a job with an employer that uses AI to support hiring decisions.
The second figure is the uncomfortable one. Resistance does not stop at AI making the decision. It extends to AI merely helping.
The survey is American and was conducted in December 2022, the same weeks ChatGPT was released. One might have expected attitudes to have softened since, as more people have used AI themselves. Pew's later measurements point the other way.
At executive level, the candidate usually already has a good job, and the threshold for declining a process that feels impersonal is low. That is a risk worth taking seriously.
Five questions worth asking
The questions below are not traps. They are simply the ones that tend to make a process better, whoever you work with – us included. Feel free to bring them to your next meeting.
Has a person seen every candidate before anyone is ruled out? Most mistakes in a recruitment sit among the rejections, and those mistakes are the hardest to detect afterwards.
What is the list based on? A market mapping is prioritised and can be justified name by name. A long list can be good raw material, but it is something else.
Are the interviews structured and documented? A free-flowing conversation often feels nicer. A structured one is more than twice as accurate.
Where in the process is AI used, and does the candidate know? Openness costs nothing here. Trust, once spent, is hard to win back.
Can you tell us why these particular people are on the shortlist? A good answer is about the people: what they have done, what they want and why the role fits them right now.
How OGR works
We use the machine for mapping and structure, and spend the time it frees up on conversations instead. The machine produces the raw material; we make the prioritisation and the assessment. The interviews are structured and documented, because the research says that is where the accuracy sits.
But the methodology is not what we are proudest of. It is that every candidate who reaches your shortlist has spoken to a person at our firm, and that we can tell you who they actually are. What they have done, of course. But also what chafes, and why this particular role would matter right now.
That kind of knowledge only arises in a conversation, and only if someone cares enough to ask one more follow-up question.
A name can be found by machine. A person, you have to meet.
Sources
- Sackett, P. R., Zhang, C., Berry, C. M. & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology. pubmed.ncbi.nlm.nih.gov
- Sackett, P. R., Zhang, C., Berry, C. M. & Lievens, F. (2023). Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors. Industrial and Organizational Psychology, 16, 283–300. cambridge.org — the source of the re-estimated value for assessment centres (0.33) and the spread interval for the structured interview.
- Schmidt, F. L. & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124, 262–274. The summary corrected by the recalculation above.
- Pew Research Center (2023). AI in Hiring and Evaluating Workers: What Americans Think. Survey of 11,004 American adults, 12–18 December 2022. pewresearch.org