AI technical screening, and what the first round has to measure now
September 8, 2026

The first round was never about knowledge. It was about cost.
Somewhere between the application pile and the onsite loop, a company has to remove the people who cannot do the work, and it has to do that before a senior engineer spends four hours finding out. That is the whole job. The phone screen, the shared editor, the forty-five minute algorithm question: every convention of the technical screen is a cost-control device wearing the costume of an assessment.
It worked while its inputs were expensive to produce. Recall was a decent proxy for experience, because remembering how a rolling deploy handles a failed health check, live, with someone listening, correlated with having watched one fail. The correlation was never causal. It was just costly enough to be useful.
Recall costs nothing now, and a filter whose input costs nothing stops filtering. We have written before about how that collapse shows up inside a screening funnel. This is about what to build in its place. The expense the first round existed to prevent did not disappear when the filter stopped working. It moved into the onsite loop, which is the most costly hour in the building.
Start by measuring it on your own funnel
Before accepting any of the above, check whether it is true where you work. Two numbers, both already sitting in your ATS:
- Your first-round pass rate, quarter by quarter, for the last two years.
- Your onsite-to-offer rate over the same quarters.
If the first has climbed while the second has fallen, your screen has stopped removing people and started forwarding them. If both lines are flat, none of this is your problem yet and you can stop reading. Where they have diverged, the divergence is the cost, and it is being paid by whoever runs your onsite loop.
There is no industry benchmark to compare against here, and anyone offering you one is guessing. The honest version of this measurement is longitudinal against yourself. Two years of your own data beats a published average you cannot audit.
Recall stopped being evidence
Seventy-six per cent of applicants use AI to tailor their applications. Eighty-one per cent of hiring managers believe they can identify an AI-generated application. Both figures come from the same 2026 survey of roughly 60,000 people, with an India sample of 3,552.
The first number is unremarkable. Candidates picked up a tool that was placed in front of them, which is what candidates have always done with tools. The second is the one worth sitting with, because it is a confidence claim and nothing in that survey tests it.
Grant it anyway. Assume the managers are right and they catch most of what is generated. The screen still fails, because detection is aimed at the wrong target. A tailored application is not a lie. A fluent, well-structured answer about Kubernetes upgrade strategy, from someone who has only ever read about Kubernetes upgrade strategy, is not a lie either. Nothing has been fabricated. The candidate knows the answer. Knowing the answer simply stopped carrying information.
Which is why the industry reflex, more detection and tighter proctoring, solves a problem adjacent to the real one. You can catch every candidate who pastes from a second tab and still have no idea which of the remaining ones has done the job.
Volume is the other half of the problem
Recall collapsing would be manageable if the pile were small. It is not, and in India it is nowhere near small.
Eighty-two per cent of Indian employers report difficulty filling roles against a global figure of 72 per cent, and the hardest skill to find is AI model and application development at 39 per cent, per ManpowerGroup's 2026 talent shortage survey with an India sample of 3,051. More applications arriving, and fewer verifiable candidates inside them, is what turns an imprecise first round into an expensive one. We have argued separately that the shortage is a screening problem before it is a supply problem.
At low volume you can absorb a weak first round by simply interviewing more people. At a few thousand applications against one listing, the first round is the only thing standing between an engineering org and a calendar full of interviews that were never worth having. The precision of the filter matters in proportion to the size of the pile.
Say versus do
One signal survives all of this, and it requires detecting nothing.
Ask a candidate to explain something. Then ask them to do a small piece of that same thing. Compare the two.
The comparison is the measurement, and it works because the halves fail in opposite directions. People who have operated a system often explain it badly. They hedge, they qualify, they get lost in a particular incident from 2023 that turns out not to be representative. Then you put them in front of the task and they move. People who have only studied a system explain it cleanly, in ordered points, trade-offs named. Then the task starts and there is nowhere to begin.
Either half alone will mislead you. The fluent explainer reads as strong in a conversation-only screen; the halting operator reads as weak. Held together, the gap between what someone says and what they do is the most reliable thing a first round can produce, and it is the one measurement that does not care whether a model helped with the explanation.
You can run it manually, this week, in forty-five minutes:
- Pick one claim from the résumé that the role actually depends on. One, not five.
- Spend fifteen minutes probing it conversationally, descending rather than moving sideways. Each question takes a noun from the previous answer and asks what it cost.
- Spend twenty minutes on a small task in the same territory. Not a puzzle. Something the job contains: read this config and tell me what breaks, sketch the pipeline for this release, work out why this rollout stalled.
- Spend the last ten on why they did what they did.
- Write down two ratings rather than one. How well they explained, and how well they performed. Look at the difference before you look at either number.
The honest problem with this is that it is expensive. Forty-five minutes of senior attention per candidate is precisely the cost the first round was invented to avoid, which is why almost nobody does it at volume. That tension is the actual state of first-round screening right now, and it does not get resolved by deciding to care more.
What replaces recall
Two things, load-bearing in different ways.
Demonstrated work. Something the candidate builds, debugs or designs, where an artifact exists at the end that either holds up or does not. The argument for it is structural rather than empirical: a task with an artifact is harder to fake than a question with an answer, because the artifact has to survive inspection while the answer only has to survive the next thirty seconds. That is a claim about the shape of the two things, not a measured result, and it should be read as one. A hands-on task also surfaces the skill almost nobody screens for well yet, which is how someone works with an assistant rather than around it.
Probed history. Depth-testing the claims already on the page, one layer below where the candidate prepared. This is cheaper than it sounds, and most interviewers do it badly, not because it is difficult but because it requires descending instead of moving on. The probe table is the mechanical version: a claim, then a chain of follow-ups only a participant can answer.
Neither leg stands alone. Demonstrated work without history-probing tells you what someone can do in an hour under observation, which is not the same as what they have done. History-probing without demonstrated work is a conversation, and conversations are exactly what stopped being evidence.
Standardised where it counts, adaptive where it helps
The fairness objection arrives immediately. If the interview adapts to the candidate, are you still comparing like with like?
Yes, provided you are careful about which layer adapts. Hold the competencies constant. Hold the difficulty constant for a given role and level. Hold the rubric constant. Let the surface questions move. Every candidate is then measured against the same things at the same depth, and only the route through differs.
Identical question lists have two failure modes a standardised rubric does not. They leak, and once leaked they measure preparation. And they force the same question on a candidate whose experience is adjacent rather than matching, producing a low score that means nothing. The rubric is where the fairness lives. The questions are just how you get there.
This is also the more defensible position when somebody asks you to justify a decision six months later, which they will. Same competencies, same bar, evidence attached to each is an answer. Everyone got the same list stops being an answer the moment the list is on a forum.
The eight dimensions worth holding constant across a first round is where this gets concrete. It is the framework the rest of this cluster references.
A band is not a difficulty setting
The most common way a structured screen still produces noise is running one interview and turning a dial for seniority.
A junior and a staff engineer are not the same interview at two difficulties. They are different interviews, because what you weight changes. Technical recall carries most of the score early and almost none of it late. Architecture appears only once there is enough scope behind someone for it to mean anything. The ability to make a defensible trade-off is worth little at two years, since nobody has let them make one, and becomes most of the assessment at ten.
What actually changes across four experience bands covers the general shape. For the worked version, with the weightings and the reasoning behind every shift, the DevOps guide runs the whole structure end to end for one role family.
Evidence is the deliverable
The output of a first round is not a score. It is a case.
A number nobody can explain is worse than no number, because it launders a judgement into something that resembles a measurement and then resists being argued with. If a screen returns 72, the only useful follow-up is which moments produced the 72, and a system that cannot answer that has handed you a verdict dressed as data.
This matters more each quarter for reasons unrelated to quality. Article 50 of the EU AI Act has been in force since 2 August 2026: people must be told, clearly and from the start, when they are interacting with an AI system. The heavier obligations covering recruitment did not arrive with it. Annex III was deferred to 2 December 2027, and most pages currently ranking on this have the two dates the wrong way round. The disclosure duty that actually binds an AI interview today is live. The one everyone wrote about is sixteen months out.
The through-line is that a decision has to be attributable to a person and supportable with evidence. That is where regulation is converging, and it happens to be what a hiring manager wanted anyway. Nobody has ever wanted a score. They want to know what the candidate did and what it suggests, so that they can decide.
Which is the constraint we are building the interview engine around: the system assembles the evidence and makes the case, and a person makes the call.
This is heavier, not lighter
None of the above is a lighter process than what it replaces. Probing a claim properly takes longer than asking a question with an answer. Setting a real task takes preparation. Writing down evidence takes longer than writing down a number.
The reason to do it anyway is that the alternative has stopped working, and the cost of a first round that forwards everyone is not paid in the first round. It is paid four interviews later, by the people whose time you least want to spend.
Hiring for one of these roles right now?
Recio runs the first round for cloud, DevOps and AI roles and hands back the evidence behind every score. Bring one live role and we build the labs around it.