First-round screening
The eight things worth scoring in a first round
August 15, 2026

The feedback box after a first round usually gets three sentences. Strong technically. Communicates well. Would move forward. Two weeks later somebody asks why this candidate and not the other one, and those three sentences cannot answer the question. Nobody was careless. The interview simply never decided what it was measuring, so at the end there was nothing to compare.
That is the actual failure, and it is structural rather than a matter of rigour. Evaluation criteria chosen after the conversation, out of whatever the conversation happened to surface, are not criteria. When the dimensions get picked in retrospect you are not evaluating the candidate. You are summarising them.
Fix the dimensions before the candidate speaks
Two candidates, two interviews, two implicit rubrics assembled in retrospect from two different conversations. The comparison at the end is not a comparison. It is a preference, expressed in the vocabulary of assessment.
Deciding in advance costs you something real, and it is worth naming: you lose the pleasant discovery that a candidate turns out to be excellent at something you weren't looking for. That happens, and it is genuinely useful information. Take the trade anyway. The discovery is available in the second round, where there is time for it and where a human is deliberately looking. The first round exists to produce a defensible comparison across a lot of people, and it cannot do that on a rubric that changes shape per candidate.
The count matters as much as the commitment. Score three dimensions and everything collapses into "technical" — you lose the ability to tell an engineer who knows from an engineer who has done, which is the distinction the whole exercise is for. Score fifteen and you get a form nobody completes honestly; the last six inherit whatever the fifth got. Eight is roughly what survives an hour with real evidence behind each line.
The eight, and what each one measures
- Technical knowledge — command of the core tools and concepts
- Practical experience — has done the work, not just read about it
- Problem solving — debugging and reasoning under ambiguity
- Decision making — sensible trade-offs, justified
- Architecture thinking — designs for scale, failure and cost
- Communication — explains clearly, structures thinking
- Résumé authenticity — claims hold up under depth-probing
- Production readiness — would you trust them on call
Three of these get merged in practice, and each merge costs something specific.
Technical knowledge and practical experience are the most commonly collapsed pair and the most expensive. One asks whether they know it. The other asks whether they have done it. A candidate can be full marks on the first and near zero on the second, and that exact combination is what a trivia bank is structurally incapable of seeing — it is testing recall, and recall is what the well-prepared candidate arrives holding. Separating the two is most of the reason to run a first round at all.
Problem solving and decision making look adjacent and are not. Problem solving is finding an answer when the situation is ambiguous. Decision making is choosing between two answers that are both defensible, and being able to say why one and not the other. Junior work is mostly the first. Senior work is mostly the second, and an interview that only scores the first will rate a strong senior identically to a strong mid.
Architecture thinking is the only one of the eight that legitimately does not apply at every level. It is not scored at all in the junior bands and becomes the heaviest component at the top, which is part of a broader pattern in how a technical interview should change by experience level.
The two you cannot ask for
Six of the eight can be elicited. You ask something, you get an answer, you score it. Two cannot, and they happen to be among the most decision-relevant on the list.
Production readiness is a conclusion, not an answer. There is no question that returns it. You assemble it from everything else: whether the debugging story had a real first step or started at the fix, whether they mentioned blast radius before they mentioned the solution, whether they can distinguish what they would check at three in the morning from what they would check at three in the afternoon. Ask someone directly whether they can be trusted on call and you will learn only how they talk about themselves.
Résumé authenticity is the other one, and it is the entire reason depth-probing exists as a technique. You never ask whether a claim is real. You drill one layer past the point the candidate prepared for, and the texture of what comes back does the work. Specific and lived-in reads as genuine. Textbook but vague reads as shallow, which is a flag rather than a verdict — the technical and scenario rounds have to compensate. Wrong or self-contradictory reads as fabricated. The probe table method is that technique written down so it survives being handed to somebody else.
Both of these are inferences. That is worth stating plainly, because a scorecard assembled purely from question-and-answer pairs has two holes in it exactly where the expensive mistakes live.
Weighting components is not the same as scoring dimensions
This is the distinction that separates a scorecard that holds up from one that degrades into a single number, and it gets missed almost universally.
There are two axes, not one.
The components are what you run: résumé verification, technical assessment, scenario work, hands-on labs, behavioural. These carry the weights. They sum to a hundred, and the split shifts with seniority.
The dimensions are what you report: the eight above.
They are not the same list, and a single component feeds several dimensions at once. Walking a candidate through a production scenario produces evidence on problem solving, decision making, communication and production readiness simultaneously. A hands-on lab produces practical experience and problem solving, and at senior levels architecture thinking as well. One activity, four scores.
Collapse the axes and the report becomes a description of which part of the interview went well, rather than a description of the candidate. The hiring manager reads "scored well on the technical assessment" and learns nothing about whether this person can be trusted with the system, because the technical assessment was never measuring that.
The part that has to be made explicit is the mapping — which component contributes evidence to which dimension. If it stays implicit, two interviewers will map it two different ways, and the reports stop being comparable across candidates, which was the point. Write the mapping down once, per role. It is a twenty-minute exercise that you do a single time.
Why an average is the wrong last step
Eight weighted dimensions produce a score out of a hundred, and a score maps to a recommendation: strong hire above 85, hire from 70 to 84, borderline from 55 to 69, no hire below that, with the cutoffs tuned to how high a given company sets its bar.
Then you stop, because an average hides two failure modes and both are worse than the errors it catches.
The first is dishonesty. If authenticity is only a weighted line, a strong enough performance everywhere else buys it back. A candidate who invented a credential and scored ninety on the remaining seven dimensions averages out to a hire, and the arithmetic quietly resolves something no interviewer would resolve that way in conversation. An authenticity gate refuses the trade: fabrication caps the recommendation at borderline and flags the case for human review no matter what the total says.
The second is the missing core skill — the great-on-paper hire who cannot do the central thing the role exists for. If the requisition is genuinely orchestration-heavy and the candidate's orchestration score is below threshold, a critical-skill floor caps at borderline even when the weighted total is strong. Seven good dimensions do not compensate for the one the job is actually about.
Both gates cap and flag. Neither rejects anybody. That distinction matters more than it might appear: the purpose of a gate is not to automate a decision, it is to stop an average from making one silently. The case still goes to a person, with the flag attached and the evidence underneath it.
A scorecard you can defend
Five properties, and a first round that has them will survive the question two weeks later.
The dimensions are fixed before the interview and identical across every candidate at that level. Each dimension carries a score with a specific moment behind it, so the number is traceable to something that was said rather than free-floating. The weights are written down and specific to the seniority band. Gates are kept separate from scores, so honesty and the one indispensable skill cannot be averaged away. And a person makes the call at the end, with all of the above in front of them.
That is what Recio's interview engine is built to produce — per-dimension scores tied to transcript evidence, the gates applied on top, and a recommendation a hiring manager can walk through line by line rather than defend as a verdict.
The three-sentence feedback box is not a discipline problem. Interviewers write it because it is the honest summary of an interview that never committed to what it was looking for. Commit first, and the box writes itself.