First-round screening

Four seniorities, four different interviews

Four seniorities

Most first rounds are one interview wearing four different job titles. The same opening, the same three technical topics, the same production scenario — handed to a candidate eighteen months into their career and to one who has carried a pager for nine years. The senior gives a fuller answer. They score higher. Everyone agrees the process worked.

It didn't. Both candidates were measured on one axis, and the person with more years had more to say on it. That is a fluency test. It sorts people who have read more and talked more, which tracks seniority closely enough to feel like it's working, and fails precisely where the cost is highest: the articulate senior who has never actually owned anything.

Conducting a technical interview well across experience levels is not a matter of making the questions harder as the years go up. It means changing what the interview is for at each level. Below is what that looks like when you write the weights down — the calibration Recio uses across four experience bands, and the reasoning behind each shift. The reasoning is the part worth stealing.

What the weights actually do across four bands

Every interview scores the same set of components: résumé verification, technical assessment, scenario work, hands-on labs, behavioural, and — above a certain level — architecture and system design. What moves between bands is how much each one is worth. These are the reference weights for a generalist DevOps role.

0–2 years

  • Technical assessment — 40%
  • Interactive labs — 20%, one lab
  • Scenario-based assessment — 15%
  • Behavioural — 15%
  • Résumé verification — 10%
  • Architecture and system design — not scored

2–5 years

  • Technical assessment — 25%
  • Scenario-based assessment — 25%
  • Interactive labs — 25%, two labs
  • Behavioural — 15%
  • Résumé verification — 10%
  • Architecture and system design — not scored

5–8 years

  • Scenario-based assessment — 25%
  • Interactive labs — 25%, three labs
  • Technical assessment — 15%
  • Architecture and system design — 15%
  • Behavioural — 15%
  • Résumé verification — 5%

8–12 years

  • Architecture and system design — 25%
  • Interactive labs — 25%, four labs
  • Scenario-based assessment — 20%
  • Behavioural — 15%
  • Technical assessment — 10%
  • Résumé verification — 5%

Two lines carry the whole argument. Technical recall runs 40 to 25 to 15 to 10. Architecture starts at nothing, appears at 5–8, and ends up the single heaviest component. Somewhere in the middle those lines cross, and the interview stops being about what someone knows and starts being about what they would choose.

At 0–2 the emphasis on recall isn't laziness, it's honesty about what exists to measure. A candidate two years in has no production history to interrogate and no architectural decisions they owned. What they have is fundamentals, and fundamentals are worth 40% because at that level fundamentals are the job. The one lab is there to check that the knowledge survives contact with a terminal.

At 8–12 recall is the least interesting thing about the candidate. Whether they remember a specific flag is a search away. What you cannot look up is whether they will make a defensible call on a system nobody has built yet, under a cost constraint, knowing it has to survive a region going dark. That is why architecture takes a quarter of the score, and why scenarios — which at 2–5 were about diagnosing a failure — become at 8–12 about designing so the failure is survivable.

The middle bands are where scenarios dominate, and that is deliberate. A 2–5 engineer is past fundamentals and short of system ownership. The most informative thing you can do is put them in front of a production problem with more than one plausible move and watch which one they take.

The two things that don't move

Two components hold steady across all four bands, and both are more interesting than the ones that shift.

Behavioural stays at 15% at every level. How someone behaves during an incident, what they take ownership of, how they work with the people around them — that matters identically whether they have two years or twelve. What changes is not the weight but the rubric. A junior who escalates early is doing the right thing. A senior who escalates early on something they should have owned is telling you something else entirely. Same weight, same question territory, a different bar.

Labs hold roughly a quarter of the score from 2–5 upward, but the count scales one, two, three, four. The weight is flat and the evidence volume is not. That is on purpose. A single lab can go badly for reasons that have nothing to do with the candidate — an unfamiliar interface, a bad twenty minutes, a misread brief. At junior level one sample against 40% of recall-based scoring is a reasonable balance. At senior level, where the hands-on work is carrying far more of the decision, one sample is too thin to hang a hire on, so you take four.

The labs themselves also change character. At the junior and mid bands they are build exercises: make the thing work. At the senior bands they shift to design and decide, where the artifact matters less than the trade-offs made to produce it.

The same topics, at four depths

The list of topics barely changes between bands. The expected depth changes on nearly every line. Most question sets get this backwards — they swap the topics and hold the depth constant.

A calibrated interview runs a ladder instead: awareness, beginner, intermediate, advanced, strategic. Linux and troubleshooting sit at intermediate for the first two bands and advanced for the last two. CI/CD runs intermediate, advanced, advanced, strategic. Networking starts at beginner and ends at advanced. Nothing drops off the list. What counts as a complete answer moves every time, which puts the weight on the interviewer's ability to tell depth from fluency in an answer they cannot verify.

Orchestration is the clearest example of the shape. At 0–2 you are checking that they can tell the objects apart and know what runs where. At 2–5 you want the day-two mechanics: how the platform decides an instance is healthy, and what breaks when that judgement is configured wrong. At 5–8 the interesting territory is behaviour under resource pressure and what a live upgrade actually costs. At 8–12 it is design across clusters and regions, and whether they can articulate what that buys and what it breaks.

A few topics genuinely start at zero and stay there for a while. Architecture and system design isn't scored at all in the junior bands. Technical strategy and leadership doesn't appear until 5–8. Operational database work shows up at 2–5 and never climbs past intermediate, because for this role it isn't the job. Deciding which topics are absent at a band is as much of the calibration as deciding which are deep.

For seniors, honesty is a gate rather than a score

Résumé verification is worth 10% at the junior bands and 5% at the senior ones. Read quickly, that looks backwards. Senior candidates have longer résumés, more to overstate, and more incentive.

The weighting is lower because at senior level authenticity stops being something you score and becomes something you gate.

At 0–2 the claims are small and checking them is genuinely informative. Was the project real work or a course exercise? Is the certification a weekend or a year? That evidence is worth putting on the scoreboard, because there isn't much else to weigh it against.

At senior level the same finding has a different consequence. A fabricated claim doesn't cost a candidate five points. It caps the recommendation at borderline and flags the case for a human to look at, regardless of how the rest of the interview went. Severe or repeated fabrication ends it. Seniors are gated for honesty, not rewarded for it.

The distinction is not academic. If authenticity is only a weighted line, a strong enough performance everywhere else buys it back — a candidate who invented one credential and scored ninety on the rest averages out to a hire, and the arithmetic quietly makes a decision nobody would defend out loud. A gate refuses the trade. It doesn't reject anyone; it stops the average from settling something a person should settle. That principle applies well beyond résumés, and it is the same reason everyone passing your screen is a scoring problem before it is a candidate problem.

What to change before your next first round

Four things, none of which require tooling.

Write the weights down before the call, not after. If you cannot say what fraction of the decision is recall, you will default to whichever component the candidate performed best on, and you will do it without noticing.

Find the two or three components that move most between the band you are hiring for and the band below it. Those are your interview. For a 5–8 hire against a 2–5 baseline, that is architecture appearing and recall halving. Spend the time there.

Separate what you gate from what you score. At minimum: honesty, and the one skill the role genuinely cannot be done without. Everything else can average.

Hold the rubric across every candidate in a band. The surface questions can vary — they should, since fixed questions leak. The bar cannot.

None of this makes an interview longer. It makes the same forty minutes point at the thing you are actually trying to find out, which at 8–12 years is a different thing than it was at 18 months. Recio's interview engine applies this calibration automatically and returns the evidence behind each score, but the calibration itself is the useful part, and it works on a whiteboard.

The senior who aced your junior interview didn't tell you they were senior. They told you they were fluent. Those are not the same finding, and only one of them is worth a second round.

← All posts