How to build an interview scorecard with behavioral anchors
A scorecard is only as good as its anchors. How to define a competency, write a 1 to 4 scale a panel reads the same way, and tie every score to evidence.
A scorecard turns an interview into evidence a panel can compare. Pick four or five competencies, write behavioral anchors so a 3 means the same thing to everyone, tie each score to a quote and a timecode, and confirm the AI draft rather than trusting it.
A scorecard is what turns an interview into evidence a panel can compare. Its weakest part is usually the rating scale: numbers without meaning, filled in from memory. Behavioral anchors fix that. They describe what a 1 and a 4 look like, so that a score means the same thing to every interviewer. This is how to build one.
Start from competencies, not questions
A scorecard rates a small set of competencies, not a list of questions. Pick four or five for the role, no more. For a software engineer that might be system design, coding, ownership and communication. Each competency needs a one-line definition of what you are assessing, so the panel rates the same thing.
Decades of research on the structured interview point the same way. Campion, Palmer and Campion (1997) identify fifteen components of structure; the ones that shape a scorecard are anchored rating scales, rating each answer, taking detailed notes, training, and combining scores by a rule rather than by impression. Sackett, Zhang, Berry and Lievens (2022) re-ranked selection methods and put the structured interview first, with the highest mean operational validity (.42).
Write behavioral anchors
An anchor is a sentence that describes behavior at a point on the scale. A 1 to 4 scale needs an anchor at each level, phrased as what the candidate does, not how they made you feel. For ownership:
- 4: Spots the problem, owns the fix and the follow-up without being asked.
- 3: Takes an assigned fix through to the end.
- 2: Does their part; the follow-up waits for others.
- 1: Describes what the team did, not their own part.
Good anchors are observable, mutually exclusive and about the work. Avoid words like "strong" or "excellent" that only restate the number. The method goes back to Smith and Kendall (1963), who built rating scales from behaviors that raters agreed on, so two people watching the same answer land on the same score.
Tie every score to evidence
A score with no evidence is an impression. Each score should point to a moment: a quote and a timecode from the interview. This is where notes matter. If the scorecard is filled in three hours later from memory, it records the impression and loses the evidence.
In Vettasy the AI drafts the scorecard from the transcript, with a quote and a timecode for every proposed score. The interviewer confirms or corrects each line. Nothing unconfirmed counts, and nothing unconfirmed reaches the rest of the panel. The result is a scorecard where every number points to a place in the recording, and the decision log records who decided and on what evidence.
Keep it consistent across the panel
The same scorecard, with the same anchors, goes to everyone who interviews for the role. If the plan changes after the first interview has run, record the change, so a later comparison knows not everyone got the same version. A consistency index shows where a panel drifted: which candidates were not asked about a competency, and where the rubric was applied differently.
Google's re:Work guide reaches the same conclusion from practice: structured tools it calls rubrics are "more predictive than unstructured interviews", and reusable questions and rubrics save about 40 minutes per interview.
A short checklist
- Four or five competencies, each with a one-line definition.
- A 1 to 4 scale with a behavioral anchor at every level.
- Anchors that describe observable behavior, mutually exclusive, about the work.
- A quote and a timecode behind every score.
- The same scorecard for every candidate for the role, changes recorded.
- Scores combined by a rule, and a decision log that holds the evidence.
Sources
- Campion, M. A., Palmer, D. K. and Campion, J. E. (1997). A review of structure in the selection interview. Personnel Psychology, 50(3), 655 to 702. doi.org/10.1111/j.1744-6570.1997.tb00709.x.
- Levashina, J., Hartwell, C. J., Morgeson, F. P. and Campion, M. A. (2014). The structured employment interview: narrative and quantitative review of the research literature. Personnel Psychology, 67(1), 241 to 293. doi.org/10.1111/peps.12052.
- Smith, P. C. and Kendall, L. M. (1963). Retranslation of expectations: an approach to the construction of unambiguous anchors for rating scales. Journal of Applied Psychology, 47(2), 149 to 155. doi.org/10.1037/h0047060.
- Sackett, P. R., Zhang, C., Berry, C. M. and Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection. Journal of Applied Psychology, 107(11), 2040 to 2068. doi.org/10.1037/apl0000994.
- Google re:Work, A guide to structured interviewing.
- Vettasy, For employers and AI Notice.
Written by the Vettasy teamVettasy is a desktop application for job interviews. Employers run structured, recorded interviews in it with verified participants. It is free for candidates. Windows and macOS.