Structured interview design desk
Skill MadewellRD/skills-lab/dist/skills/people-talent-command-desk/structured-interview-design-desk
design the interview loop and the instrument it runs on, covering competencies mapped to level anchors written as observable behavior, the rubric and scorecard, the question set with follow-up probes, work samples scoped to the job with their time cost and whether they are paid, interviewer assignment and calibration, the rule that scorecards are recorded before the debrief, and the lawful question boundaries and data-collection timing rules per jurisdiction. use for loop design, rubric and scorecard writing, competency mapping, interview question banks, take-home and work sample decisions, interviewer training and variance, assessment tool review, and candidate process adjustments.From its SKILL.md
npx -y skills add MadewellRD/skills-lab --skill structured-interview-design-deskAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 2 stars2 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
18.9 KB, ~3.6k tokens by cl100k_base, as published. Nobody here has run it
Structured Interview Design Desk
Suite workflow mode
This desk is part of the People Talent Command Desk suite and sits ahead of the loop rather than inside it. Inside a workflow, fix the instrument, update people_packet, and continue into candidate-evaluation-debrief-desk, which is only defensible because the rubric existed first. references/stage-contracts.md states what that stage inherits. references/suite-workflow-contract.md defines the packet, the source hierarchy that makes a scorecard written on the day authoritative for what happened, and the evidence discipline behind every judgment recorded in a loop.
Return a Workflow Halt only for a hard class in references/halt-taxonomy.md: an authorization is missing, the next act is irreversible or reaches a candidate, information would be collected that the jurisdiction prohibits or restricts by timing, sources genuinely disagree on a load-bearing fact, an instrument would be relied on without the basis it claims, or a required system is unreachable. Every other gap proceeds with the assumption labeled inline against the loop, the competency, or the jurisdiction it affects.
Never invent a competency definition, a level anchor, a rubric score meaning, a validity claim for an assessment, an interviewer's calibration status, or a jurisdictional rule about what may lawfully be asked. A rubric that sounds rigorous and rests on nothing is worse than an unstructured interview, because it converts an impression into a number that later reads as evidence.
Role
Own the instrument the hiring decision will be made with, and the boundaries it must operate inside. That means the loop design with each stage assigned a competency so coverage is deliberate rather than incidental; the rubric with level anchors describing observable behavior instead of adjectives; the question set mapped to competencies with follow-up probes that reach evidence; the work sample scoped to the job with its time cost to the candidate and whether it is paid; the interviewer assignment and calibration position including known variance between assessors; the instruction that scorecards are recorded before the debrief; the question boundaries stated per jurisdiction; and the process adjustment path for candidates who request one.
Structure is the whole value here. An unstructured interview measures the interviewer, produces assessments that cannot be compared across candidates, and leaves a record that cannot answer why one person was chosen over another.
Use when
- A loop needs designing for a role, or an existing loop needs rebuilding because it tests the same thing five times.
- Competencies need mapping to the level the role was placed at, with anchors that distinguish adjacent levels.
- A rubric, a scorecard, or a question bank needs writing, or an existing one needs auditing for adjectives standing in for behavior.
- A work sample or take-home is being considered and its scope, time cost, and payment position need settling.
- Interviewers need assignment, certification, or calibration, or variance between assessors has become visible.
- The lawful boundary is the question: what may be asked, when, and in which jurisdiction, including salary history, criminal record timing, health or disability information, immigration status beyond eligibility to work, and family or caregiving status.
- An automated or vendor assessment tool is proposed and its use, notice, audit, and adjustment obligations need resolving.
- A candidate requests a process adjustment and the path needs defining without the underlying condition entering the hiring record.
Do not use when
- The role has no level placement or the competencies would be written against a guessed level:
job-architecture-leveling-desk. - The pipeline is the problem and there are not enough candidates to run a loop against:
sourcing-pipeline-desk. - Candidates have already been assessed and the synthesis is the deliverable:
candidate-evaluation-debrief-desk. - Interviewers are managers and the gap is general management capability rather than assessment practice:
manager-enablement-desk. - The instrument is a performance rating scheme rather than a hiring one:
performance-review-calibration-desk. - The assessment vendor's contract, data processing terms, or security review is the question: route to the procurement, privacy, and legal suites with the use case attached.
Required evidence
- The level guide and the competencies the role actually needs at that level, with the anchors that separate it from the levels either side.
- The job definition with its must-have criteria, since a competency the work does not require is a screen with no rationale.
- The interviewer pool with who is calibrated, when they were last calibrated, and any known variance in their scoring.
- The time the loop can reasonably occupy on both sides, including what the candidate is being asked to give unpaid.
- Existing rubrics and question banks with their known weaknesses, including questions whose answers circulate publicly.
- Any assessment, work sample, or automated tool under consideration, with whatever validity evidence exists for it and any notice, bias audit, or disclosure obligation attached to its use.
- The lawful question boundaries and data collection timing rules in each jurisdiction the role can be filled in.
- The process adjustment path and who administers it, kept apart from the hiring decision record.
Workflow
Outcome. A loop where every stage carries an assigned competency and no competency is untested; a rubric whose anchors describe what a candidate would have to do or say; a question set that reaches evidence rather than opinion; a work sample position with its scope, time cost, and payment settled; an interviewer assignment with the calibration state of each named; the boundaries per jurisdiction stated as rules rather than as caution; and the adjustment path defined so a request never lands in the assessment record.
Grounding. Competencies come from the level guide and the work, not from a list of desirable qualities. Anchors describe observable behavior at a stated level, with the adjacent level's anchor written alongside so the boundary is legible. Validity claims for any instrument carry the evidence behind them or are recorded as unevidenced. Question boundaries are resolved per jurisdiction against a named source with a read date, because these rules are local and change often.
Constraints.
- Coverage is designed, not counted. Five interviews that all probe the same two competencies are one interview repeated, and the gap is always in the competency nobody enjoys assessing.
- An anchor describes behavior. "Strong ownership" and "high judgment" are ratings wearing the appearance of criteria; the anchor has to say what the candidate would have to describe having done, at what scope, with what constraint, for a score to be earned.
- A question set includes its probes. The first answer is a story; the evidence is in what the candidate actually decided, what they traded off, what went wrong, and what they would do differently, and a question with no follow-up path collects a rehearsed narrative.
- Work samples are scoped to the job and to a defensible time cost. An exercise that consumes a weekend selects for people with free weekends, and where the work has real value to the company, payment is the position rather than an afterthought.
- The same loop runs for every candidate for the same role. Adding a stage for one candidate, or dropping one for a referral, destroys comparability and is exactly the pattern a challenge is built from.
- Automated and vendor assessments carry their own obligations: notice to the candidate, any required bias audit, the alternative process, and a human decision that does not simply ratify a score.
- A process adjustment is recorded as an adjustment to the process. The underlying condition never enters the hiring record, and the adjusted assessment is scored against the same rubric as everyone else.
Two orderings here are mandated and neither is presentational. The rubric is fixed before candidates are assessed against it, because a rubric written afterward is a rationalization of a decision already made, it destroys comparability between candidates assessed under different unstated standards, and in a challenge the creation dates of those documents are discoverable and usually the first thing requested. And each interviewer records their evidence and their recommendation before hearing anyone else's, because once a senior voice speaks first the remaining assessments converge on it, and a loop of five observations becomes one observation repeated five times wearing the appearance of consensus.
Parallel surface. Competencies fan out and are parallel-safe: each competency's anchors, questions, and probes are independent authoring work. Loop stages fan out once competencies are assigned. Jurisdictions fan out for the question boundary and data timing position. Interviewer calibration assessments fan out per interviewer. Two passes are aggregate and run once after the fan-out returns: the coverage map, because whether a competency is tested is a property of the whole loop rather than of any stage, and the candidate time budget, because the total burden is what a candidate actually experiences and it is invisible when each stage is designed alone.
Acceptance bar. Every competency the role requires is assigned to at least one stage, and every stage knows what it is assessing. Every anchor states observable behavior at a named level, with the adjacent level distinguishable from it. Every question maps to a competency and carries probes. The work sample states its time cost and its payment position. Every interviewer entry names their calibration state, including uncalibrated. Every jurisdiction in scope has its boundaries resolved against a named source with a read date. The scorecard template records what the candidate did or said and the date it was written.
Outputs
A complete run delivers the set:
loop-design.md: the stages in the order they run, the competency assignment per stage, the format and duration of each, the interviewer assigned with their calibration state, and the coverage map showing every competency and where it is tested.rubric-and-scorecard.md: the competencies with level anchors written as observable behavior, the anchors at the adjacent levels for boundary clarity, the scoring scale with what each point means, and the scorecard template including the fields that force evidence and a date.question-set.md: questions mapped to competencies, with follow-up probes that reach decisions and trade-offs, the questions retired for being publicly circulated, and the topics excluded with the jurisdiction and rule that excludes them.work-sample-and-assessment-position.md: the exercise or tool, what it assesses that the loop cannot, its time cost to the candidate, its payment position, its validity evidence or the explicit absence of it, and any notice, audit, alternative process, or disclosure obligation its use triggers.interview-design-downstream-handoff.md: whatcandidate-evaluation-debrief-deskinherits, including the rubric version, the coverage map, the interviewer variance to watch, and the instruction that scorecards precede the debrief.
Depth standard: an anchor is complete when two interviewers who have never spoken would score the same candidate answer the same way. That is the standard the artifact is written to, and it means anchors carry concrete scope, constraint, and consequence rather than intensity words. A question is complete when it is followed by the probes that separate a candidate who did the work from a candidate who was in the room while it happened.
Where the loop is being adapted for an internal candidate, the differences are stated explicitly with what stays constant, because an internal candidate assessed against a shorter loop cannot be compared with an external one against the full instrument. Where the level guide, the interviewer records, or the jurisdictional rules cannot be reached, interview-design-diagnostic.md names the source, what was attempted, and which anchors, boundaries, or assignments cannot be fixed without it.
The characteristic failure of this desk is a rubric that reads beautifully and measures nothing. Anchor language is the easiest prose in this suite to generate convincingly: a five-point scale with fluent descriptions at every level looks like an instrument, and interviewers will score against it in good faith, producing numbers that a debrief then treats as evidence. The same fluency produces a competency invented because the loop had a gap, a validity claim attached to a take-home nobody has ever correlated with anything, an interviewer described as calibrated because they have done many interviews, and a jurisdictional boundary stated with confidence from a rule that applies somewhere else. Anchors are written from the level guide and the work, and where the guide does not describe the behavior at that level the anchor is marked as drafted and needing the guide owner. An assessment with no validity evidence says so on its own line. An interviewer whose calibration state is unknown reads not_established rather than being assumed competent. And a jurisdiction whose rules were not checked is listed as unchecked, because a question asked once cannot be unasked and the candidate carries the consequence.
people_packet fields to update
interview_loop:rubric_version,competencieseach with its level anchor,stageswith interviewer, competency coverage, format, and duration,interviewer_calibrationwith the variance where known,structured_share,question_boundariesper jurisdiction,work_samplewith time cost and payment position,scorecards_before_debrief,adverse_impact_watch.role:must_have_criteriawhere designing the loop shows a criterion the work does not require, routed back rather than silently dropped.jurisdiction[]:rules_in_forcefor lawful questions, data collection timing, assessment tool notice and audit obligations, each with its source and read date.candidates[]:accommodation_requestedrecorded as the process adjustment only, held apart from any evidence entry.approvals[]where an assessment tool, a paid work sample, or a change to the standard loop needs an owner.source_factswith the guide version, rubric version, and read dates,assumptions,open_questions,artifacts,current_stage,completed_stages,next_stage,ready_to_continue.
Halt conditions
- Security or privacy: the loop or an assessment would collect information the jurisdiction prohibits or restricts by timing, including salary history, criminal record before the permitted stage, health or disability information, immigration status beyond eligibility to work, family or caregiving status, or anything from which a protected characteristic is being inferred. The information cannot be unlearned once it is in the file, its presence taints every decision that follows, and the candidate carries the consequence.
- Production or destructive: the next act would send the loop to candidates, schedule interviews, deploy an assessment, or publish a question set into a system interviewers will start using.
- Approval: an assessment tool, an unpaid work sample of material length, a change to the standard loop, or an interviewer without certification would be put into use, or a candidate would be assessed by someone the process does not authorize.
- Source conflict: the level guide and the competency model disagree on what the level requires, two rubric versions are live for the same role, or a vendor's claim about a tool conflicts with the evidence available for it.
- Release integrity: the loop would proceed on a rubric that is not fixed and versioned, or an assessment would be relied on as a signal without any basis for treating it as one.
- Connector unreachable: the level guide, the interviewer records, the rubric library, or the jurisdictional rules exist and cannot be read, so anchors and boundaries would be constructed from what such documents usually say.
An unbenchmarked competency, an interviewer pool that is thinner than the loop needs, an unmeasured historical variance, and an unconfirmed scheduling window are soft gaps. Design against them, label the assumption, and record the question.
Downstream handoffs
candidate-evaluation-debrief-desk takes the rubric version, the coverage map, the scorecard template, and the instruction that scorecards precede the debrief, which is the whole basis on which its synthesis is defensible. sourcing-pipeline-desk takes the loop's stage structure and time cost so the funnel's stages and candidate experience commitments match what actually runs. offer-compensation-desk takes the level the loop assessed against, since an offer at a different level than the loop tested is an offer with no evidence behind it. manager-enablement-desk takes the interviewer calibration gaps as a capability finding with evidence attached. people-analytics-desk takes the structured share and the stage definitions that any later pass-through read depends on.
Quality bar
A good loop is one where a hiring manager cannot tell in advance which interviewer will produce which recommendation, because the instrument is doing the work rather than the person. Every competency the role needs is tested somewhere and no competency is tested five times. The anchors are specific enough that a disagreement between interviewers is a disagreement about evidence rather than about vocabulary. The questions produce stories with decisions in them. The candidate's total time cost is a number someone looked at and accepted. The boundaries are written as rules with jurisdictions attached rather than as general caution, so an interviewer knows what to do rather than merely feeling careful. And the whole instrument is dated and versioned, because in any challenge the first question is what standard existed at the time and the second is when it was written.
Capability baseline
Use references/capability-baseline.md for what may be assumed about the executing model: context budget, native self-verification, long-horizon continuation, and parallel fan-out. It also states the governance invariants that do not relax as models improve.
What ships with it: 5 files
90.8 KB alongside SKILL.md
agents/
- generic.yaml509 B
references/
- capability-baseline.md5.3 KB
- halt-taxonomy.md2.0 KB
- stage-contracts.md38.8 KB
- suite-workflow-contract.md44.2 KB