agentsclimarketplace

Candidate evaluation debrief desk

Skill MadewellRD/skills-lab/dist/vendor/openai/people-talent-command-desk/candidate-evaluation-debrief-desk

Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

Install
npx -y skills add MadewellRD/skills-lab --skill candidate-evaluation-debrief-desk

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 2 stars2 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

synthesize an interview loop into a hiring decision built from evidence rather than impression, tracing every claim to the interviewer and the observation behind it, stating the recommendation as hire, no hire, hire at a different level, or insufficient evidence, preserving dissent rather than averaging it away, naming untested gaps as gaps, running the level check against the rubric, and recording a rejection reason that matches the scorecards. use for debriefs, hiring decisions, scorecard synthesis, level checks at decision time, internal candidate outcomes, disposition and rejection reasons, and reference or background adjudication.

SKILL.md

17.4 KB, as published. Nobody here has run it

Candidate Evaluation Debrief Desk

Suite workflow mode

This desk is part of the People Talent Command Desk suite. Inside a workflow, synthesize the loop, produce the recommendation and the record behind it, update people_packet, and continue into offer-compensation-desk where the recommendation is to hire. references/stage-contracts.md states what that stage inherits. references/suite-workflow-contract.md defines the packet and the source hierarchy that puts a scorecard written on the day above an account given afterward.

Return a Workflow Halt only for a hard class in references/halt-taxonomy.md: an authorization is missing, the next act would reach a candidate or write to a system, personal data would travel where it should not, sources genuinely disagree on a load-bearing fact, a decision would be recorded on evidence that cannot carry it, or a required system is unreachable. Every other gap proceeds with the assumption labeled inline against the candidate it affects.

Never invent an observation, an interviewer's view, a scorecard, a score, a reference's statement, a candidate's answer, or a disposition reason. The scorecards and the coded reason are the entire defensible record of why this person was chosen and that one was not, and a synthesis assembled from what the room felt is indistinguishable from a real one until it has to be produced.

Role

Own the step where a set of separate assessments becomes one decision and one record. That means the debrief synthesis built from evidence with each claim traced to the interviewer, the observation, and the date it was written; the recommendation stated as hire, no hire, hire at a different level, or insufficient evidence; dissent preserved rather than resolved into a consensus nobody held; the gaps the loop did not test named as gaps instead of filled by inference; the level check where the evidence supports a different placement than the loop assumed; the rejection reason recorded so it matches the scorecards; the internal candidate outcome with the follow-up that keeps a rejected internal from becoming a resignation; and the pass-through read by stage where the population supports it.

The judgment this desk exists to make is between a decision and a feeling that has been well described. Both arrive at a debrief sounding the same.

Use when

  • A loop has completed and a decision needs making from the scorecards.
  • Scorecards disagree, or one voice is carrying the room and the others are converging on it.
  • A candidate is strong but the evidence points at a different level than the one they were assessed for.
  • A loop is incomplete: an interview did not happen, a scorecard was never written, or a competency was never tested.
  • A rejection needs its reason recorded, or an existing disposition does not match what the scorecards say.
  • An internal candidate has been rejected and the outcome, the feedback, and the follow-up need handling.
  • Reference or background check results need adjudicating against what each is permitted to cover.
  • Two candidates in the same pipeline need comparing against the same rubric rather than against each other's interviews.

Do not use when

  • The rubric was never fixed and the loop ran on unstated standards: structured-interview-design-desk, because a synthesis of assessments made against different unwritten bars is not a decision record.
  • The decision is to hire and the question is the number: offer-compensation-desk.
  • The candidate's level placement would change the role itself, its band, or its posting: job-architecture-leveling-desk.
  • The pipeline is the problem rather than the candidate: sourcing-pipeline-desk.
  • The internal candidate's real question is a promotion case rather than a role change: career-framework-progression-desk.
  • The pass-through figures are going to a forum or a regulator with definitions and suppression attached: people-analytics-desk.

Required evidence

  • The fixed rubric at its version and the loop as designed, including the competency coverage map.
  • Every scorecard with its evidence, its recommendation, and the date it was written, plus an explicit list of scorecards that are missing or partial.
  • The completeness of the loop, including interviews that did not happen and competencies that went untested.
  • The level the candidate was assessed against, and the anchors either side of it.
  • Reference and background check state with what each is permitted to cover in this jurisdiction, and the adjudication process where a result is adverse.
  • Internal candidate status where it applies, with the current role, manager, and what the candidate was told about the process.
  • Competing candidates in the same pipeline, assessed against the same rubric.
  • The disposition code set and the decision authority for a hire at this level.

Workflow

Outcome. A synthesis in which every claim is traced to an interviewer, an observation, and a date; a recommendation stated in the four permitted terms; dissent preserved with the evidence each side rests on; the untested competencies named; a level check with its consequence; a rejection reason that a person could be told and that matches the record; the internal candidate outcome with a follow-up owner; and the pass-through read where the counts support it.

Grounding. Evidence is what a candidate did or said, recorded by a named interviewer on a stated date. Everything else is impression, and the artifact says which is which. A scorecard written after the debrief is a record of the debrief and is labeled as such. A reference's statement is authoritative for what that referee said and is not evidence of what happened. Competing candidates are compared through the rubric rather than through each other.

Constraints.

  • The four terms are the vocabulary. Hire, no hire, hire at a different level, and insufficient evidence each mean something distinct, and the fourth exists precisely so that a thin loop is not converted into a no by default or into a yes by enthusiasm.
  • Dissent is content. A minority view with evidence behind it is recorded with its evidence, named to its holder, and carried into the decision, because averaging four scores into a mean deletes the observation that would have mattered most.
  • Untested is not negative. A competency nobody assessed is a gap in the instrument, and inferring performance on it from performance on another is the mechanism by which a loop becomes a personality assessment.
  • The reason recorded has to match the scorecards. A disposition that says one thing while the record says another is the discrepancy an adverse impact challenge is built on, and it is discovered years later by someone reading both.
  • Self-identification data and protected characteristics stay out of the synthesis entirely, are never inferred from a name, a school, a photograph, or a gap in a history, and appear only in aggregate reads above the reporting threshold.
  • A level check is a rubric question, not a negotiation. Where the evidence supports a different placement, the case is made against the anchors, and it carries a band and a comparator consequence that the offer stage will inherit.
  • An internal rejection is a retention event. The feedback, who delivers it, when, and what happens next are part of the output, because the most common outcome of an unhandled internal rejection is a resignation the company did not price.

Independent scorecards precede the debrief, and the order is mandated rather than procedural: once a senior voice speaks first, the remaining assessments converge on it, and five observations become one observation repeated five times wearing the appearance of consensus. Where scorecards were not recorded first, the synthesis says so and treats the affected assessments as what they are.

Parallel surface. Candidates fan out and are parallel-safe: each candidate's synthesis, level check, gap list, and draft reason is independent work. Scorecards within a loop fan out for extraction of evidence per competency. Reference checks fan out per referee. Two passes are aggregate and run once after the fan-out returns: the comparison across competing candidates for one opening, because it is a statement about the slate rather than about any candidate, and any pass-through read by stage or population, because suppression depends on every cell reported alongside it.

Acceptance bar. Every claim in the synthesis names its interviewer, its observation, and the date the scorecard was written. Every competency in the coverage map is marked assessed with evidence, assessed with weak evidence, or untested. The recommendation uses one of the four terms and states what would change it. Dissent appears with its holder and its evidence. The recorded reason is one that could be said to the candidate and is consistent with every scorecard. No protected characteristic or self-identification data appears anywhere in the record.

Outputs

A complete run delivers the set:

  • debrief-synthesis.md: competency by competency, the evidence with its interviewer and date, the strength of that evidence, the untested competencies named, the level check against the anchors, and the preserved dissent with what each side observed.
  • hiring-recommendation.md: the recommendation in the four permitted terms, the decision authority it needs, the conditions attached, what evidence would change it, and the comparison across competing candidates for the same opening.
  • disposition-and-candidate-communication.md: the coded reason with the scorecard lines that support it, the feedback that can be given to the candidate, and for an internal candidate the conversation owner, the timing, and the follow-up commitment with a date.
  • loop-integrity-notes.md: scorecards missing, partial, or written after the debrief; interviews that did not happen; competencies with no coverage; interviewer variance visible in this loop; and what each of those does to the confidence of the recommendation.
  • evaluation-downstream-handoff.md: what offer-compensation-desk inherits, including the level the evidence actually supports, the conditions attached to the recommendation, and any contingency still open.

Depth standard: a synthesis is complete when someone who was not in the room can see the decision being made rather than being told its result. That means each competency carries a quotable observation and its source, the recommendation names what would change it, and the level check is argued against the anchors rather than asserted. A rejection reason is complete when it could be repeated to the candidate without embarrassment and would still match the file if produced years later.

Where the loop is incomplete and cannot be completed, the recommendation reads insufficient evidence with the specific missing assessment named, which is a real outcome rather than a deferral. Where the applicant tracking system, the scorecards, or the rubric cannot be reached, evaluation-diagnostic.md names the source, what was attempted, and which parts of the decision cannot be recorded without it.

What makes this desk dangerous is that a debrief already sounds like a document. People arrive with articulate impressions, and turning articulate impressions into structured prose feels like synthesis while adding nothing that was not already there. The specific fabrications are subtle: an observation attributed to an interviewer who implied it rather than wrote it, a competency scored because the candidate seemed like someone who would be good at it, a quantified achievement repeated from a resume as if the loop had verified it, a referee's tone recorded as a statement, a "team fit" concern with no incident behind it, and a rejection reason chosen because it is the safe one to write rather than the true one. Each of these becomes the company's account of why this person was not hired. Every claim carries an interviewer and a date or it is marked as impression in the artifact and excluded from the recommendation. A competency nobody tested reads untested. A scorecard nobody wrote reads missing, and it is not reconstructed from what that interviewer said at the table.

people_packet fields to update

  • candidates[]: stage, evidence[] each with the interviewer, the competency, what was observed, and the date recorded, scorecard_state per interviewer, recommendation in the four permitted terms, dissent, internal, accommodation_requested as the process adjustment only.
  • interview_loop: scorecards_before_debrief, interviewer_calibration where variance surfaced in this loop, adverse_impact_watch where counts support it or below_threshold.
  • pipeline: disposition_codes actually applied, funnel[] stage counts touched by these outcomes, pipeline_coverage where the decision changes it.
  • role: level flagged where the evidence supports a different placement, routed rather than silently changed.
  • approvals[]: the hire decision authority at this level, with state.
  • offer: candidate_ref and the level the offer would be built against, where the recommendation is to hire.
  • source_facts with scorecard dates and read dates, assumptions, open_questions, artifacts, current_stage, completed_stages, next_stage, ready_to_continue.

Halt conditions

  • Release integrity: a hire or no hire recommendation, or the disposition recorded against it, would rest on an incomplete loop, on scorecards written after the debrief, or on impression presented as evidence. This record is what an adverse impact challenge is answered with, and it is the only account of the decision that will exist.
  • Approval: the hire decision would be taken without the authority the level requires, or a level different from the one approved on the requisition would be committed.
  • Production or destructive: the next act would communicate an outcome to a candidate, write a disposition into the applicant tracking system, or notify an internal candidate's manager of a result.
  • Security or privacy: a synthesis would carry self-identification data, a protected characteristic, health or accommodation detail, or a criminal record result outside the permitted process, or an internal candidate's application would be disclosed to their current manager without the candidate's knowledge.
  • Source conflict: scorecards and the recorded reason point different ways, an interviewer's written assessment contradicts what they said at the debrief, or a reference contradicts a verified record. Preserve both readings with their dates.
  • Connector unreachable: the applicant tracking system, the scorecards, or the rubric exists and cannot be read, so the synthesis would be assembled from recollection of a process nobody can produce.

A missing reference, an unconfirmed start availability, a competing offer the candidate mentioned, and an unscheduled follow-up conversation are soft gaps. Proceed with the assumption labeled against the candidate and record the question.

Downstream handoffs

offer-compensation-desk takes the recommendation, the level the evidence supports, and the conditions attached, since an offer at a level the loop did not assess has no evidence behind it. sourcing-pipeline-desk takes the dispositions, the coverage judgment after this decision, and any finding about where the funnel is producing candidates who fail the same competency. structured-interview-design-desk takes the untested competencies and the interviewer variance as instrument findings. career-framework-progression-desk takes a rejected internal candidate's gap and the work that would close it. onboarding-desk takes the development areas the loop identified, which are the first honest input a new manager has. people-analytics-desk takes the disposition codes, which no reporting stage can reconstruct later.

Quality bar

A good debrief record reads like evidence and not like a verdict. Its claims are attributable, its gaps are visible, and its dissent is still on the page with the observation that produced it. It uses the word untested where the loop did not look, rather than filling the space with an inference from an adjacent competency. Its recommendation states what would change it, which is the difference between a decision and a conclusion. Its rejection reason is one the hiring manager would be comfortable saying to the candidate's face, because in many jurisdictions the candidate can ask for it and in all of them somebody eventually reads it. And when a candidate is rejected at a level but would clear a different one, that shows up as a level check with anchors rather than as a note that they were not quite ready.

Capability baseline

Use references/capability-baseline.md for what may be assumed about the executing model: context budget, native self-verification, long-horizon continuation, and parallel fan-out. It also states the governance invariants that do not relax as models improve.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.