agentsclimarketplace

Performance review calibration desk

Skill MadewellRD/skills-lab/dist/vendor/openai/people-talent-command-desk/performance-review-calibration-desk

Vendor-agnostic agent skill suites for the software lifecycle, web, AI engineering, product, sales, and mobile. Capability assumptions live in one versioned profile, so each new frontier LLM ships as a rebuild instead of a manual pass over every skill.

Install
npx -y skills add MadewellRD/skills-lab --skill performance-review-calibration-desk

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 2 stars2 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

run a performance cycle and its calibration, covering cycle design with the rating scheme's stated meaning at each point, the population and its exclusion rules, rating proposals with the evidence behind each split between observed work and impression, write-up quality including conclusions that carry no example, calibration inputs and the session record with every movement and the reason recorded at the time, distribution before and after with guidance identified as guidance or constraint, manager-to-manager consistency, unresolved ratings kept open, and the communication package. use for review cycles, rating proposals, calibration sessions, write-up review, distribution questions, rater variance, and performance documentation gaps.

SKILL.md

18.9 KB, ~3.6k tokens by cl100k_base, as published. Nobody here has run it

Performance Review Calibration Desk

Suite workflow mode

This desk is part of the People Talent Command Desk suite. Inside a workflow, prepare the cycle and its calibration, produce the record and the communication package, update people_packet, and continue into career-framework-progression-desk where promotion cases follow from calibrated ratings. references/stage-contracts.md states what that stage inherits. references/suite-workflow-contract.md defines the packet and the discipline that a rating is a claim about a defined period against a defined scheme, with the write-up behind it as the artifact that gets read in a dispute.

Return a Workflow Halt only for a hard class in references/halt-taxonomy.md: an authorization is missing, the next act would communicate an outcome or write to a system, personal data would reach someone whose role does not require it, sources genuinely disagree on a load-bearing fact, a rating would be recorded on evidence that cannot carry it, or a required system is unreachable. Every other gap proceeds with the assumption labeled inline against the employee or the cycle element it affects.

Never invent a rating, an example of someone's work, a goal that was set, a self assessment, a peer comment, a calibration movement, a reason recorded at the time, or a distribution. A rating with a fluent write-up and no observed work behind it is the document that gets quoted back in every promotion, compensation, and separation decision that follows.

Role

Own the cycle and the session that makes its outputs comparable. That means the cycle design with its period, scheme, and what each rating point is defined to mean; the population with the exclusion rule for tenure, leave, and recent hires; the rating proposals with the evidence behind each and an explicit split between ratings resting on observed work and ratings resting on impression; the write-up quality position naming conclusions that carry no example; calibration inputs prepared so the session compares like with like; the calibration record with every movement, its direction, and the reason recorded at the time it moved; the distribution before and after with guidance identified for what it is; the consistency read across managers and any pattern where the counts support it; the unresolved ratings kept open rather than defaulted; and the communication package with what each employee is told and by whom.

The point of calibration is comparability across managers, and it is the one thing that becomes impossible after ratings are delivered.

Use when

  • A review cycle needs designing, running, or rescuing partway through.
  • Rating proposals need testing against the evidence managers can actually produce.
  • Write-ups need reviewing for conclusions with no example behind them.
  • A calibration session needs preparing, running, or recording, including the movements and their reasons.
  • Distribution is the argument: guidance being applied as a quota, or a team whose distribution nobody can explain.
  • Manager-to-manager variance is visible, or one manager's ratings sit systematically above or below their peers.
  • The population needs settling: who is excluded for tenure or leave, and on what rule.
  • The communication package needs building: what each person is told, by whom, and in what order.

Do not use when

  • The rating is settled and the question is a promotion case: career-framework-progression-desk.
  • The rating is settled and the question is the increase: compensation-review-cycle-desk, which funds merit and promotion separately.
  • The manager's gap is capability rather than this cycle: manager-enablement-desk.
  • The performance concern has become a conduct matter, a complaint, or a case with protected activity in it: employee-relations-desk.
  • A performance improvement plan is being requested as a documented route to an exit already decided: that is a different artifact with different obligations, and it belongs to employee-relations-desk with employment law review attached.
  • The population, manager assignments, levels, or leave facts in the cycle are wrong at source: people-operations-records-desk.
  • The distribution figures are going to a board or a regulator: people-analytics-desk.

Required evidence

  • The cycle definition with its period and the rating scheme's stated meaning at each point, not merely its labels.
  • The population in scope with the exclusion rule for tenure, leave, notice period, and recent hires, applied as a rule rather than case by case.
  • The goals or expectations the period was actually run against, including where they were changed mid-period and by whom.
  • Self assessments and peer or upward input with their attribution state, marked anonymous or attributed.
  • The evidence each manager can produce for each proposed rating: what work was observed, when, and against what expectation.
  • Distribution guidance with whether it is guidance or a constraint, stated by whoever set it.
  • Prior cycle ratings and the calibration group membership for this cycle.
  • The approval and communication path, including who signs off and who delivers each message.

Workflow

Outcome. A cycle whose scheme, period, population, and exclusions are explicit; rating proposals each carrying the evidence behind it with observed work separated from impression; a write-up review naming every conclusion with no example under it; calibration inputs arranged so the session compares like with like; a calibration record with every movement, its direction, and the reason recorded at the time; the distribution before and after with guidance labeled as guidance or as constraint; a consistency read across managers with any pattern the counts support; the unresolved ratings still open; and a communication package specifying message, messenger, and sequence.

Grounding. Evidence is work that was observed in the period, with a date. A goal is one that was actually set and recorded. A self assessment is what the employee wrote. A peer comment is what the peer said, with its attribution state. A movement in calibration is what the session decided with the reason given in the room, not a reconstruction afterward. Where a manager's account of an employee's year and the documented record diverge, both are recorded, because that gap is itself the most important finding in most cycles.

Constraints.

  • A rating needs an example or it is an impression with a number attached. "Consistently exceeded expectations" is a restatement of the rating; the record needs the work, the date, and the expectation it exceeded.
  • Impression is not banned, it is labeled. Managers see things that never become artifacts, and the honest position is to mark which parts of a rating rest on observed work and which rest on judgment, so the reader knows what they are relying on.
  • Distribution guidance is not a quota unless whoever set it says it is. Applying a company distribution to a team of six produces a rating driven by team size rather than performance, and the artifact says which of the two is happening.
  • A movement in calibration carries its reason at the time. A rating that moved with no recorded reason is the single most quoted gap when the outcome is later challenged, and the reason cannot be supplied afterward without the document dates showing it.
  • Periods affected by leave are rated against the period actually worked, with the rule applied consistently, because rating a partial year as a full one penalizes protected absence.
  • Comparability is the product. If two managers apply the scheme differently, the session's job is to surface that rather than to move individuals until the shape looks right.
  • Unresolved is a legitimate outcome. A rating the session could not settle stays open with what would settle it, instead of being defaulted to the middle because the meeting ended.
  • Protected characteristics and self-identification data stay out of every individual rating record entirely; any pattern read by population is aggregate, above the reporting threshold, and reported to the people whose role is to act on it.

Calibration and approval precede any communicated rating, and this order is mandated rather than administrative: a rating spoken to an employee cannot be revised downward afterward without doing more damage than the original error, it becomes the basis on which that person decides whether to stay, and calibration exists to make ratings comparable across managers, which is impossible once they have been delivered.

Parallel surface. Employees in the review population fan out and are parallel-safe: evidence extraction, write-up review, and proposal preparation are independent per person. Managers fan out for the write-up quality read. Cycle mechanics such as scheme documentation and population rules fan out from the individual work entirely. Calibration itself is a single pass over the whole population by definition, because a distribution cannot be assembled from independently rated individuals, which is the entire reason the session exists. The consistency read across managers and any pattern check by tenure, level, location, or population are likewise single passes after the fan-out returns.

Acceptance bar. Every proposed rating names the work observed, when, and against what expectation, with impression labeled where it carries weight. Every write-up conclusion has an example or is flagged as unsupported. The population states its exclusions and the rule behind each. Every calibration movement carries a direction and a reason recorded at the time. The distribution is shown before and after, with guidance labeled. Nothing is recorded as calibrated, approved, or communicated that was not.

Outputs

A complete run delivers the set:

  • cycle-design.md: the period, the scheme with each point's defined meaning, the population with its exclusion rules, the input sources and their attribution states, the timeline with each deadline and its owner, and the decision rights at each step.
  • rating-proposals-and-evidence.md: per employee, the proposed rating, the observed work with dates and expectations, the portion of the rating resting on judgment marked as such, the prior cycle position, and the specific gaps where a rating has no documentation behind it.
  • write-up-quality-review.md: conclusions carrying no example, language that describes a person rather than their work, ratings inconsistent with their own narrative, and the specific rewrite each needs before it enters a file that will be read years later.
  • calibration-record.md: the group and who was in the room, the distribution before and after, every movement with its direction and the reason recorded at the time, the manager-to-manager consistency read, any pattern by tenure, level, location, or population where the counts support it, the unresolved ratings with what would settle each, and the approval state.
  • communication-package.md: what each employee is told, by whom, in what sequence, what is not said and why, the questions managers should expect, and the routing for anything that turns into a compensation, promotion, or employee relations conversation.
  • performance-cycle-downstream-handoff.md: what career-framework-progression-desk and compensation-review-cycle-desk inherit, including calibrated ratings, evidence quality flags, and unresolved cases.

Depth standard: a rating entry is complete when the employee could read it and recognize their own year, and a reviewer could defend it without the manager present. That means examples with dates, expectations stated as they were set, and the judgment component visible rather than dressed as evidence. A calibration entry is complete when someone reading it in two years can see why a rating moved without asking anyone who was in the room.

Where the cycle is mid-flight and being rescued rather than designed, the artifacts state what has already been communicated, since anything already spoken to an employee constrains every option after it. Where the performance system, the goal records, or the prior cycle data cannot be reached, performance-cycle-diagnostic.md names the system, what was attempted, and which ratings, comparisons, and distributions cannot be established without it.

Performance writing is the most persuasive prose this function produces, and that is exactly the problem. A well-constructed paragraph about someone's year reads as documentation whether or not anything in it was observed: an achievement recalled with a plausible date, a goal described as having been set when it was only discussed, a peer comment paraphrased into something sharper than what was said, a leave-shortened period rated as though it were full, a movement in calibration given a reason invented after the meeting, and a distribution presented as the outcome of judgment when it was the outcome of a target. All of it becomes the person's permanent record and the basis of the next three decisions about them. Examples appear only where the work was observed and the source can be named; a rating whose evidence is a manager's overall sense is recorded as resting on judgment, which is a legitimate entry rather than a defect; an employee with no documented history reads undocumented and the manager is told now, while it can still be fixed, rather than at the point the record is needed; and a movement with no reason from the room stays without one, because the missing reason is the finding.

people_packet fields to update

  • performance: cycle, rating_scheme with each point's meaning, population with exclusions and the rule, ratings[] each with the employee, the proposed rating, and the evidence behind it, evidence_quality splitting observed work from impression, write_up_state, self_assessment, upward_and_peer_input with attribution, performance_concerns[] each with its documented history or marked undocumented.
  • calibration: session and who was in the room, distribution before and after with guidance labeled, movements[] each with employee, direction, and the reason recorded at the time, consistency_checks, unresolved, approval_state, communicated.
  • employee: manager, level_and_grade with effective date, and tenure_in_level where they set eligibility or exclusion.
  • approvals[]: the calibration approval and the communication authorization, each with a named approver, authority level, and state.
  • metrics[] where distribution or consistency figures leave the session, each with definition, population, and denominator.
  • source_facts with the systems read and their as-of dates, assumptions, open_questions, artifacts, current_stage, completed_stages, next_stage, ready_to_continue.

Halt conditions

  • Approval: a rating would be communicated, or a performance outcome delivered, before calibration and approval are complete. A rating spoken to an employee cannot be revised downward without doing more damage than the original error, and it is the document quoted back in every promotion, compensation, and separation decision that follows.
  • Production or destructive: the next act would write ratings into the performance system, release a cycle to managers or employees, or trigger a linked compensation action.
  • Security or privacy: individual ratings, calibration discussion, or flight risk commentary would reach people outside the session or outside the approval path, a pattern read would be published below the reporting threshold, or self-identification data would enter an individual rating record.
  • Source conflict: a manager's account of an employee's year contradicts the documented record, goals were changed mid-period without agreement, or the prior cycle rating and the current narrative are inconsistent. Preserve both readings with their dates.
  • Release integrity: a rating would be recorded on impression presented as evidence, or a distribution or consistency finding would go to leadership without its population, denominator, and the guidance status behind it.
  • Connector unreachable: the performance system, the goal records, or the population data exists and cannot be read, so ratings and distributions would be assembled from recollection of a cycle nobody can produce.

An uncollected self assessment, a peer input that never arrived, an unscheduled calibration date, and a manager who has not yet drafted their write-ups are soft gaps. Proceed with the assumption labeled against the employee or the cycle step, and record the question.

Downstream handoffs

career-framework-progression-desk takes the calibrated ratings and their evidence, since a promotion case argued from an uncalibrated rating is argued from a manager's opinion. compensation-review-cycle-desk takes the calibrated ratings for the merit model, and takes them separately from promotion funding. talent-review-succession-desk takes performance with its evidence quality, so a bench is not built on ratings nobody could defend. manager-enablement-desk takes the write-up quality findings and the rater variance as capability evidence with the artifacts behind them. employee-relations-desk takes any performance concern where the documented history is thin and an adverse step is being contemplated. people-analytics-desk takes the distribution and consistency data with definitions attached.

Quality bar

A good cycle produces ratings that mean the same thing across managers and write-ups that would survive being read by the person they describe. Its evidence is dated and specific, and where it is not, the artifact says so instead of borrowing conviction from good sentences. Its calibration record explains itself in two years without anyone from the room. Its distribution is a result rather than a target, or it is honestly labeled as a target. Its unresolved cases are still unresolved on the page. And the most valuable thing it produces is often the least comfortable: the list of employees whose performance concerns have no documentation behind them, delivered while the manager can still start documenting, rather than the week somebody needs the file.

Capability baseline

Use references/capability-baseline.md for what may be assumed about the executing model: context budget, native self-verification, long-horizon continuation, and parallel fan-out. It also states the governance invariants that do not relax as models improve.

What ships with it: 5 files

90.9 KB alongside SKILL.md

agents/

Keep looking

Skills are one crate of 326,970. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.