This guide walks through the performance review calibration process step by step, from setting scope to documenting the outcome, along with a ready-to-use checklist for HR, managers, and senior leadership. The six steps are:
- Decide what calibration is allowed to change
- Group employees so the comparisons are fair
- Distribute the rating data before the meeting
- Brief managers and set the ground rules
- Run the session by performance segment
- Document every change and close the loop with employees
The aspects that make or break the process are a pre-meeting data pack, a structured meeting agenda, and a written record of what changed and why.
Key Takeaways
- Performance review calibration brings managers together to compare notes and adjust employee ratings before they are finalized, so a high rating means the same thing across every team.
- Done well, calibration reduces bias by forcing managers to defend ratings with evidence. Done poorly, it introduces new bias through dominance dynamics and groupthink, so how the session is designed matters as much as whether you run one.
- Most performance review processes aren't working. Only 2% of CHROs feel their system inspires improvement, and about 1 in 5 employees see the process as fair, per Gallup (2024).
What Is Performance Review Calibration?
Performance review calibration is the process of bringing managers together to compare, discuss, and adjust employee ratings before they are finalized and communicated. It exists because performance management, broadly, is not working. Only 2% of CHROs surveyed by Gallup strongly agree their performance management system inspires employees to improve, and about one in five employees strongly agree that their performance review process is fair and transparent.
Calibration will not fix every part of that on its own, but it is a structured process in which managers and HR leaders compare employee ratings across teams and departments before those ratings are finalized. The goal is to ensure that a '4 out of 5' means the same thing whether it was given by a lenient rater in marketing or a strict rater in engineering.
What Does Review Calibration Look Like in Practice?
In practice, calibration sits between review submission and rating release. After managers have submitted their individual ratings and before any rating is communicated to an employee, the calibration session runs. Ratings can be adjusted up or down as a result. The final rating an employee receives is the post-calibration rating, not the manager's original submission.
Does Calibration Actually Reduce Bias?
The honest answer is: sometimes, and it depends on how the session is run. Research by Demeré, Sedatole, and Woods found that calibration committees improve the consistency of ratings across supervisors and help mitigate leniency bias. The mechanism is straightforward: when managers must explain and defend ratings in front of peers, vague impressions give way to specific evidence.
That said, the case for caution is more recent. A 2024 HBR article by Raafiya Ali Khan, Rachel M. Korn, and Joan C. Williams found that calibration meetings can introduce bias as well as remove it through dominance dynamics, centrality bias, and 'tightrope bias.'
The practical implication is not to abandon calibration but to design it carefully. Automated outlier detection tools can flag distribution anomalies before the meeting, and structured facilitation with explicit ground rules addresses the social dynamics that allow bias to enter. Done well, calibration is a meaningful check on individual manager judgment. Done poorly, it can amplify the problems it was meant to solve.
The Performance Review Calibration Process: 6 Steps

Each step below covers what it is, who owns it, when it happens, and what goes wrong without it. Work through them in order: skipping or compressing earlier steps makes later ones harder to run.
Step 1: Decide What Calibration Is Allowed to Change
- Owner: HR leadership
- Timing: 2 to 3 weeks before the session
- Output: Written scope statement distributed to all managers before the session
The most reliable way to make a calibration session adversarial is to leave its scope undefined. Managers who arrive not knowing whether the session can affect merit increases, promotion nominations, or PIPs will defend their ratings as though all of those things are at stake. Defining scope in advance turns a defensive conversation into a calibration conversation.
To do this well, decide and document whether calibration can change final ratings, merit pay, promotion nominations, and PIP status, or ratings only. If a distribution target applies, state whether it is a guideline or a quota. Managers will ask.
What goes wrong: Scope is communicated verbally or assumed rather than written down. Managers arrive with different mental models of what the session can touch. The facilitator spends the first 20 minutes resolving a scope dispute that should have been settled three weeks earlier.
Step 2: Group Employees So the Comparisons Are Fair
- Owner: HR Business Partner
- Timing: 2 weeks before the session
- Output: Calibration group assignments shared with managers in the pre-meeting data pack
Calibration only works when the people being compared are actually comparable. Grouping by reporting line evaluates employees relative to their own team rather than across the organization at the same level and role family.
Instead, group by level and role family. Handle employees with no natural peer group explicitly: calibrate them in the closest available group and note the limitation, or have HR review them separately. For mid-cycle role changes, calibrate based on the role held for the majority of the cycle, or split and calibrate each portion.
What goes wrong: Employees are grouped by reporting line rather than level and role family. The comparison becomes 'who is the best person on my team' rather than 'what does strong performance at this level look like across the organization.' Rating inflation or deflation by department compounds across cycles.
Stepr 3: Distribute the Rating Data Before the Meeting
- Owner: HR Business Partner
- Timing: 3 to 5 business days before the session
- Output: Pre-meeting data pack sent to all managers in the calibration group
Managers who see the rating distribution for the first time during the session spend the opening reacting to data they haven't processed. This is where dominance dynamics take hold: the most prepared or vocal person in the room frames the conversation before others have oriented themselves.
At minimum, the pack should include each employee's name, level, role, manager, goals, submitted rating, and performance notes, along with the rating spread by manager and by team. Include demographic-level distributions only where sample sizes are sufficient to protect privacy and where permitted under applicable privacy and employment rules. HR analytics platforms can generate these distributions automatically.
What goes wrong: The data pack is sent the day before or shared on screen at the start of the session. Managers who haven't reviewed the distribution cannot contribute meaningfully. The session becomes a reactive negotiation rather than a structured comparison.
Step 4: Brief Managers and Set the Ground Rules
- Owner: HR Business Partner
- Timing: Brief: 1 week before; ground rules: opening of the session
- Output: Written ground rules read aloud and displayed at the start of the meeting
Managers who believe their ratings are already final will defend rather than discuss. The brief before the session should state in writing that all ratings are provisional, the session can move a rating in either direction, and evidence rather than seniority determines outcomes.
With that in mind, read the ground rules aloud at the start of the session and display them throughout:
- All ratings submitted are provisional. Nothing is final until the session closes.
- Every rating discussion requires specific behavioral evidence, not general impressions.
- Any participant can request a rating discussion. No manager's rating is immune from review.
- Seniority does not determine outcomes. The most senior person in the room does not have a veto.
- Treat calibration discussions and employee information as confidential, and share only what is necessary with authorized participants.
- A rating may move as a result of this session. That is the point of being here.
What goes wrong: Managers understand the session as a formality. When a rating is challenged, they treat it as a criticism rather than a calibration question. The session produces cosmetic adjustments rather than genuine cross-manager alignment.
Step 5: Run the Session by Performance Segment
- Owner: HR Business Partner (facilitator)
- Timing: The calibration meeting
- Output: A running log of ratings reviewed, ratings moved, and the evidence cited for each change
Work through one performance segment at a time rather than one manager at a time. Starting with all employees rated Exceeds Expectations, then Meets Expectations, then Below Expectations keeps the comparison horizontal across the group rather than vertical within one manager's team.
From there, for each employee discussed, the presenting manager states the rating and the primary evidence in two to three sentences. The facilitator opens the floor for questions. If no one challenges, the rating is confirmed. If a challenge is raised, the facilitator requires specific evidence from the challenger before discussion proceeds.
What goes wrong: The session runs manager by manager rather than segment by segment. Managers front-load defense of their own employees and disengage when other teams are discussed. The comparison stays vertical when it needs to be horizontal.
Step 6: Document Every Change and Close the Loop with Employees
- Owner: HR Business Partner
- Timing: Immediately after the session; employee communication within 5 business days
- Output: Written record of all rating changes with evidence cited; final ratings entered in the HRIS
The calibration record is an important operational record that may also be relevant in employment disputes, audits, or litigation. Record every rating reviewed, every rating changed, who requested the change, and the evidence cited. Retain calibration records according to applicable employment, privacy, document-retention, and litigation-hold requirements. Confirm the retention period with employment counsel.
On the communication side, managers can tell employees their rating reflects cross-organizational calibration and should communicate the evidence behind the final rating. Managers should not share other employees' ratings or confidential discussion details. Employee engagement survey tools can measure whether employees perceive the process as fair after the cycle, providing input for future process improvement.
What goes wrong: Rating changes are entered in the system but the reasoning is not documented. Months later, when an employee raises a dispute or a manager asks why a rating was moved, there is no record.
Note: This guide is operational guidance, not legal advice. Consult employment counsel before finalizing record-retention, confidentiality, compensation, promotion, PIP, or rating-override policies.
The Calibration Meeting Agenda
Time the session to the size of the group. These estimates assume a prepared group that has reviewed the data pack in advance. Add 20% for a group running calibration for the first time.
Time guidance by group size:
- Up to 15 employees: 90 to 120 minutes
- 16 to 30 employees: 2.5 to 3.5 hours with one break
- 31 to 50 employees: Full day with two breaks; consider splitting by segment
- 50+ employees: Split into multiple sessions by level or role family
Remote and hybrid calibration: Use a shared screen throughout, encourage video participation where appropriate while providing accessible alternatives for those who cannot use video, and manage the speaking floor more actively than in person. Use breakout rooms for large groups with a designated note-taker and structured report-back. Avoid asynchronous calibration for contested ratings.
The Performance Review Calibration Checklist

The following three checklists are organized by role. Copy and use directly.
HR Business Partner / HR Lead
- Written scope statement drafted and approved by HR leadership
- Calibration groups defined by level and role family, not reporting line
- Mid-cycle role changes handled explicitly with a documented approach
- Pre-meeting data pack prepared: employee list, submitted ratings, goals, distribution by manager and team
- Data pack sent 3 to 5 business days before the session
- Managers briefed: session has authority to move ratings; all ratings are provisional until session closes
- Facilitator prepared with ground rules to read aloud at session open
- Running log template prepared for real-time documentation during the session
- Session time confirmed against group size with break schedule
- All rating changes documented with evidence cited immediately after the session
- Final ratings entered in the HRIS within 48 hours of session close
- Manager communication timeline confirmed (within 5 business days)
Manager
- Pre-meeting data pack reviewed before the session
- Rating distribution for your team compared against peers at the same level and role family
- Specific behavioral evidence prepared for every rating in your group, especially top and bottom performers
- Session scope understood: know what calibration is allowed to change
- Prepared to articulate evidence rather than impressions when defending or challenging a rating
- Prepared to listen to evidence rather than defend reflexively when your rating is challenged
- Clear on confidentiality: what you can and cannot tell employees about the session
- Employee communication delivered within the stated timeline after session close
Senior Leadership
- Scope confirmed and approved before the process begins
- Distribution guidelines or targets clarified as guidelines or quotas, not left ambiguous
- Ground rules include explicit statement that seniority does not determine outcomes
- Calibration record retention policy confirmed with legal or HR counsel
- Post-cycle fairness perception data reviewed and used as input for process improvement
- Calibration outcomes reviewed at aggregate level for demographic patterns over time
Common Calibration Mistakes
Most calibration failures trace back to the same handful of process gaps. These are the ones that appear most consistently, along with the warning sign that each one produces.
- Defining scope too late or not at all. Managers arrive not knowing what the session can change. The first 20 minutes become a scope negotiation.
- Grouping by reporting line instead of level and role family. Creates within-team comparisons that cannot produce cross-organizational consistency.
- Distributing data at the session, not before it. Managers cannot calibrate data they haven't reviewed. The most vocal person in the room frames the conversation.
- Treating sessions as a formality. If managers believe ratings are effectively final before the session, the session produces theater rather than calibration.
- No escalation process for disagreements. Without a pre-established decision-maker or escalation path, a disputed rating can stall the session.
- Running manager by manager rather than segment by segment. Produces within-team adjustments rather than cross-team consistency.
- No documentation of what changed and why. The session has no audit trail. Disputes months later cannot be resolved.
- Assuming calibration fixes bias automatically. Poorly facilitated sessions can entrench dominance dynamics and centrality bias. Design the session to address bias, not just assume the process does it.
Do You Need Calibration Software?
Yes, in most cases. Running one small calibration session on a spreadsheet is manageable, but the outlier detection, distribution charts, and audit trail that make calibration defensible get harder to maintain by hand as the group grows. Top-of-the-line performance management and review software often builds in extensive calibration features that handle this work automatically.
Our performance management software roundup covers which platforms include calibration functionality and how to evaluate them, and our comparison pages let you evaluate platforms against your requirements. If you're assessing options, our individual platform reviews include calibration feature coverage alongside pricing and implementation details.
Look for AI-powered HR software in particular, since that layer flags distribution anomalies and drafts the pre-meeting data pack automatically, cutting real prep time from the HR Business Partner's plate without replacing the facilitated discussion itself.
Frequently Asked Questions (FAQs)
How long should a performance calibration session take?
Plan meeting time by employee segment and group complexity: roughly 3 minutes for straightforward core-performer reviews, 5 minutes for top-performer discussions, and 8 minutes or more for below-expectations or contested cases. For overall group sizing: up to 15 employees runs 90 to 120 minutes; 16 to 30 employees runs 2.5 to 3.5 hours with one break; 31 or more should be split by level or role family. Add 20% for a first-time group.
Should calibration use a forced distribution or a bell curve?
Forced distribution and bell curve targets are calibration's most contested design decisions. Proponents say they prevent rating inflation; critics say they produce arbitrary outcomes when the actual distribution doesn't match the target. If you use a guideline, treat it as a flag for discussion, not a quota to hit at the cost of accuracy, and document any deviation.
What is the difference between performance calibration and talent calibration?
Performance calibration aligns employee ratings for consistency across a review cycle. Talent calibration identifies high-potential employees, succession candidates, and development priorities, typically using the nine-box grid. The two often run in the same cycle but ask different questions: performance calibration asks 'are our ratings consistent?' and talent calibration asks 'who are we developing, and into what?'
How do you run calibration for a remote or distributed team?
The process is the same; the logistics require deliberate design. Use a shared screen throughout, encourage video participation where appropriate while providing accessible alternatives for those who cannot use video, and manage the speaking floor more actively than in person. Use breakout rooms for large groups with a designated note-taker and structured report-back. Avoid asynchronous calibration for contested ratings: real-time discussion and documentation are required to resolve challenges.
Do employees find out their rating was changed during calibration?
Managers generally should not share other employees' ratings or confidential session details. They should explain the employee's final rating and the evidence supporting it.

.avif)

