Calibration

I remember the year I was expecting my best performance review. Not because someone told me I was exceptional. Because the evidence was everywhere.

Why I Stopped Believing in Calibration

A software engineer’s story (A common story this time of year. Character is fictitious)

By every metric I could see, it had been a great year.

  • I had shipped every major feature on my roadmap.
  • Production incidents were down.
  • Customer support tickets related to my services had dropped significantly.
  • My pull requests were getting merged faster than ever.
  • I had mentored two junior engineers who were now contributing independently.
  • Architecture proposals I wrote had been adopted across multiple teams.

So when review season arrived, I wasn’t worried. I assumed performance reviews worked the same way engineering did.

Measure the work. Inspect the outcomes. Learn from the results. Improve the system.

Simple.

I was wrong.


The Meeting I Wasn’t In

Review season had started in full swing. I was feeling great about my chances of a promotion or at least recognized for the work that I have been doing.

However, “the meeting” took place. A calibration meeting. A room full of managers discussing people. Not projects, metrics, cusomer impacts, or engineering outcomes, were being discussed. Subjective conversations about about reputations. It felt like a Hamiltonian moment of “In the room where it happened”. I won’t know because I was not in the room, only my reputation and a representative of that.

Knowing my manager, I know he was doing his best to represent. I also knew that odds were stacked against other managers that didn’t have the same ethical standards, integrity, and sense of fairness that my manager did (but that is another story).


Wait… What Am I Being Measured Against?

Several weeks after submitting my self-review, came the annual discussion. As the conversation unfolded, confusion started replacing confidence.

  • I knew the goals I had committed to.

  • I knew the systems I owned.

  • I knew the results I delivered.

What I didn’t know was this: What was the actual benchmark?

Was I being evaluated against:

  • My commitments?
  • Business outcomes?
  • Engineering excellence?
  • Team impact?

Or was I being evaluated against other engineers?


Metrics Versus Comparisons

As engineers, we are trained to value measurement.

  • We build dashboards.

  • We create alerts.

  • We define objectives.

  • We use data because feelings alone are unreliable.

Yet calibration often introduces a different model. Instead of asking:

“Did the engineer achieve the expected outcomes?”

The question quietly becomes:

“How does this engineer compare with everyone else?”

That sounds harmless. Until you realize those are entirely different systems. In one system, everyone can succeed. In the other system, success becomes scarce by design.


The Strange Problem with Relative Performance

Imagine telling a development team:

  • Every service met its uptime goals.
  • Every sprint objective was achieved.
  • Every customer commitment was delivered.

But only a few teams can be considered successful. An engineer would immediately identify the flaw. The results and the ratings no longer align. Yet calibration sometimes creates exactly this situation.

People who met expectations are forced into different categories because the process requires a spread. The measurement moves from:

Did you perform well?

to

Did you perform better than someone else?


Objective Data Meets Subjective Opinion

This is where things become uncomfortable.

Some parts of engineering are measurable.

  • Deployment frequency
  • Reliability
  • Availability
  • Performance improvements
  • Customer outcomes

These are quantitative. They are visible. They can be debated using evidence. Other discussions are different.

  • Leadership presence
  • Strategic thinking
  • Executive readiness
  • Influence
  • Visibility

These matter too. But they’re harder to measure. And unlike latency graphs or uptime reports, two people can look at the same engineer and reach entirely different conclusions.

That’s when engineers start asking difficult questions.Not because they reject feedback. Because they reject ambiguity.


An Engineering Mindset

We are not “Brains In Jars”. We look for concrete values to measure us. A performance system should resemble a well-designed system. Engineers are particularly sensitive to poorly designed systems. We look for:

Clear inputs. Clear expectations. Clear outputs.

If engineers don’t know:

  • What success looks like
  • How success is measured
  • Which measurements matter most

I stopped caring about ratings

Over time, I stopped caring about ratings and started focusing on different questions.

My measurements are no longer concrete, they are becoming relational. I am measuring soft-skills. I am measuring empathy. I am measuring commitment and passion. I am mesauring squishiness because it is more than a number. Questions that align more closely with the PLACES Framework I have created. Here are some areas where I am reflecting

Planning

  • Did I understand the problem?
  • Did I prioritize the right work?

Learning

  • Did I grow as an engineer?
  • Did I improve my craft?

Attitude

  • Did I take ownership?
  • Did I respond constructively when things went wrong?

Communication

  • Did I create clarity?
  • Did I help others succeed?

Execution

  • Did I deliver?
  • Did I finish what I started?

Service

  • Did I create value for customers, teammates, and the business?

These questions help me become a better engineer.

Comparing myself against another engineer rarely does.


The Real Question

Today, when review season arrives, I don’t ask:

“Where do I rank?”

I ask:

“Did I leave the system better than I found it?”

Did customers benefit? Did teammates grow? Did the software improve? Did the business move forward?

Because great engineering has never been about defeating the engineer sitting next to you. It has always been about creating something valuable together. And perhaps that’s why calibration feels uncomfortable to so many engineers.

We’re taught to optimize systems. Calibration often asks us to optimize comparisons. Those are not the same thing.


Questions Worth Reflecting On

  • Do employees know exactly what they are measured against?
  • Are metrics weighted more heavily than opinions?
  • Can everyone achieve excellence, or is excellence artificially limited?
  • Is performance evaluated against a standard or against peers?
  • Does the review process encourage growth or competition?
  • If the process were a software system, would engineers trust its design?

These questions may reveal whether a performance system is helping people do their best work or merely helping organizations rank them.