Skip to content

09 · Measuring Leadership Effectiveness

Everything so far in this track asks you to trust your own read of whether it's working — did that conversation land, is the team actually safer, is delegation actually sticking. Your own read is biased in a predictable direction: leaders systematically overrate their own approachability, fairness, and psychological safety relative to how their team rates them. This module gives you concrete, gatherable signals for whether your leadership is actually working, beyond how it feels from where you sit.

1. Why gut feel is the wrong instrument

Three structural reasons your own sense of "how am I doing" is unreliable: people self-censor more in front of you than with each other, so you get a filtered sample of the real feedback; recency bias means one great conversation this week outweighs a slow accumulation of small frustrations you never heard about; and status makes disagreement costly for the other person, so silence reads as agreement when it might be caution. None of this means your judgment is worthless — it means it needs external signals to check against, the same way a pilot doesn't fly on feel alone once instruments are available.

2. The Leadership Signal Set

A set of concrete, gatherable indicators, grouped by how hard they are to fake.

LEADERSHIP SIGNAL SET

HARD TO FAKE (trust these most)
  Regrettable attrition — people you wanted to keep, leaving
  Internal referral rate — do people recommend friends join your team?
  Who asks to transfer ONTO vs. OFF your team when they have a choice
  Anonymous engagement survey trend over time (not one snapshot)
  Whether bad news reaches you before it's already a crisis

MODERATELY RELIABLE
  Skip-level conversations (your manager or their manager talking
    to your reports without you in the room)
  360 feedback, if genuinely anonymous and aggregated (n≥4 minimum
    to protect anonymity, or individual comments become identifiable)
  Promotion/growth rate of your direct reports vs. peer teams

EASY TO GAME OR MISLEADING ALONE
  Direct 1:1 feedback given straight to your face
  Team meeting energy/mood as you perceive it in the room
  Your own retrospective sense of "the team seems happy"

The ranking isn't about ignoring the bottom tier — direct feedback still matters — it's about weighting your confidence correctly. If the top-tier signals and the bottom-tier signals disagree, believe the top tier.

3. Worked example: the gap between self-perception and signal

Oscar believes he runs an open, low-hierarchy team — he takes pride in his door-open policy and gets warm feedback in 1:1s. His skip-level report from his own manager, gathered from Oscar's team without him present, tells a different story.

Oscar's manager, Divya (relaying anonymized themes): Three separate people brought up the same thing, independently — that disagreeing with you in the team meeting has a cost. Not that you're harsh. That you go quiet and re-litigate the decision one-on-one with whoever disagreed, right after the meeting.

Oscar: I didn't realize I was doing that. I thought I was just... following up to understand their concern better.

Divya: From your side, sure. From theirs, disagreeing in the room means a private conversation afterward that feels like a consequence, even if you don't intend it that way. So now people just don't disagree in the room. Which is probably why your 1:1s feel so smooth — the disagreement's been filtered out before it reaches you.

Oscar: That explains something. My team meetings have felt very agreeable for months, and I read that as alignment.

Divya: It might be alignment. It also might be exactly this pattern. The way to find out is to stop following up privately for a month after someone disagrees publicly, and see if disagreement starts happening in the room instead.

Oscar's direct feedback (bottom tier — warm 1:1s) told him one story; the skip-level (moderately reliable, aggregated, gathered without him present) told him a more accurate one, and it only surfaced because a structure existed for it to reach him at all.

4. Building the instruments, not just waiting for signal

Most of the Signal Set doesn't arrive on its own — you have to build the channel:

  • Request a skip-level from your own manager on a cadence (quarterly is reasonable), explicitly asking them to probe for what people wouldn't say to you directly.
  • Run 360s with a real anonymity floor. Below about four respondents, individual comments become identifiable by writing style or specifics — either aggregate further or don't run it at that group size.
  • Track exit interview themes over time, not just the individual exit reason, which is often diplomatically softened. A pattern across several regrettable departures is a signal even when any single one looks like "better opportunity."

5. The trap of optimizing the metric instead of the thing

Once you start measuring, the temptation is to manage the number instead of the underlying reality — pushing for higher engagement survey scores directly, rather than fixing what the low scores were describing. Any metric in section 2 that becomes a target stops being a reliable measure (Goodhart's Law). The discipline: treat every signal as a prompt to go find the real story behind it, the way Oscar had to have an actual conversation about his own behavior, not just log the finding and move on.

How It Actually Works

Gut feel is an unreliable instrument for measuring one's own leadership effectiveness for a specific, well-documented reason: self-assessment suffers from the same motivated-reasoning bias covered in the Level 1 project module, compounded by an availability problem — a leader's impression of "how I'm doing" is built disproportionately from the interactions they were most present and attentive for, which systematically excludes the moments (a report's private frustration, a quiet erosion of trust) they weren't in the room for by definition. A leader cannot observe their own blind spots directly — that's what makes them blind spots — so any purely internal measurement instrument is structurally incapable of detecting the exact failures most worth catching.

Why a designed "signal set" (specific, externally-sourced indicators) outperforms impression. Externally sourced signals — retention among high performers specifically, upward feedback, whether people bring problems early or only after they've escalated — are less subject to the leader's own motivated reasoning because they're generated by other people's independent behavior, not the leader's self-report. This mirrors why 360-degree feedback and objective retention data are used in leadership research instead of self-rating alone: aggregating independent, externally generated signals reduces the single-source bias any one measurement (especially self-assessment) carries.

Why optimizing the metric instead of the underlying thing it measures is a near-guaranteed failure mode, not just a risk. Any proxy metric is, by construction, an imperfect stand-in for the actual outcome it's meant to represent (this is Goodhart's Law: when a measure becomes a target, it ceases to be a good measure). If a leader starts managing to the specific number — pushing retention up through counter-offers rather than genuine engagement, or manufacturing positive feedback scores through subtle pressure — the gap between the metric and the real underlying leadership quality widens precisely because effort shifted from the target outcome to the measurement of it, which is why the module treats the signal set as diagnostic input to be interpreted, not a scoreboard to be maximized directly.

Exercise

Pick one hard-to-fake signal from section 2 that you are not currently tracking — regrettable attrition, referral rate, or a skip-level cadence — and set it up this month. If you already get 1:1 feedback that feels uniformly positive, treat that as a prompt to check it against a harder signal rather than as reassurance, the way Divya's skip-level surfaced something Oscar's own 1:1s had filtered out.