Skip to content
Kumar Chandrachooda
Engineering Practice

What a Scorecard Cannot See

The honest retrospective on two weighted competency frameworks - the blank templates at their centre, the qualities no weight can hold, the situations where a weighted model is the wrong tool, and what I would build differently.

By Kumar Chandrachooda 18 Apr 2026 8 min read
A lens over a scorecard, with the important part outside the frame

Part 11 closed the loop: rate, evidence, score, prioritise, commit, calibrate, repeat. Eleven parts of this series have argued for weights, defended specific numbers and built a planning mechanism on top of them. Every series on this blog ends with an honest retrospective — strengths, sharp edges, when not to use the thing, what I would do differently — and this one has more to confess than most, because the most important fact about my two competency frameworks is not in any of the tables. It is this: both self-assessment templates are, as I write this, blank.

The forms nobody has filled in

Meera, the engineer whose scorecard part 11 walked through, does not exist. I invented her to demonstrate the arithmetic, and I told you so at the time — but the deeper admission is that there was no real example to use instead. No completed assessment of a real engineer exists against either framework. No calibration session has put a second signature on a real set of ratings. No 90-day loop has gone around even once and produced the trend line I promised it would. The templates are polished, the weights are argued, the formula is sound on paper — and the framework's biggest untested assumption is the assumption that it works.

That is worth sitting with, because this series has been written in creator voice throughout, and the creator register earns its authority by owning defects. In the code-based series on this blog, the honest ledger lists shipped bugs; the research equivalent is shipped confidence. Let me itemise what is actually unverified:

  1. That evidence fields discipline ratings in practice. The design bets that an empty evidence cell will shame an inflated rating downwards. Plausible; unproven. People are resourceful about narrating themselves generously.
  2. That gap-times-weight matches judgement. The formula ranks development priorities mechanically. I believe a good mentor looking at the same engineer would broadly agree with its ordering. I have never once put the two side by side.
  3. That ninety days is the right cadence. Long enough for a rating to genuinely move, short enough to stay honest — that was the reasoning. It is a guess with a round number on it.
  4. That scores are comparable between people. The fixed 500-point ceiling only makes scores comparable if two assessors mean the same thing by a 3. Without calibration data across real pairs of raters, that is a hope, not a property.

Part 9 documented the one defect the frameworks have already yielded — a sentence in my own document that its own tables contradict, caught because writing this series forced a re-reading. That is the pattern I expect to continue: the model will be improved by being used and argued with, and it has so far been argued with far more than it has been used.

Taste, trust and timing

Suppose the loop runs perfectly. Suppose an engineer rates honestly, evidences everything, calibrates with a sceptical manager and works the weighted gaps for a year. What has the scorecard still not measured? Three things at least, and they may be the three that matter most at the senior end of a career.

Taste. Two designs both work, both scale, both pass review — and one of them is right. The faculty that tells you which is the one every strong engineer I have worked with possesses and none of them can decompose. It does not appear in my frameworks, and not for lack of trying: every candidate sub-discipline I drafted for it collapsed into either a skill I had already weighted or a tautology. You cannot rate taste on a five-point scale, because judging the rating requires exactly the faculty being rated.

Trust. A review comment from one engineer lands as an attack; the identical words from another land as a favour. The difference is years of accumulated credibility — being right when it cost something, being wrong and saying so first. Trust is real competency in the sense that matters most for influence, and it is structurally invisible to a self-assessment because it does not live in the self. It lives in other people's ledgers about you.

Timing. Knowing what to build is covered — decomposition, prioritisation, trade-off reasoning all have rows. Knowing when is not. When to ship the imperfect version, when to wait a week, when to raise the concern that will derail the sprint, when to let it go. Timing is competency applied at the right moment, and a scorecard is a photograph of capability with the clock cropped out.

I considered forcing rows for all three and decided against it, and I stand by that. A weighted model should refuse to measure what it cannot observe, because a bad number displaces attention that an honest blank would leave free. The cost of that refusal is permanent: the frameworks describe the trainable, evidencable core of the job, and the job is larger than that core.

When not to weight anything

A retrospective owes you the anti-recommendations. Situations where I would not reach for either framework:

  • Hiring. The model measures development over time against evidence accumulated in context. An interview samples neither. Repurposing the scorecard as a screening rubric invites candidates to perform the rows, and rewards articulate self-narration — the exact bias part 7 warned about — over the competencies themselves.
  • Anything wired to pay. The moment a weighted score feeds a compensation formula, every incentive tilts toward inflating ratings and negotiating targets downward, and the calibration conversation turns adversarial. The score stops informing development the day it starts pricing it.
  • A team of two or three. The overhead is not worth it. If you can see everyone's work directly and talk about it over coffee, the framework is a bureaucratic re-statement of conversations you should simply have.
  • An organisation that will not argue. Part 1 said the disagreement is the value: an arguable number beats an unarguable blank. In a culture that adopts frameworks by decree, the weights calcify into dogma — 22 becomes true rather than argued — and the model's one genuine virtue is destroyed at the door.
  • A confidence crisis. An engineer who is struggling does not need a document that renders their situation as 251 out of 500. They need a conversation. The framework is a good servant of that conversation later; it is a terrible opening move.

What I would do differently

If I were starting the pair of frameworks again, four changes.

The worked example first. I built taxonomies, weights, level definitions, progression tables and templates — the complete apparatus — before running the assessment on its most available subject: me. That was backwards. One honest, filled-in form, scars visible, would have tested more assumptions than any amount of definitional polish, and it is the obvious next artefact this series owes. The commitment is on record now, which is one of the reasons to write retrospectives.

Evidence as a habit, not an event. The templates ask for evidence at assessment time, which sends you digging through months of work for artefacts you half-remember. The better design is a running evidence log — a line per artefact as it happens, tagged by sub-discipline — that the quarterly assessment merely summarises. Ten minutes a week would make the form's hardest section nearly free.

Version the weights like code. Every number in both frameworks was argued for once, in my head, and written down. The arguments that produced 22 and 8 and 4.40 deserve a changelog: when a weight moves, record what evidence or argument moved it. A model that claims to be falsifiable should keep receipts of its falsifications.

Collect the calibration deltas. The gap between self-rating and calibrated rating is itself the most interesting data the process generates — per person, per discipline, over time. Aggregated, those deltas would show where the framework's descriptions mislead raters, and whether the overrate-the-visible pattern the templates warn about actually shows up row by row. Designing the form without designing the collection of that data was a missed trick.

And a fifth, half-held: both frameworks might be one size too detailed. Thirty-three sub-disciplines in the AI model, thirty-one in the software one. The bottom rows carry absolute weights near a single point, where the measurement noise is plausibly larger than the signal. A leaner model would give up some completeness for more honesty per row. I have not made that cut, because every row still earns its argument — but I notice that defending completeness is exactly what the committee-built matrices in part 1 did.

The argument was the product

So what survives the confession? More than it might seem. The frameworks made me commit, in public, to what I believe matters in this profession and by how much — twelve parts of arguable numbers instead of a grid of equal blanks. The weighting mechanism is sound arithmetic wrapped around a genuinely useful discipline: say what you value, in proportion, and let people push back. The stack-versus-weight distinction, the evidence rule, the priority formula — those I will defend as they stand. The blank templates I can only promise to fill.

If you have read all twelve parts, you know the model's numbers better than most engineers know their own competency matrix — and you know exactly where its author has and has not earned your trust. That was the design. A model that refuses to decide cannot guide anyone, and a model that hides its untested assumptions cannot be trusted when it decides. This series tried to be the first kind of model, honestly. If you are arriving here first, the argument starts where the series did: why weight a competency model at all.