Skip to content
Kumar Chandrachooda
Engineering Practice

Your Self-Assessment Is Lying to You

Two research results explain why self-ratings drift - METR's randomised trial of AI-assisted developers and Kruger and Dunning's calibration studies - plus the evidence fields, peer calibration notes and built-in warnings my templates use to correct for both.

By Kumar Chandrachooda 08 Apr 2026 8 min read
A mirror reporting a taller figure than the one standing before it

The evidence tables in part 6 exist because of an uncomfortable fact: in any self-assessment, the person holding the pen is the least reliable observer in the room. That is not a character judgement — it is a measured property of human calibration, and in AI-assisted engineering it has been measured with unusual precision. This part rests on two pieces of external research, both credited and linked, and then turns to what I actually built into the assessment templates to compensate. The frameworks are mine; the evidence that self-perception drifts is emphatically not, and it deserves its named sources.

Sixteen developers, two hundred and forty-six issues

The first pillar is METR's randomised controlled trial, published in July 2025 under the title “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity”. Sixteen experienced open-source developers completed 246 real issues in large, mature repositories they knew well. Before the study, the developers forecast that AI would speed them up by 24 per cent. When allowed to use early-2025 AI tools — primarily Cursor Pro with Claude 3.5/3.7 Sonnet — they in fact took 19 per cent longer to complete their tasks. And even after finishing, they still believed AI had sped them up by about 20 per cent.

Read those numbers slowly, because they are three distinct measurements taken at three distinct moments. Plus 24 per cent is the forecast, made before any work happened. Nineteen per cent longer is the measured result, from the stopwatch. Plus 20 per cent is the post-hoc belief, reported after the work was done and the slowdown had already been experienced. The forecast was wrong, which is forgivable — forecasts usually are. The devastating finding is the third number: living through the slowdown barely moved the needle. Perception did not merely lag reality; it pointed the other way, before and after.

METR's own caveats travel with the result, and I will not strip them. The participants were experienced developers working in repositories they knew intimately, using early-2025 tools, and METR itself cautions against generalising the headline number beyond that setting — tooling has moved a long way since. What generalises, for my purposes, is not the 19 but the shape: a gap between measured performance and perceived performance that survives direct experience. If experienced professionals can misjudge the direction of their own productivity while it is being measured, an unaided rating of “I'm about a Level 3 in Context Engineering” deserves exactly as much trust as you would place in any other unverified instrument.

The warning label I wrote wrong

Here is this part's entry in the honest ledger. My AI self-assessment template has carried a warning block about this study since version 1, and version 1 states it wrongly. It says developers “predicted AI made them 24% faster while actually being 19% slower — and they still believed it after the study ended”. The sentence welds the forecast to the aftermath: it implies the developers still believed the 24 per cent. They did not — the post-study belief was about 20 per cent, a different number from a different moment. My template also says “19% slower” where METR's framing is “took 19% longer”, and those are not the same arithmetic.

I am documenting the error rather than silently patching it because it is almost too on-the-nose: the number I deployed to warn people about miscalibration was itself miscalibrated. I compressed three measurements into two and produced a sentence that was punchier than the truth. The corrected version is, if anything, worse news for self-assessment — the belief did not stubbornly persist unchanged, it was re-formed after the experience and still landed on the wrong side of zero. But the corrected version is what the study found, and a framework built on “ratings without evidence should be challenged” has to hold its own citations to the same standard.

Confidence pools where skill is thinnest

The second pillar is older and broader. Kruger and Dunning's classic 1999 studies — “Unskilled and Unaware of It”, published in the Journal of Personality and Social Psychology — found that the least-skilled overestimate their ability most, while top performers slightly underestimate theirs. A pattern, I suspect, that every calibration exercise in engineering will recognise. Two honest qualifications before I use it. First, the magnitude of the effect is debated in the replication literature — how much of the classic curve is psychology and how much is statistics is a live argument — so I lean on the direction of the asymmetry, not its size. Second, a confession: the warning block in my software progression document claims the effect is “well-documented in software engineering”, junior engineers overestimating and seniors underestimating. I could not trace a software-specific literature that says that. Kruger and Dunning studied general populations; the application to engineering careers is my extrapolation, and it should have been labelled as one. Second scar, same ledger.

What makes the asymmetry structurally nasty for competency models is its mechanism: assessing your skill in a discipline is an exercise of that discipline. Recognising that your specification was ambiguous requires exactly the specification-reading skill you lack at Level 1. The practitioners with the largest gaps are the least equipped to see them — the error is not noise scattered evenly across the scorecard, it concentrates where the scorecard matters most.

My design assumption — stated as my expectation, not a research finding — is that in AI engineering this concentrates along visibility lines. Prompt Craft is high-visibility: you see the output seconds after the instruction, so your self-model updates constantly and tends to flatter. Specification and Context Engineering are high-leverage but low-visibility: their failures surface later, somewhere else, wearing the model's name — the agent “got it wrong”, not “my context was missing the schema”. So I expect over-rating to pool at the bottom of the stack and under-recognised gaps to pool at the top, which is the worst possible distribution, because the top is where the weights are.

Designing for an unreliable witness

Given all that, the naive conclusion is to abandon self-assessment. I kept it — but I built the templates on the assumption that the assessor is an unreliable witness, and wired in four correctives.

Evidence fields on every row. Each sub-discipline rating sits next to a mandatory evidence cell, and the instruction is explicit: concrete artefacts, not assertions. “Designed the context architecture used across three projects”, not “good at context engineering”. A rating with an empty evidence cell is not a low-confidence rating; it is a hypothesis awaiting rejection.

Reflection prompts that demand episodes. Every discipline section ends with questions of the form “when did you last…” — when did you last identify an agent optimising for the wrong objective, when did a domain-specific exception produce wrong output and what did you do. These questions cannot be answered with an opinion. Either a specific memory exists or the silence is informative.

Calibration notes. The template ends with a section that does not belong to the assessor at all: a table where a peer or manager records a calibrated rating beside each self-rating. The deliverable is the diff. Where the two columns agree, move on; where they diverge, something — the self-model or the peer's visibility — is wrong, and finding out which is the actual assessment.

The warning printed into the form. The METR result and the Kruger–Dunning asymmetry are quoted inside the template itself, positioned so the assessor reads them at the moment of rating, not in a preface nobody revisits. (This is the block I corrected above — the corrective needed correcting.)

Above all four sits the arbitration rule: where self-assessment and peer assessment diverge, the artefacts arbitrate. A developer who rates themselves Level 3 in Specification Engineering while their specifications routinely need agent re-runs and human correction is not Level 3, regardless of what they believe — because part 6's definition of that level is specifications complete enough to produce correct output on first execution, and that is checkable against execution history, not testimony.

A starting point, not a verdict

Why keep the self-assessment at all, given everything above? Because it does two jobs nothing else does. For the individual, it forces structured reflection across every discipline — including the ones they would never think to rate themselves on, which are reliably the interesting ones. And for a team, the aggregate is a strategic instrument even when every individual number is wrong: the gap between a team's perception of its profile and its artefact-based reality is itself a finding. A team that believes it is strong at specification while its specifications keep failing needs a different intervention from a team that knows its gap and lacks the practice — the first needs its mirror fixed before any training will land.

The point of a self-assessment is not the score; it is the diff between the score and the evidence. The score alone is testimony from a witness two research programmes have impeached. The diff is where development actually starts.

Both templates — the seven-discipline one and the eight — share this machinery, which raises the question this series has been circling since part 1: how did seven disciplines turn into eight in the first place? Next, the hinge of the whole arc: seven disciplines become eight.