Measuring the Developer
Most competency models fail the same way - everything matters equally, so nothing does. The case for a two-level weighting system, and a first look at the two mirrored frameworks this series unpacks.
26 articles filed under this topic.
Most competency models fail the same way - everything matters equally, so nothing does. The case for a two-level weighting system, and a first look at the two mirrored frameworks this series unpacks.
Dashboards go green while codebases rot - the opener of a fourteen-part series on measuring code quality and software delivery without staging the theatre that ruins both.
The series retrospective - what the AI-era evidence actually says with every figure dated and attributed, the nine characteristics of ISO 25010 as revised in 2023, and what I now gate differently when much of the diff was machine-written.
Eleven numbers that still haunt engineering dashboards - what killed each one, from Goodhart to queueing theory to changing fashion, and the one-line replacement carved on every headstone.
The team, the team lead and the executives need different slices of the same truth - what each audience should actually see, the numbers I put in front of a boardroom, and dashboard principles that skip the red-amber-green pantomime.
What Azure DevOps will measure for you and where it stops - the two clock definitions, the reactivated-item rule, how to read a cumulative flow diagram, and why Microsoft's own guidance points past its widgets to OData.
Why no single metric captures developer productivity - the SPACE framework's five dimensions in plain terms, the measure-at-least-three rule, and what I do about the two dimensions no dashboard can hold.
The queue mathematics underneath delivery - one equation from 1961, Scaled Agile's six flow metrics rewritten in plain terms, and why the biggest delivery gains hide in the waiting rather than the working.
DORA as it actually stands - five metrics since the 2024 report, benchmarks dated to their vintage, the 2025 archetypes, and the comparison warning printed on the tin.
Vulnerability density, severity-tiered remediation clocks and hotspot review coverage - running the findings backlog as an operational queue with owners and deadlines, not an annual scare.
Code churn, defect density and hotspot scoring - reading a repository's git history to find the files most likely to break next, and the thresholds I use to decide when movement has become rework.
Line, branch and condition coverage each assert less than you think - mutation testing is the honesty check that reveals whether your suite would actually notice a bug, and the thresholds worth enforcing are conventions, not laws.
Measuring the lines between the boxes - afferent and efferent coupling, the instability ratio, Robert C. Martin's main sequence, and the cohesion metric that catches a class doing two jobs, plus where every one of them needs a human to overrule it.
The composite score that hides more than it shows and the percentage that translates code quality into money - where the maintainability index formula came from, why its logarithms flatten real problems, and how the technical debt ratio earns its place as a trend.
Two complexity metrics that look interchangeable and answer different questions - McCabe's 1976 path count is a testing instrument, SonarSource's cognitive score is a readability instrument, and only one of them earns a place in the merge gate.
Dashboards go green while codebases rot - the opener of a fourteen-part series on measuring code quality and software delivery without staging the theatre that ruins both.
The honest retrospective on two weighted competency frameworks - the blank templates at their centre, the qualities no weight can hold, the situations where a weighted model is the wrong tool, and what I would build differently.
How the framework becomes a form and the form becomes a plan - per-sub-discipline ratings with evidence, a 500-point ceiling, a gap-times-weight priority formula and a 90-day goal loop, walked through with a worked example.
The four disciplines that operate across my AI competency stack rather than inside it - agent architecture, evaluation, governance and domain translation - and why 37 of the framework's 100 points live outside the core.
The four core disciplines of my AI engineering framework and the argument behind their weights - why the prompt sits at the bottom of the stack at 8 points, why the specification sits at the apex at 22, and the sentence in my own document that its own table contradicts.
The transformation map between my two frameworks - what each of the seven traditional disciplines became among the AI-era eight, what split, what transferred with its weight intact, what dissolved into sub-disciplines, and the fifth learning-stack discipline that is deliberately not on the list.
Two research results explain why self-ratings drift - METR's randomised trial of AI-assisted developers and Kruger and Dunning's calibration studies - plus the evidence fields, peer calibration notes and built-in warnings my templates use to correct for both.
Novice to Architect - the five-level ladder both of my frameworks share, why every cell of it demands evidence you can point at rather than a feeling about ability, and the Level 0 my own summary tables invented by mistake.
Leadership weighs 15 in my competency framework and not one point of it requires authority - decision making, alignment, mentoring, process thinking and facilitation as evidenced skills, plus the two quiet disciplines that complete the model.
Feedback weighs 13 and Collaboration weighs 12 - together they match Engineering Craft point for point, and this part defends that arithmetic sub-discipline by sub-discipline, from effective communication to handling disagreement.
Delivery carries 18 points in the framework and none of them is a deadline - incremental value, work breakdown, prioritisation and ambiguity as four trainable skills that turn engineering capability into shipped outcomes.
Engineering Craft carries 25 of the framework's 100 points and splits into nine sub-disciplines - why writing code earns only 18 of them, why architecture is the most leveraged skill in the model, and why reading code overtakes writing it by the time seniority arrives.