Evergreen note

A Mirror That Argues Back

7 min read

Self-assessment, left unchecked, tends to be flattering. Not dishonest, just generous. We remember our best week and quietly forget the three months either side of it.


The Problem With Frameworks

Every capability framework has the same weakness: it tells you what "good" looks like, but not how well you're actually doing against it. That gap gets filled by self-assessment, and self-assessment is a genre of fiction most of us write flatteringly, without realising we're doing it.

I ran into this doing competency mapping as part of one of my regular objectives, checking myself against a framework rather than assuming. It's easy to spend that kind of exercise doubling down on what you're already good at, because that's comfortable and the evidence is easy to find. The harder, more useful version is finding the actual gaps and building the evidence to close them. I had a framework, the DDaT capability framework, which sets out what "good" looks like for delivery roles, level by level. I had a rough sense of where I sat on it. What I didn't have was a repeatable process for turning "I think I'm doing okay at this" into something I'd actually stand behind. This quarter, that's what I've been working through. I haven't finished the exercise, but I've landed on the approach: pick the skill, write the evidence properly, then get something other than my own judgement to argue against it.

I broke it into three steps.


1. Pick the Framework, Pick the Skill

Start with whatever capability framework applies: DDaT for government digital delivery, your own organisation's competency framework, or just a role profile if that's all there is. Don't assess yourself against the whole thing at once. Pick two or three skill areas that matter for where you're actually heading (a promotion case, an appraisal, an application) and decide, honestly, which level you think you're at for each.


2. Write the Evidence Properly: STAR, Not Vibes

"I'm good at stakeholder management" isn't evidence. It's an opinion wearing evidence's coat.

STAR (Situation, Task, Action, Result) forces the claim to anchor in something that actually happened: what the situation was, what you were trying to achieve, what you specifically did (not "we"), and what changed as a result, ideally with a number attached. If you want the method itself explained simply first, the National Careers Service's guide to the STAR method is a clean two-minute read with worked examples. GOV.UK's brief guide to competencies is a five-minute refresher if it's been a while since a job application forced you to use it.

One full STAR paragraph per skill area. Not a bullet point.


3. Get Uncomfortable: Ask AI to Argue Against You

This is the step that turns "self-assessment" into an actual reality check. Take the STAR example and the level you've claimed, and feed both to Copilot or ChatGPT:

"Here is my evidence for [skill area], and I've assessed myself at [level]. Act as a devil's advocate. Challenge this evidence against the level I've claimed: what's missing, what am I assuming the reader will infer, where's the result vague rather than specific, and what would a sceptical reviewer ask that I haven't answered here?"

It will not be kind. That's the point: a model trained to agree needs an explicit instruction to stop, and once given one, it's good at finding the gaps, the result asserted but never measured, the "I" that was actually a "we," the level claimed that the evidence only supports one tier down.


A Worked Example (Mine)

Easier to show this than describe it, so here's an actual pass through it, using my own evidence rather than a hypothetical.

Role and level: I'm a Head of Delivery. This is the role I'm doing day to day, and that's exactly why the check matters: it's easy to assume you're meeting the standard for a role you're already in without ever actually testing it. The DDaT framework's top tier for this track is Head of (Agile) Delivery Management. One of the skill areas at that level is Planning, which at Expert level means you "lead a continual planning process in a very complex environment, plan beyond product delivery, identify dependencies in plans across services and co-ordinate delivery, coach other teams as the central point of expertise."

My claim: Expert level, on the strength of leading a quarter's planning cycle.

The STAR example:

Situation: Quarterly planning needed to pull together multiple workstreams into one coherent view, with real dependencies between them that hadn't previously been mapped in one place. Task: Own the planning process end to end: structure it, get the right people looking at the right questions, and produce something that showed not just what each stream was doing but how they depended on each other. Action: Defined the prep structure for the cycle, aligned cluster leadership on it, built the deep dives for each priority area, and pulled every lead's input into a single integrated planning view rather than leaving it as separate documents. Result: The planning was independently and publicly praised. The Deputy Director wrote in their own weeknote afterwards: "Appointments quarterly planning was a strong step forward... it was good to see the level of planning that had gone into it, and the thinking on dependencies and value... Thanks to Sean and others who are working hard to keep teams aligned."

The AI pushback, run for real rather than simulated:

This is stronger evidence than most people bring to a self-assessment, and it's worth saying why: the Result isn't self-reported. The Deputy Director's quote is unprompted, dated, and specific about what they saw, "the thinking on dependencies and value," which lines up almost exactly with the Expert descriptor's own language. That's rare. Most STAR evidence has to ask the reader to trust the writer's account of the result; this doesn't. But read the descriptor again in full: "lead a continual planning process... coach other teams as the central point of expertise." The evidence proves the first half convincingly and doesn't touch the second. What's shown is that you personally ran one planning cycle well. What Expert actually asks for is evidence that the capability now lives beyond you: that cluster leadership could run the next cycle to the same standard without you holding the structure together. Is there anything in the log showing that handover starting, or is the "central point of expertise" still just you? If it's still just you, this is excellent evidence of doing Expert-level work, not yet of building Expert-level capability elsewhere, and a panel that knows the framework will ask the difference on purpose.

So the honest read isn't "not Expert." It's specific: the planning quality is proven, external, and dated. What's missing is evidence of transfer, and closing that gap is what I'm working on this quarter. My own part in quarterly planning is shifting too: from running the structure myself to owning the outcome, plan quality and alignment across the piece, while cluster leadership runs the process directly. Not a finished case, just the approach I've landed on: hand the running of the next cycle to cluster leadership, and write down what that actually looks like against the same standard.


Why the AI Step, Specifically

A good colleague or coach does this better. They know the context, and they can push with a nuance AI can't. But they're busy, and asking someone to seriously stress-test your self-assessment is a bigger favour than it sounds. The AI version is available at 11pm the night before, doesn't tire of your evidence, and has no incentive to be nice. Use it as the cheap first pass; take what survives to the actual human for the real conversation.


The Test

Try it once, on the piece of evidence you're most confident about. If it comes back clean, that's a good sign. If it doesn't, better to find out from a chatbot than in the room.