Evergreen note
Will It Be Alright on the Night?
The question everyone actually means when they ask "how's it going" isn't are we there yet. It's will we get there on time.
The title comes from Denis Norden's It'll Be Alright on the Night, decades of blooper-reel television built on the reassuring idea that however much goes wrong backstage, the show somehow comes together anyway. The question that actually matters in delivery is the less comforting version. Given what we know now, will it actually be alright on the night?
The problem with a status colour
Most programmes carry several legitimate but disconnected signals: milestone dates, commercial RAG, sprint progress, a risk register, and eventually some evidence that the thing went live and got used. None of those signals on its own answers the question anyone actually wants answered, which is whether the outcome is likely to land by the point it matters.
A single RAG rating tries to compress all of it into one colour, set once and then defended. The colour usually tracks comfort rather than evidence, and green stays green until something forces the conversation.
The Infrastructure and Projects Authority, which tracks the UK's major government projects, has a proper definition for this. A Delivery Confidence Assessment is:
"An assessment of the likelihood of a project delivering its objectives to time and cost." (IPA Annual Report 2023-24)
Since 2021 the IPA has used three bands only: green, amber, red. It used to run a wider scale with hybrid amber/green and amber/red ratings, and dropped them, because in practice more bands blurred exactly the judgement they were meant to sharpen. An organisation whose entire job is tracking delivery confidence at scale concluded that a longer scale makes the call easier to dodge, not easier to make.
Where I started
My first attempt at this was a six-factor weighted model: outcome clarity, plan credibility, delivery evidence, gate readiness, capacity and ownership, and residual risk, each scored and weighted into one number. I also built in a second point. A score means something different depending on how much you actually know when you set it. An early guess and a late-stage assessment backed by live data shouldn't be dressed up as the same kind of confidence, even when the score is identical.
There's some truth in both of those. The model itself wasn't wrong either, but it asked a lot of anyone using it consistently, six judgement calls, made fresh, at every review, for every milestone. In practice that meant it either didn't get used properly, or one person kept all six factors in their head and did it themselves, which defeats the point of writing any of it down.
Two inputs, not six
What I use now fits in one sentence:
Progress against plan, plus categorised risk and issue exposure, equals delivery confidence, expressed as a number and a direction of travel.
Progress against plan is the easy half. It's just where the work actually stands against what was committed. Risk and issue exposure is where confidence models usually fall apart, because not all risks are the same, and lumping them into one register with one number hides which kind you're actually looking at.
HM Treasury's Orange Book, the standard government guidance on risk management, splits that into two separate questions: what type of risk is it, and what's actually being done about it. A risk gets typed by its source: governance, a dependency on someone else, people and capacity, a technical blocker, or something genuinely outside anyone's control. Separately it gets categorised by response: being actively treated, deliberately retained and monitored because there's nothing more useful to do right now, or shared with a third party who owns that exposure contractually.
That split matters more than it sounds. A risk outside anyone's control and correctly being watched shouldn't drag a score down the same way an overdue internal decision that's supposedly being actioned but isn't moving should. Applying the same fixed rule every time, rather than a fresh judgement each review, is also the honest answer to a temptation every delivery leader recognises, to make things look a little better or a little worse than they are, depending on what someone needs to hear that week. A rule that doesn't move with who's presenting it removes that temptation rather than relying on everyone's good faith.
The evidence-maturity point from the first version still applies here. A progress score built on live delivery data carries more weight than one built on an early plan, even when the two numbers happen to match.
Where the score sits
Confidence gets calculated once, at the level where the evidence actually is, then rolled up using the same rule at every step above it.
Weight by contribution, not effort. When several deliverables roll up into a milestone or an outcome, the temptation is to weight them by size, cost, or ticket count. That's the wrong axis. The real question is how much each one genuinely contributes to the outcome, and whether the outcome could still land without it.
Don't average away a failed gate. A weighted average is a very effective way of hiding the one thing that makes an outcome impossible. Ten healthy supporting workstreams cannot compensate for a missing approval, a supplier who hasn't committed, or a decision that hasn't been made. An unowned, undated blocker on the critical path caps the whole score, whatever the rest of the average says.
Trend over snapshot
A single score on a single day shows where things stand. It doesn't show whether they're getting better or worse, and that's usually the more useful fact.
Amber sliding from 71 to 62 is a real regression, even though the colour hasn't moved. The same logic runs the other way. Amber climbing from 61 to 72 is real progress, even while it's still technically amber. Without the trend, both of those just read as "still amber," which hides exactly what a good delivery conversation needs to catch early.
The minimum useful statement
A naked number is nearly always misread. The smallest honest unit of a confidence statement is closer to this: score and band, evidence maturity, trend, the principal reason, and the next event that would actually change it. That's more work than a traffic light, and a more useful one.
None of this promises certainty. The honest answer to "will it be alright on the night" is usually somewhere between yes and we'll know more soon. Said clearly, with the reasons attached, that's a genuinely useful thing to hear. Far more useful than a guess dressed up as a fact.
Lesson learned
A confidence score needs its reasons shown, or it's just a guess with a number on it. Six factors taught me what mattered. Two inputs, typed properly, turned out to be enough to say it.
Sources
- IPA Annual Report 2023-24, Delivery Confidence Assessment definition and banding history
- HM Treasury, The Orange Book: Management of Risk (May 2023), risk typing and response categorisation
- Institute for Government, Major projects in government, the original spark for this piece
- This started as, and still draws on, a longer internal working paper with worked examples against live programme milestones. This is the public, principles-only version of that thinking, kept up to date as the thinking itself moves on.
This is a live line of thinking, not a finished position. It started as a six-factor model, it's a two-input one now, and it'll keep changing as it gets tested against real examples.