The Ember Codex
The wrong-instrument null: why a client improves and the symptom score says nothing
A 63-trial review found behavioural work moves function, not symptoms, with the gains still there at three years. Every trial in it enrolled 10-to-19-year-olds.
Behavioural work for ADHD, cognitive-behavioural therapy included, moves function, not symptom counts. The biggest review of it, by Sibley and colleagues in the Journal of Child Psychology and Psychiatry, pooled 63 randomised trials. The review found “inconsistent impact on ADHD symptoms but strong and consistent impact on impairment and executive function skills.” The gains were still there at three years.
I wanted that finding to be about adults. Every participant in those 63 trials was aged 10 to 19.
What did the 63-trial review find?
A split. Across 63 randomised trials, each finding graded on the GRADE system for how much weight it can take, the symptom measures and the function measures did not move together.
| What was measured | What the 63 trials showed |
|---|---|
| ADHD symptoms | inconsistent impact |
| Impairment and executive function skills | strong and consistent impact, maintained up to 3 years |
The age band is not in the sentence everyone quotes, which is how an adolescent result ends up in copy written for adults. Sibley’s review also recommends that medication and behavioural work start together. So the three-year line is a poor argument for skipping the prescriber, and that misreading bothers me more than the age band does. It puts words in an author’s mouth.
What is the wrong-instrument null?
The wrong-instrument null is a session where the client’s week changed and the score did not, so the write-up says no progress. It is common. The 2025 Lancet Psychiatry network meta-analysis of 113 trials and 14,887 adults shows the two outcomes coming apart at scale: stimulants and atomoxetine cut core symptoms, and no medication improved quality of life.
López-Pinar’s 2026 review of CBT points the other way, with moderate and sustained gains in everyday functioning, and quality-of-life gains that are smaller and not maintained at follow-up.
So symptoms, function and quality of life are three separate readings, taken on three different clocks. The Solanto review of CBT for adults says plainly what the work aims at: the executive functions of time management, organization and planning. A symptom count will not notice that the invoices went out.
What should a coach track instead?
Pick the function, then re-rate the same one. The outcomes that held across Sibley’s 63 trials were impairment and executive function skills. So the measure has to be something the client did this week: the invoice sent, the handover done.
Goal attainment scaling does that, and the honest zero is how to write down a week with no movement without turning it into a judgement of the client. When a rating does move, the missing denominator is the check on what else changed that month.
One r/ADHD poster, in a thread about symptoms and age, splits the two readings without the vocabulary for it: the symptoms are not worse, but adulthood hands you more that can go wrong.
Borrow the measurement instruction from that review, and leave its three-year number where it was measured, with the 10-to-19-year-olds.
Questions people ask next?
Is CBT effective for treating ADHD in adults?
It depends which outcome you name. Yang’s 2025 review in Frontiers in Psychiatry calls CBT highly effective on the core symptoms of adults. López-Pinar’s 2026 review finds everyday functioning improves moderately and lasts, with quality-of-life gains smaller and fading by follow-up.
Does the three-year maintenance finding apply to adults?
Not as evidence. Every one of the 63 trials enrolled 10-to-19-year-olds, so using the three-year number in adult coaching is an argument you are making, not a result you are citing.
What is the strongest long-term evidence in adults?
Nakashima’s 2026 eight-year follow-up in Scientific Reports: 48 adults randomised, 24 assessed at eight years, 6 of 14 CBT completers showing reliable improvement in inattention against 1 of 10 controls. The authors call it hypothesis-generating, with power at 0.45.
The goal ladder is free to fill and free to send; your client reads their own result and keeps the PDF.
Ember