Measurement as the interventiondraft

Draft, not signed off

Draft last pushed: 31 August 2026, 07:12 AEST

Every figure on this page was read from the document named beside it on 30 August 2026. Five of the cited papers' publishers refuse an automated request; each citation says so and each page opens for a person in a browser. This page reports findings from psychotherapy research; no equivalent study exists in financial advice, and the page says where that absence comes from. Christoph signs off as licensee.

In psychotherapy, measuring each client's progress with a short standard questionnaire and showing the practitioner the result has been tested as an intervention in its own right, in randomised trials. Financial advice has no equivalent study. In psychotherapy it was found that pharmaceutical interventions often did not have better outcomes than placebo, chemically inert, medicines, except in the most severe cases: the pooled analysis of the antidepressant trials submitted to the United States drug regulator, published and unpublished alike, found the drug's advantage over placebo reaching the conventional criterion for clinical significance only for patients at the upper end of the very severely depressed category (Kirsch and colleagues, PLoS Medicine, 2008). Hence there was a need to establish whether a practitioners' effect on clients' mental health existed. This research resulted in showing that the practitioners' effect exists and varies widely between individual therapists: Wampold and Imel's chapter on therapist effects, cited below, assembles that research, and the author's own review of doctors states it in one sentence, that psychotherapists can have an effect on their patients' mental health "that equals the strength of pharmaceutical interventions". This page sets out what the therapy literature reports, where its reviews disagree, and what has and has not been tested in advice.

What the therapy literature reports

The information is as of 30 August 2026.

What routine outcome measurement is

In this literature, a client answers a short standard questionnaire about symptoms and functioning at each session, most often the Outcome Questionnaire-45, a 45-question measure that takes about five minutes. Software compares the client's scores with the recorded course of many earlier clients who started at the same level, and flags the client whose course has fallen off that track. Feeding that flag back to the practitioner, sometimes with problem-solving tools attached, is the intervention the trials test; the treatment is otherwise unchanged. The research programme is assembled in two chapters of the field's standard handbook and in Wampold and Imel's book; every figure below is cited to the journal study or meta-analysis it comes from.

Sources Castonguay, Barkham, Lutz and McAleavey, "Practice-Oriented Research: Approaches and Applications", and Lambert, "The Efficacy and Effectiveness of Psychotherapy", chapters 4 and 6 of Bergin and Garfield's Handbook of Psychotherapy and Behavior Change, 6th edition, edited by Michael J. Lambert, John Wiley and Sons, 2013. The publisher's page for that book does not answer a plain request, so the citation stands without a link. Wampold and Imel, The Great Psychotherapy Debate: The Evidence for What Makes Psychotherapy Work, 2nd edition, Routledge, 2015, the quality-improvement discussion in its closing chapter.

The two books assemble the same body of trials; they are the map to the literature rather than additional evidence. The chapters' editor and one chapter's author, Michael Lambert, also leads the research group behind the Outcome Questionnaire system, which matters below.

Detection of deterioration, by judgement and by algorithm

Some clients end therapy worse than they began it, and the question of who notices has been measured. In a study at a large United States university counselling centre, therapists were asked to predict which of their current clients would leave treatment worse than they arrived. The published summary states that clinicians rarely accurately predict who will not benefit from psychotherapy, and that the formal monitoring methods, the questionnaire and its decision rules, identified 100% of the clients whose condition had deteriorated by the end of treatment, and 85% of them by their third session. A separate pair of studies, one reviewing therapy progress notes against recorded score changes and one surveying therapists, reports that therapists had considerable difficulty recognising client deterioration. And a survey of mental health professionals found 25% rating their own clinical skill at the 90th percentile of their profession and none rating themselves below average, while overestimating their clients' improvement and underestimating deterioration against the published rates.

Sources Hannan, Lambert, Harmon, Nielsen, Smart, Shimokawa and Sutton, "A lab test and algorithms for identifying clients at risk for treatment failure", Journal of Clinical Psychology 61(2), 2005; Hatfield, McCullough, Frantz and Krieger, "Do we know when our clients get worse?", Clinical Psychology & Psychotherapy 17(1), 2010; Walfish, McAlister, O'Donnell and Lambert, "An Investigation of Self-Assessment Bias in Mental Health Providers", Psychological Reports 110(2), 2012. All three publishers refuse an automated request and each page opens for a person in a browser; the figures were read from the published abstracts as indexed by Europe PMC on 30 August 2026.

The first study is one counselling centre, and the comparison is between practitioners' unaided predictions and an algorithm built from thousands of earlier score courses: a comparison of detection, not of treatment. The survey figures are practitioners' self-reports. Two of the three papers are from the research group that built the questionnaire being compared.

Fed back to the practitioner, the measurement has a small measured effect, and the reviews disagree on how firm it is

The feedback trials have been pooled four times, on four different sets of rules, and the four results belong together. The largest pooling took 58 randomised and non-randomised studies with 21,699 patients across routine care and found a small effect of feedback on symptom reduction: d =  0.15, and "small" is the authors' own word for it. The d is the difference between the two groups' average outcomes, measured in units of the spread of outcomes between patients (the standard deviation), so d = 0.15 means the average outcome of the patients whose practitioners received feedback ended 0.15 of one such unit ahead of the average of the patients whose practitioners did not. The 95% confidence interval runs from 0.10 to 0.20 – the exact definition of a confidence interval is disputed among statisticians. One explanation may be that there is a 95% probability or likelihood that the true value is between 0.10 and 0.20 and a 5% chance that the true value is lower than 0.10 or higher than 0.20, with the largest probability that the true value is fairly close to the measured value of 0.15. A 95% confidence interval that is wholly above or below zero typically denotes a statistically significant effect; an interval that straddles zero denotes that the measured effect may have been pure chance. For the patients flagged as not on track the effect was d = 0.17 (0.11 to 0.22). An earlier pooling of twelve studies found a short-term effect of d = 0.10 (0.01 to 0.19) that did not persist at longer follow-up. The research group behind the Outcome Questionnaire reanalysed the six trials of its own system, 6,151 patients in two settings, and reports the feedback interventions effective in enhancing outcomes, especially for the patients flagged as at risk of treatment failure, with two of the three feedback forms also reducing treatment failure. Against those three, the Cochrane review of routine outcome measurement for common mental health disorders admitted only randomised trials and pooled twelve of them, 3,696 participants: a standardised mean difference, the same kind of unit as d above, of -0.07 (95% confidence interval -0.16 to 0.01), which could not be distinguished from no difference; it graded the evidence as low quality, because it judged every included study at high risk of bias, and concluded there was insufficient evidence to support the routine use of these measures in those settings. In the earlier years of Cochrane reviews, they constituted the gold standard for establishing the validity of scientific findings. In recent years that standing has been contested: in 2018 the organisation's own groups described a governance crisis, after the governing board expelled a founding member by a six-to-five vote and four further board members resigned (the Cochrane Work group's own statement), amid published accusations that the organisation had grown too close to the pharmaceutical industry.

What can and cannot be concluded: wherever the effect is measured it is small; it is most consistent for the clients the algorithm flags as off track; and the strictest of the four reviews could not tell the effect apart from zero and judged the underlying trials weak.

Sources de Jong, Conijn, Gallagher, Reshetnikova, Heij and Lutz, "Using progress feedback to improve outcomes and reduce drop-out, treatment duration, and deterioration: A multilevel meta-analysis", Clinical Psychology Review 85, 2021 (the publisher refuses an automated request; the abstract was read through Europe PMC); Knaup, Koesters, Schoefer, Becker and Puschner, "Effect of feedback of treatment outcome in specialist mental healthcare: meta-analysis", British Journal of Psychiatry 195(1), 2009; Shimokawa, Lambert and Smart, "Enhancing treatment outcome of patients at risk of treatment failure", Journal of Consulting and Clinical Psychology 78(3), 2010 (the publisher refuses an automated request and the page opens for a person in a browser); Kendrick, El-Gohary, Stuart, Gilbody, Churchill and others, "Routine use of patient reported outcome measures (PROMs) for improving treatment of common mental health disorders in adults", Cochrane Database of Systematic Reviews, 2016, CD011119, linked at its freely readable PubMed Central copy because the Cochrane Library's own address performs a bot check that did not clear in a browser here on 30 August 2026.

The four poolings differ in what they admit, which is most of why they differ in what they find: the largest counts randomised and non-randomised studies across curative care, the Cochrane review only randomised trials in common mental health disorders, graded for quality. The outcomes are symptom and functioning scores over weeks to months. Part of the trial evidence and one of the four poolings come from the research group that developed the system being tested: Michael Lambert is an author of the Shimokawa review and of the Hannan and Walfish studies above. The Knaup, de Jong and Cochrane reviews are from other research groups.

No equivalent study exists in financial advice

No study was found that tests routine outcome measurement in financial advice: a standard measure of the client's position, collected on a schedule and fed back to the adviser or to the client, compared against advice without it, on any client outcome, in any jurisdiction. What psychotherapy had before its trials could run is also what advice does not have: an agreed outcome questionnaire, and a recorded expected course for clients who start in the same position. What the advice record itself contains is treated on the State of Research page.

Source that no such study was found is this practice's own search, of 30 August 2026.

About

Built by In Your Interest Financial Planning. In Your Interest Financial Planning Pty Ltd, ABN 28 094 300 464 is Authorised Rep. No 308161 of Fiduciary Duty Advisers Pty Ltd AFSL No 527434. Nothing on this page is advice to you or to your clients, and nothing on it is a recommendation of any product, any strategy or any adviser. Corrections are welcome and wanted – contact us.