Addresses a documented definitional problem
The paper targets the one-versus-two antidepressant-trial inconsistency that limits comparison across TRLLD studies and proposes a common minimum threshold.
↳ Introduction; Table 8
Crunching the numbers. Responsibly.
BACKGROUND: Treatment-Resistant Late-Life Depression (TRLLD) represents a significant clinical and research challenge. Studies on TRLLD use heterogeneous diagnostic criteria limiting the interpretation of results and the development of clinical guidelines.This study aimed to establish expert consensus on the core components and definition of TRLLD through a Delphi process conducted by a European Task Force of clinicians and researchers with experience in late-life depression.
METHODS: We conducted an electronic Delphi study with 30 European experts to identify key definitional elements of TRLLD. In the 1st survey round (SR), participants responded to open-ended questions on definitions of TRLLD. Responses informed the development of 70 structured items for 2nd SR, involving 58 closed (yes/no) items and 12 multiple-choice questions across six domains: symptom presentation, cognitive impairment, comorbidities, pharmacotherapy, treatment adherence, and psychosocial factors. Consensus was defined as ≥70% agreement. In 3rd SR, participants reviewed comments and validated the final categorical definition, the operational staging model, and the derived decision algorithm for TRLLD.
RESULTS: Consensus was reached for 72.4% of the 58 closed items. Experts agreed that the categorical definition of TRLLD should correspond to major depressive disorder in individuals aged ≥65 years who show insufficient response to two adequate antidepressant treatments, in the absence of dementia and medical conditions that could account for depressive symptoms. An operational staging model and decision-making algorithm were developed from consensus items.
CONCLUSION: This study provides the first European consensus definition of TRLLD and highlights the need for age-adapted diagnostic criteria that reflect the clinical complexity of depression in older adults.
Experts agreed that the categorical definition of TRLLD should correspond to major depressive disorder in individuals aged ≥65 years who show insufficient response to two adequate antidepressant treatments, in the absence of dementia and medical conditions that could account for depressive symptoms.
core items reached consensus, but final synthesis and third-round endorsement are not fully quantifiable
Consensus was reached for 72.4% of the 58 closed items.
the reported proportion is consistent with the item counts in the consensus and non-consensus tables
Most panellists agreed that one should wait approximately 4 weeks to observe the effects of antidepressants.
only 46% selected four weeks, narrowly exceeding the 42% selecting six weeks
Derived from the full evaluation — not a separate score.
Strengths
The paper targets the one-versus-two antidepressant-trial inconsistency that limits comparison across TRLLD studies and proposes a common minimum threshold.
↳ Introduction; Table 8
The study uses open-ended item generation, a prespecified 70% consensus threshold, aggregate feedback between rounds, and a sensitivity analysis for second-round non-completion.
↳ Methods, Delphi Survey Rounds and Participant Feedback; Results, Sensitivity Analysis
The categorical definition is linked to an adapted Thase and Rush staging model and a decision algorithm, creating a concrete framework for research classification.
↳ Tables 8–9; Figure 2
Limitations
Table 8 includes criteria such as specific scale cutoffs, no upper age limit, inclusion of MCI, and exclusion of acute physical comorbidities without clear matching second-round consensus items. Third-round endorsement is not quantified.
↳ Tables 4–6 and 8; Results, Third Survey Round
The Discussion calls the 46% four-week preference the view of “most panellists” and describes a lack of psychotherapy consensus despite two threshold-reaching psychotherapy items.
↳ Discussion, Global Definition and Clinical Presentation; Tables 4 and 6
The framework comes from a small convenience panel with uneven national and disciplinary representation and has not been tested for feasibility, reliability, or prognostic validity.
↳ Table 2; Strengths and Limitations; Conclusion
The paper earns credit for a structured Delphi process with open item generation, a prespecified threshold, feedback between rounds, and explicit sensitivity analysis. Its contribution is useful but incremental, extending Patrick et al. and Thase and Rush with a European definition and staging framework. Scores were constrained by incomplete mapping from second-round votes to several Table 8 criteria, unquantified third-round endorsement, and imprecise interpretation of the four-week and psychotherapy results. The framework is therefore most defensible as a research harmonisation proposal requiring independent validation, not as a validated clinical standard.
Nabu’s assessment, alongside the field’s view.
Are you an author of this paper?
Sound3.7
Confidence mediumThe paper makes an incremental but meaningful contribution by proposing a European categorical definition, adapted staging model, and decision algorithm for a field with heterogeneous resistance thresholds. Its advance remains consensus refinement rather than empirical resolution of diagnostic validity.
Consensus was reached for 72.4% of the 58 closed items.
The Delphi uses open item generation, a prespecified threshold, structured feedback, and attrition sensitivity analysis. Execution is limited by the small convenience panel, unquantified third-round endorsement, experience-skewed attrition, and incomplete traceability between item-level votes and Table 8.
all expert panel members (who are co-authors of this consensus paper) reviewed, refined, and endorsed the final consensus
The progression from survey rounds to definition, staging model, and algorithm is generally followable and the outputs are explicitly framed as expert consensus. Precision is reduced by calling a 46% plurality “most panellists” and by describing psychotherapy findings inconsistently.
Most panellists agreed that one should wait approximately 4 weeks
The paper positions the framework against Patrick et al., Thase and Rush, heterogeneous trial definitions, and relevant clinical evidence, while acknowledging geographic, disciplinary, anchoring, and validation limitations. The application language remains somewhat stronger than warranted by fragile items and absent empirical testing.
consensus represents agreement among panellists and does not necessarily indicate an objectively “correct” position
Caveats4 of 4 checks
The overall participation counts and 72.4% consensus calculation are coherent, but some narrative statements and final criteria cannot be fully reconciled with the reported item-level results. The third-round endorsement also lacks response counts or agreement percentages.
Participant selection and aggregate data availability are described, but the supplied article does not state an ethics-review or consent status for the expert survey. Conflict disclosure is incomplete because it states only that most panellists reported no relevant conflicts.
Flags: 2 declared / 5 total
79 references in manuscript 73 of 73 checkable references found in an index 6 have no canonical index record — counted, but not index-checkable 2 references confirmed by manual review
No retraction notice found in Retraction Watch.
Sources: Retraction Watch ✓
Where this paper’s evidence sits on the path from initial observation to real-world use.
The definition, staging model, and algorithm are operational outputs, but they remain expert-derived and untested for feasibility, inter-rater reliability, prognostic validity, or clinical benefit. Readiness is therefore preliminary rather than deployment-ready.
future studies should evaluate the proposed TRLLD criteria in independent samples
AI-generated, human-governed. Something look off? Contact us to request a review.