Abstract
Individual differences are a foundational construct in psychological science, yet clinical psychology has applied this principle inconsistently to the selection and evaluation of its own practitioners. This perspective article selectively synthesizes evidence for reliable variability in client outcomes across individual clinicians (the “therapist effect”) and reviews candidate clinician-level correlates of that variability, including facilitative interpersonal skills, therapist empathy, alliance rupture-repair capacity, and engagement in deliberate practice. It then explicitly distinguishes statistical therapist-effect estimates from stable traits, trainable competencies, therapist–client matching effects, and contextual influences on performance — a distinction the broader literature does not always maintain. Current admissions and licensure practices in United States clinical psychology doctoral training are described within an explicit jurisdictional scope, contrasted with the constructs reviewed, and evaluated against a set of conditions that any candidate selection variable would need to satisfy before informing admissions or licensure decisions. A staged, fairness-constrained research agenda for prospective validation of performance-based interpersonal assessment is proposed. The article concludes that current evidence justifies rigorous, multisite validation research, not the implementation of new selection criteria.
1. Introduction
Psychology's disciplinary identity rests on the premise that people differ in stable, measurable, and consequential ways. Selection systems built on this premise are common across the field's applied branches, from personnel psychology to educational placement. It is therefore worth asking how consistently the discipline applies this same premise to the pipeline that produces its own practitioners. In United States clinical psychology, doctoral admission relies chiefly on academic indicators and interviews of varying structure, and licensure relies on supervised-hour requirements and a written knowledge examination (Nye & Ryan, 2023; Wood, 2025). Neither stage incorporates a standardized, validated assessment of the interpersonal capacities that a growing body of psychotherapy-outcome research associates with clinician effectiveness.
This article is offered as a perspective and selective synthesis, not a systematic review: it does not follow a registered search protocol or formal eligibility criteria, and its purpose is to develop an argument and a research agenda rather than to exhaustively catalogue the relevant literature. Its central claim is deliberately modest. The evidence reviewed below does not yet justify implementing new admissions or licensure criteria. It does, in this author's view, justify a structured program of prospective, multisite validation research into whether performance-based interpersonal assessment can improve the selection and development of clinical psychologists, conducted with explicit attention to incremental validity, fairness, and resistance to gaming. The remainder of the article proceeds in six steps: (a) distinguishing therapist effects from therapist characteristics, (b) reviewing the empirical case for clinician-level outcome variability and its limits, (c) synthesizing candidate individual-difference correlates of that variability, (d) specifying the additional evidence needed before any correlate could justify a selection decision, (e) describing what current gatekeeping systems assess within an explicit jurisdictional scope, and (f) addressing the fairness and ethical constraints that any such proposal must satisfy.
Scope and jurisdiction. This article focuses on doctoral-level clinical psychology training and licensure in the United States. Licensure is administered independently by state, provincial, and territorial psychology boards rather than as a single national process; the Association of State and Provincial Psychology Boards (ASPPB) coordinates a common written examination and mobility programs across more than sixty member jurisdictions, but specific requirements vary by jurisdiction (ASPPB, n.d.). Admissions practices likewise vary across the roughly two hundred accredited doctoral programs in the United States. Statements below about “current practice” describe general patterns documented in the cited sources, not a uniform national standard, and should not be assumed to generalize to other countries' training and licensure systems or to other mental health professions (e.g., counseling, social work, psychiatry, marriage and family therapy).
2. Distinguishing Therapist Effects From Therapist Characteristics
A “therapist effect” as typically estimated in this literature is a variance component derived from a multilevel model: the proportion of variability in client outcomes attributable to therapist identity after accounting for client-level factors (Baldwin & Imel, 2013). This is a statistical quantity describing a dataset, not a direct measurement of any single stable psychological trait, and the distinction matters for the argument that follows. Several documented phenomena can inflate or deflate an estimated therapist effect without necessarily reflecting an enduring individual difference in skill. Non-random allocation of clients to therapists can matter: Saxon and Barkham (2012) found that estimated therapist effects were larger among more severe patients, suggesting that case mix, not only clinician skill, shapes the size of the effect. Clinical setting and treatment format also matter: Johns et al. (2019) found smaller weighted therapist effects in university counseling-center samples than in primary-care or specialist settings. Outcome measurement choices matter as well: Kraus et al. (2011) found that individual therapists were differentially effective across symptom and functioning domains within the same caseload, such that a clinician's apparent effectiveness partly depended on which domain was measured. Caseload size and statistical estimation error matter too, since Johns et al. (2019) noted that therapist sample sizes in much of the literature remain below recommended thresholds for stable multilevel estimation. Therapist-by-client interaction effects are a further complication, since the Kraus et al. (2011) domain-specificity finding is also consistent with certain clinicians being more effective with certain client presentations rather than uniformly more skilled. Finally, performance is not necessarily fixed over time: deliberate-practice research (Chow et al., 2015) implies that a clinician's effectiveness can change with training activity rather than remaining a stable trait-level quantity.
Because of these sources of variability, this article distinguishes five conceptually separate constructs that public and academic discussion of “therapist effects” sometimes elides: (a) observed outcome variability across therapists in a given dataset, a descriptive and sample-specific finding; (b) stable individual-difference traits, hypothesized enduring characteristics of the person; (c) trainable competencies, skills shown to change with instruction, feedback, or practice; (d) therapist–client matching effects, in which effectiveness depends on the pairing rather than either party alone; and (e) contextual and structural influences on performance, such as caseload, supervision quality, and organizational climate. Only categories (b) and (c) bear directly on a selection-and-training argument; category (a) alone — evidence that therapists differ — cannot establish either. Section 4 examines, for each proposed predictor, which of these categories the evidence most plausibly supports, and Section 5 specifies what further evidence would be needed to justify treating any of them as a selection criterion.
3. The Therapist Effect: Evidence and Its Limits
With that distinction in place, the empirical case that individual clinicians differ in the outcomes their clients achieve is well replicated, if modest in average magnitude. Baldwin and Imel's (2013) review and Johns et al.'s (2019) systematic update converge on a weighted average therapist effect of approximately 5% of outcome variance in practice-based studies, rising to a weighted average of 8.2% within randomized controlled trials — settings explicitly designed to standardize treatment delivery. Individual studies report considerably more variability: Johns et al. (2019) found therapist-effect estimates ranging from 0.2% to 29% across the twenty studies they reviewed, and cautioned that a single overall statistic may lack precision given this heterogeneity.
Naturalistic outcome data make the practical stakes concrete, with the caveats noted above about case mix and measurement. Saxon and Barkham (2012), studying 119 therapists and 10,786 clients in United Kingdom primary-care services, found that individual recovery rates ranged from 23.5% to 95.6%. Kraus et al. (2011), analyzing 6,960 clients across 696 therapists, classified a substantial minority of clinicians as reliably associated with client deterioration, with effect sizes for these “harmful” therapists (d = −0.91 to −1.49) comparable in magnitude, though opposite in direction, to those for the most effective therapists (d = 1.00 to 1.52). Wampold and Brown (2005) reported similar naturalistic variability in a large managed-care sample, and Feliciano et al. (2026), examining 8,145 clients treated by 44 therapists delivering internet-supported cognitive behavioral therapy, again found reliable, though modest, therapist-attributable variance — indicating the effect is not confined to face-to-face delivery. Okiishi et al. (2003) additionally found that supervisor ratings of competence and therapist theoretical orientation did not reliably distinguish more from less effective clinicians in their sample, a result they described, somewhat wryly, as “waiting for supershrink”: a search for markers of superior performance that conventional training and credentialing indicators did not readily supply.
This evidence base has an important boundary condition that qualifies any general claim about the size or inevitability of therapist effects. King et al. (2017), meta-analyzing 15 studies in which participants were randomized to receive the same cognitive behavioral protocol through either self-help materials or a therapist, found broadly equivalent treatment completion and outcomes across the two formats and, contrary to their own hypothesis, broadly equivalent outcome variability as well — suggesting that in this context, differences among individual therapists were not sufficient to make therapist-delivered outcomes more variable than minimally supported self-help. This finding does not overturn the broader therapist-effects literature, which spans many more studies and settings, but it indicates that the magnitude and even the presence of clinician-attributable variability is context-dependent rather than a fixed property of psychotherapy as such.
A further caution concerns how this literature is compared to research on treatment modality. Work motivating the “contextual model” of psychotherapy (Wampold & Imel, 2015; Wampold, 2015) has found that differences between bona fide treatment approaches often account for less outcome variance than relational and therapist-attributable factors considered in aggregate. It is tempting to compress this into a claim that “whom a client sees matters as much as, or more than, which treatment they receive,” but this comparison should be made cautiously: therapist-effect estimates, treatment-difference estimates, alliance–outcome associations, and aggregated “common factor” estimates typically come from different research designs, different outcome measures, and different levels of analysis, and are not directly interchangeable in magnitude. Alliance and empathy, moreover, can vary across clients and across sessions for the same therapist, so they are not simply restatements of a fixed, person-level therapist effect. The defensible claim from this literature is narrower: clinician-associated variability in outcomes persists and remains consequential even when treatment protocols are standardized, as in randomized controlled trials (Johns et al., 2019). This is sufficient to motivate the inquiry that follows without requiring the stronger, less defensible claim that existing evidence has settled the relative importance of clinician identity versus treatment selection in general.
4. Candidate Individual-Difference Correlates of Clinical Effectiveness
If clinicians differ reliably in their outcomes in at least some settings, the natural next question is what accounts for the difference, and which of the five categories introduced in Section 2 — stable trait, trainable competency, matching effect, or context — each candidate correlate most plausibly represents. Four literatures offer candidate answers with meaningfully different evidentiary weight; a fifth, more speculative literature is treated separately and explicitly flagged as such.
4.1 Facilitative Interpersonal Skills
Anderson and colleagues developed the Facilitative Interpersonal Skills (FIS) task, a performance-based measure in which clinicians respond to videotaped simulations of difficult client moments (e.g., hostility, disengagement) and are rated by independent observers on relational responsiveness. In an initial study of 25 therapists and 1,141 clients, Anderson et al. (2009) found that FIS, but not therapist age, theoretical orientation, or a self-report social-skills measure, accounted for significant variance in the therapist effect. A subsequent randomized clinical trial found that FIS predicted alliance and outcome, whereas formal clinical training status did not (Anderson et al., 2016b). Most consequential for the argument developed here, Anderson et al. (2016a) administered the FIS task to clinical graduate students before the start of their clinical training and found that pre-training scores prospectively predicted their treatment effectiveness with real clients years later. This is the strongest existing evidence bearing directly on the selection question: a relational-skills assessment administered before training predicted clinical effectiveness after training, which places FIS closer to category (b) or (c) in the Section 2 framework than most other candidates reviewed here. Two caveats are nonetheless warranted. First, this prospective design has, to this author's knowledge, been reported primarily by one research group using overlapping training samples; independent replication across other training sites has not been established in the sources reviewed here. Second, prospective prediction from a single cohort does not by itself establish the further properties — incremental validity, cross-population fairness, and resistance to coaching — addressed in Section 5.
4.2 Therapist Empathy
Elliott et al.'s (2018) meta-analysis, conducted for the American Psychological Association's Interdivisional Task Force on Evidence-Based Therapy Relationships, synthesized 82 independent samples (N = 6,138) and found a moderately strong association between therapist empathy and client outcome (r = .28, equivalent to d = .58). The construct under study is intended to capture accurate perception and communication of a client's internal experience rather than generalized warmth, but two measurement issues deserve emphasis rather than the more specific label “empathic accuracy,” which refers to a narrower construct with its own separate measurement tradition not directly tested in the studies aggregated here. First, empathy in this literature has been rated from multiple sources — client, therapist, and independent observer — and associations with outcome are not necessarily uniform across rating source; client-rated empathy in particular is vulnerable to the concern that a client who is already improving, or who is more broadly satisfied with treatment, may also rate the therapist as more empathic, creating a common-method and temporal-ordering ambiguity that observer-rated measures are somewhat better positioned to avoid. Second, because empathy is assessed session-by-session or case-by-case rather than as a fixed pre-training characteristic, existing evidence speaks most directly to a trainable, in-session competency (category c) rather than to a stable trait present in an applicant before training (category b), which limits its direct relevance to an admissions-stage selection instrument as distinct from an in-training assessment and feedback target.
4.3 Alliance Rupture-Repair Capacity
Because the therapeutic relationship fluctuates rather than remaining stable, the capacity to recognize and repair strains in the alliance — a “rupture,” manifested as disagreement, disengagement, or a breakdown in collaboration — has emerged as a distinguishable clinical skill. Eubanks et al.'s (2018) meta-analysis of 11 studies (N = 1,314) found a moderate association between rupture-repair episodes and positive outcome (r = .29, d = .62). Like empathy, this construct has been studied primarily as an in-session process variable in already-practicing or already-selected clinicians rather than as a pre-training individual difference; the evidence therefore speaks most directly to what training and supervision should target (category c) rather than to what an admissions committee could screen for at intake.
4.4 Deliberate Practice Versus Accumulated Experience
A persistent and uncomfortable finding in this literature is that clinical experience, defined as years in practice or caseload volume, is at best weakly related to client outcomes (Goldberg et al., 2016; Tracey et al., 2014). Tracey et al. (2014) argue that psychotherapy, unlike domains with well-established expertise effects, lacks the rapid, unambiguous outcome feedback that experience needs in order to translate into improved performance. Chow et al. (2015), studying 69 therapists and 4,580 clients, found that the amount of time therapists spent in deliberate practice — structured, effortful, feedback-guided rehearsal of specific clinical skills, distinct from ordinary caseload accumulation — significantly predicted client outcomes, with the most effective therapists reporting disproportionately more time spent reviewing session recordings. This body of work (see also Miller et al., 2020; Rousmaniere, 2016) reframes clinical excellence as substantially a function of a cultivable orientation toward self-scrutiny and structured improvement (squarely category c, a trainable competency) rather than of tenure alone, with direct implications for how graduate training and continuing-education requirements are structured, independent of any admissions-stage selection question.
4.5 Personal Psychological Integration: A Plausible but Underdetermined Factor
A separate and more speculative literature examines how clinicians' own histories of psychological difficulty shape their clinical work. Zerubavel and O'Dougherty Wright (2012) describe “the dilemma of the wounded healer”: personal experience of distress can deepen empathic attunement and tolerance for a client's pain when sufficiently processed, but can also generate blind spots and countertransference reactivity when it has not been. This literature does not yet identify a measurable predictor of client outcomes; it is composed largely of qualitative and theoretical work rather than the outcome-linked quantitative designs reviewed in Sections 4.1–4.4, and it should not be read alongside them as evidentiarily comparable. It is better understood as a theoretically plausible but empirically underdetermined research question than as an established predictor. Separately, and more modestly, Norcross et al. (2008) found that a substantial proportion of practicing therapists in their survey had never undergone personal therapy themselves, despite widespread professional endorsement of its value, and personal therapy has not been a uniform requirement of accredited doctoral training in the United States. This establishes that practice is inconsistent, not that requiring personal therapy would improve client outcomes; that causal question remains open and is not treated here as resolved.
5. From Correlates to Selection Criteria: What Additional Evidence Is Needed
Evidence that a variable correlates with outcomes in already-practicing clinicians does not by itself establish that it can be used to screen applicants before training. The inferential distance between these two claims is the central methodological issue this article addresses, and it is treated here as a conceptual model rather than a closing caveat. Before any candidate variable from Section 4 could reasonably inform an admissions or licensure decision, evidence would be needed on each of the following:
● Stability before training — that the characteristic is measurable in applicants prior to clinical training, not only in clinicians already selected into and shaped by a training program.
● Reliable measurement — that scores show adequate inter-rater and test–retest reliability when administered under selection conditions rather than only in research samples.
● Prospective predictive validity — that scores obtained before or early in training forecast later, independently assessed client outcomes, as opposed to concurrent correlations among already-practicing clinicians.
● Incremental validity — that the measure predicts outcomes beyond what existing admissions information (grades, letters, research record, standard interviews) already predicts.
● Cross-population and cross-setting validity — that predictive relationships replicate across client populations, presenting problems, clinical settings, and, ideally, multiple independent training programs rather than a single research group's samples.
● Absence of unacceptable demographic or cultural bias — that the measure does not produce adverse impact or systematically disadvantage applicants from particular racial, cultural, linguistic, disability, or socioeconomic backgrounds independent of later clinical competence.
● Resistance to coaching, impression management, and repeated exposure — that scores cannot be substantially inflated by applicants who have been coached on, or repeatedly exposed to, the specific task materials, a live concern for any standardized performance task used at scale in high-stakes admissions.
Measured against these conditions, FIS is the only candidate reviewed here with direct evidence bearing on stability-before-training and prospective predictive validity (Anderson et al., 2016a), and even that evidence is limited to a single research group's samples, leaving incremental validity, cross-population fairness, and coaching-resistance essentially untested. Therapist empathy (Elliott et al., 2018) and rupture-repair capacity (Eubanks et al., 2018) have substantial evidence as concurrent correlates of outcome among practicing clinicians, but have not, in the sources reviewed here, been evaluated as pre-training selection instruments at all; they currently inform what training and supervision should target far more than what an admissions committee could screen for. Deliberate practice (Chow et al., 2015) is explicitly a trainable orientation rather than a pre-existing trait, and is therefore a training-design variable rather than a selection variable in the framework used here. None of the candidates reviewed has published evidence on subgroup fairness or resistance to coaching in a clinical-psychology admissions context specifically. This is the evidentiary gap that Sections 7 and 8 address.
6. Current Selection and Evaluation Practices: A Scope-Limited Description
Within the United States doctoral-admissions scope defined in Section 1, available evidence indicates that clinical psychology programs weigh undergraduate and graduate grade point average, standardized test scores where still required, research productivity, letters of recommendation, and interviews of varying structure (Nye & Ryan, 2023; Woo et al., 2023). A first-person account from a current doctoral student (Wood, 2025), published as commentary rather than as peer-reviewed research, describes considerable variability in how individual faculty and programs weigh these criteria and argues that admissions decisions are not always made in a manner consistent with the field's own standards of transparent, data-driven decision-making — a useful illustration of practice, though not a substitute for systematic evidence. Nye and Ryan (2023) and Woo et al. (2023) provide evidence on graduate admissions predictors more broadly across psychology and other disciplines rather than on clinical psychology doctoral admissions specifically; the more granular, peer-reviewed evidence this article would ideally draw on — for example, on interview structure, practicum evaluation criteria, or internship-match criteria specific to clinical psychology — is comparatively sparse in the sources located for this synthesis, and this sparsity is itself a notable gap rather than a settled absence. Accreditation standards for doctoral training specify broad competency domains to be developed and evaluated over the course of training, but the granularity and standardization of any admissions-stage assessment of relational or interpersonal competencies specifically is not well documented in the peer-reviewed literature reviewed here.
Licensure, in the jurisdictions coordinated through ASPPB, is governed by supervised-hour requirements and the Examination for Professional Practice in Psychology, an examination of declarative, ethical, and applied knowledge (ASPPB, n.d.); this article did not locate peer-reviewed evidence that this examination, or supervised-hour requirements as typically implemented, have been validated against the client-outcome constructs reviewed in Section 4. Post-licensure evaluation is discussed in the professional-issues literature as comparatively limited, often relying on utilization metrics, the absence of complaints or disciplinary action, and informal professional reputation (Tracey et al., 2014); this characterization should be read as reflecting that literature's discussion rather than as a fully quantified empirical claim, since this synthesis did not conduct an independent audit of licensing-board evaluation practices across jurisdictions.
7. Fairness, Equity, and Ethical Constraints on Interpersonal Screening
Any proposal to assess applicants on interpersonal or performance-based characteristics carries fairness risks that must be addressed directly rather than left implicit, particularly given that current guidance in graduate admissions more broadly favors holistic review specifically because it can reduce, but does not automatically eliminate, disparities associated with reliance on any single quantitative metric (Kent & McCarthy, 2016).
Cultural and linguistic variation in communication style is a first concern. Raters evaluating a videotaped interpersonal response may systematically favor applicants whose relational style matches a dominant professional or cultural norm, penalizing stylistic difference that is not in fact clinically disqualifying. Any rubric used for a task such as FIS would need explicit validation for measurement invariance across cultural and linguistic groups before being used in a decision-relevant way.
Disability and neurodivergence raise a related concern. A task requiring rapid, in-the-moment verbal responsiveness to simulated distress may disadvantage applicants with certain disabilities or neurodivergent communication styles who could nonetheless become effective clinicians using different, but clinically adequate, relational strategies. Standard reasonable-accommodation obligations would apply to any such instrument, and its design should anticipate this rather than treat it as an implementation afterthought.
Adverse impact more generally is an empirical property that must be demonstrated, not assumed away. The personnel-selection literature offers a relevant, if imperfect, analogy: structured, behaviorally anchored interview formats have been found to produce smaller subgroup score differences and to be somewhat more resistant to certain rater biases than unstructured interview formats (Levashina et al., 2014). This argues, if anything, for behaviorally anchored, multiply rated performance tasks over unstructured admissions interviews as currently practiced — but it does not establish that FIS or any similar instrument would be free of adverse impact in a clinical-psychology applicant pool specifically; that would need to be tested directly, as specified in Section 5.
Privacy and mental health history raise a boundary that this article treats as non-negotiable rather than as a matter of degree: nothing proposed here should be read as license to screen applicants on the basis of disclosed mental health history, treatment history, or diagnosis. The constructs reviewed in Section 4 concern observable, professionally relevant interpersonal behavior in a structured task, not a person's private health status or personal history, and conflating the two would raise both fairness and legal concerns and would be inconsistent with the profession's own obligations regarding disability and mental health nondiscrimination.
Finally, due process and construct validity are linked concerns. Any pilot instrument should include a mechanism for score review and applicant appeal and should not be used in an exclusionary, pass/fail manner unless and until independent, multisite validity and fairness evidence justifies that use. Rating rubrics and rater training should themselves be validated across demographic subgroups to guard against the risk that a task nominally measuring clinical competence in fact measures conformity to a particular, culturally specific expressive style.
8. A Staged Research Agenda
Given the gaps identified in Sections 5 through 7, this article proposes a staged program of research rather than a ready-made selection protocol. Each stage builds on the last and is gated by evidence, not by convenience or by the interim results of an earlier stage alone.
Stage 1 (admission, research-only). A standardized, behaviorally anchored performance task modeled on FIS would be administered to incoming cohorts purely as a research measure, with informed consent, kept fully separate from actual admission decisions, so that its properties can be studied without yet affecting anyone's outcome.
Stage 2 (training). The task would be re-administered at defined training milestones, paired with structured, deliberate-practice-informed feedback (Chow et al., 2015; Miller et al., 2020), to examine whether and how scores change with targeted training rather than assuming they are fixed.
Stage 3 (practicum). Task scores would be linked to practicum-level process measures — alliance ratings, client retention and dropout, and session-level symptom trajectories — adjusting statistically for client severity and case mix, consistent with the case-mix concerns raised by Saxon and Barkham (2012).
Stage 4 (post-training outcomes). Prospective prediction of client outcomes during internship and early independent practice would be tested using risk-adjusted, multilevel outcome models consistent with therapist-effects methodology (Baldwin & Imel, 2013), rather than simple unadjusted comparisons.
Stage 5 (formal validation). This stage would examine, as a coordinated package rather than piecemeal: inter-rater and test–retest reliability under selection-like conditions; incremental validity over existing admissions criteria; measurement invariance and subgroup score differences across race, ethnicity, gender, disability status, and linguistic background; susceptibility to coaching or repeated-exposure effects; and replication across multiple training sites, extending beyond the single research group responsible for most current FIS evidence.
Stage 6 (decision, contingent). Only if Stages 1 through 5 produce evidence of adequate reliability, incremental validity, and subgroup fairness across independent, multisite samples should such an instrument be considered for any role in actual admissions or licensure decisions — and even then, this article recommends a compensatory, non-exclusionary combination with existing criteria rather than a stand-alone pass/fail gate, consistent with the due-process concerns raised in Section 7.
9. Limitations
This is a perspective and selective synthesis rather than a systematic review; it did not follow a registered search protocol, formal eligibility criteria, or dual independent screening, so relevant literatures may be underrepresented and the balance of evidence presented reflects the author's selection choices. Its scope is limited to United States doctoral-level clinical psychology, and its claims should not be assumed to generalize to other licensure jurisdictions, other countries, or other mental health professions. The strongest direct evidence for the central selection-relevant claim (Anderson et al., 2016a) derives from a single research group using overlapping training samples, and independent multisite replication has not been established in the literature reviewed here. The boundary-condition finding from King et al. (2017) indicates that the size, and possibly the generality, of therapist effects is context-dependent, which complicates confident population-level claims about the practical stakes of selection reform. Finally, several constructs reviewed in Section 4 (therapist empathy, rupture-repair capacity) are measured predominantly in already-practicing or already-selected clinicians, which limits what can currently be inferred about their value as pre-training selection criteria specifically, as distinct from their better-supported role as targets for in-training assessment and feedback.
10. Conclusion
Clinical psychology has produced a substantial and increasingly sophisticated body of evidence that individual clinicians differ, at least in many settings, in the outcomes their clients achieve, and that this variability is partly associated with measurable interpersonal characteristics rather than solely with credentials, theoretical orientation, or years of experience. At the same time, the inferential distance between this evidence and any concrete admissions or licensure reform is presently large: stability before training, incremental validity, cross-population fairness, and resistance to coaching remain largely untested for every candidate variable reviewed here. This article's conclusion is accordingly moderated relative to a stronger policy claim it could have made. The evidence reviewed is sufficient to justify a coordinated, multisite, fairness-audited program of prospective validation research on performance-based interpersonal assessment in clinical psychology training. It is not yet sufficient to justify treating any such assessment as ready for use in actual selection or licensure decisions, and this article should be read as an invitation to that research program rather than as a case for its conclusion being foregone.
Declarations
Conflict of Interest: [To be completed by author(s). No known conflicts of interest are declared at this time.]
Funding: [To be completed by author(s). This work received no dedicated funding at the time of drafting.]
Data Availability: No new empirical data were generated for this conceptual/perspective article.
Author Contributions: [To be completed by author(s) if submitted with co-authors, using a standard contributorship taxonomy such as CRediT.]
Acknowledgments: [Optional — to be completed by author(s).]
References
Anderson, T., McClintock, A. S., Himawan, L., Song, X., & Patterson, C. L. (2016a). A prospective study of therapist facilitative interpersonal skills as a predictor of treatment outcome. Journal of Consulting and Clinical Psychology, 84(1), 57–66. https://doi.org/10.1037/ccp0000060
Anderson, T., Ogles, B. M., Patterson, C. L., Lambert, M. J., & Vermeersch, D. A. (2009). Therapist effects: Facilitative interpersonal skills as a predictor of therapist success. Journal of Clinical Psychology, 65(7), 755–768. https://doi.org/10.1002/jclp.20583
Anderson, T., Crowley, M. E. J., Himawan, L., Holmberg, J. K., & Uhlin, B. D. (2016b). Therapist facilitative interpersonal skills and training status: A randomized clinical trial on alliance and outcome. Psychotherapy Research, 26(5), 511–529. https://doi.org/10.1080/10503307.2015.1049671
Association of State and Provincial Psychology Boards. (n.d.). About ASPPB. Retrieved September 17, 2026, from https://www.asppb.net
Baldwin, S. A., & Imel, Z. E. (2013). Therapist effects: Findings and methods. In M. J. Lambert (Ed.), Bergin and Garfield’s handbook of psychotherapy and behavior change (6th ed., pp. 258–297). Wiley.
Chow, D. L., Miller, S. D., Seidel, J. A., Kane, R. T., Thornton, J. A., & Andrews, W. P. (2015). The role of deliberate practice in the development of highly effective psychotherapists. Psychotherapy, 52(3), 337–345. https://doi.org/10.1037/pst0000015
Elliott, R., Bohart, A. C., Watson, J. C., & Murphy, D. (2018). Therapist empathy and client outcome: An updated meta-analysis. Psychotherapy, 55(4), 399–410. https://doi.org/10.1037/pst0000175
Eubanks, C. F., Muran, J. C., & Safran, J. D. (2018). Alliance rupture repair: A meta-analysis. Psychotherapy, 55(4), 508–519. https://doi.org/10.1037/pst0000185
Feliciano, I. F. G., Staples, L., Scott, A., Jones, M. P., Hadjistavropoulos, H., Titov, N., & Dear, B. F. (2026). Therapist effects in internet-delivered cognitive behavior therapy for anxiety and depression. Journal of Consulting and Clinical Psychology, 94(2), 88–100. https://doi.org/10.1037/ccp0000994
Goldberg, S. B., Rousmaniere, T., Miller, S. D., Whipple, J., Nielsen, S. L., Hoyt, W. T., & Wampold, B. E. (2016). Do psychotherapists improve with time and experience? A longitudinal analysis of outcomes in a clinical setting. Journal of Counseling Psychology, 63(1), 1–11. https://doi.org/10.1037/cou0000131
Johns, R. G., Barkham, M., Kellett, S., & Saxon, D. (2019). A systematic review of therapist effects: A critical narrative update and refinement to review. Clinical Psychology Review, 67, 78–93. https://doi.org/10.1016/j.cpr.2018.08.004
Kent, J. D., & McCarthy, M. T. (2016). Holistic review in graduate admissions: A report from the Council of Graduate Schools. Council of Graduate Schools.
King, R. J., Orr, J. A., Poulsen, B., Giacomantonio, S. G., & Haden, C. (2017). Understanding the therapist contribution to psychotherapy outcome: A meta-analytic approach. Administration and Policy in Mental Health and Mental Health Services Research, 44(5), 664–680. https://doi.org/10.1007/s10488-016-0783-9
Kraus, D. R., Castonguay, L., Boswell, J. F., Nordberg, S. S., & Hayes, J. A. (2011). Therapist effectiveness: Implications for accountability and patient care. Psychotherapy Research, 21(3), 267–276. https://doi.org/10.1080/10503307.2011.563249
Levashina, J., Hartwell, C. J., Morgeson, F. P., & Campion, M. A. (2014). The structured employment interview: Narrative and quantitative review of the research literature. Personnel Psychology, 67(1), 241–293. https://doi.org/10.1111/peps.12052
Miller, S. D., Hubble, M. A., & Chow, D. (2020). Better results: Using deliberate practice to improve therapeutic effectiveness. American Psychological Association.
Norcross, J. C., Bike, D. H., Evans, K. L., & Schatz, D. M. (2008). Psychotherapists who abstain from personal therapy: Do they practice what they preach? Journal of Clinical Psychology, 64(12), 1368–1376. https://doi.org/10.1002/jclp.20531
Nye, C. D., & Ryan, A. M. (2023). Improving graduate-school admissions by expanding rather than eliminating predictors. Perspectives on Psychological Science, 18(1), 54–60. https://doi.org/10.1177/17456916221105359
Okiishi, J. C., Lambert, M. J., Nielsen, S. L., & Ogles, B. M. (2003). Waiting for supershrink: An empirical analysis of therapist effects. Clinical Psychology & Psychotherapy, 10(6), 361–373. https://doi.org/10.1002/cpp.383
Rousmaniere, T. (2016). Deliberate practice for psychotherapists: A guide to improving clinical effectiveness. Routledge.
Saxon, D., & Barkham, M. (2012). Patterns of therapist variability: Therapist effects and the contribution of patient severity and risk. Journal of Consulting and Clinical Psychology, 80(4), 535–546. https://doi.org/10.1037/a0028898
Tracey, T. J. G., Wampold, B. E., Lichtenberg, J. W., & Goodyear, R. K. (2014). Expertise in psychotherapy: An elusive goal? American Psychologist, 69(3), 218–229. https://doi.org/10.1037/a0035099
Wampold, B. E. (2015). How important are the common factors in psychotherapy? An update. World Psychiatry, 14(3), 270–277. https://doi.org/10.1002/wps.20238
Wampold, B. E., & Brown, G. S. (2005). Estimating variability in outcomes attributable to therapists: A naturalistic study of outcomes in managed care. Journal of Consulting and Clinical Psychology, 73(5), 914–923. https://doi.org/10.1037/0022-006X.73.5.914
Wampold, B. E., & Imel, Z. E. (2015). The great psychotherapy debate: The evidence for what makes psychotherapy work (2nd ed.). Routledge.
Woo, S. E., LeBreton, J. M., Keith, M. G., & Tay, L. (2023). Bias, fairness, and validity in graduate-school admissions: A psychometric perspective. Perspectives on Psychological Science, 18(1), 3–31. https://doi.org/10.1177/17456916211070833
Wood, M. (2025, April 17). Student notebook: Is the clinical psychology PhD admissions process scientific? APS Observer. https://www.psychologicalscience.org/publications/observer/student-notebook-admissions-wood.html
Zerubavel, N., & O’Dougherty Wright, M. (2012). The dilemma of the wounded healer. Psychotherapy, 49(4), 482–491. https://doi.org/10.1037/a0027824