← D2 overviewSection 1 of 7

The Screener Landscape: What Schools Measure Today

What should a school behavioral health performance management system count, and how does it know the counts are trustworthy? Before answering either question, it helps to know what schools already count. Districts entering behavioral health screening do not start from a blank page: a substantial instrument landscape already exists, built across four decades of clinical and school psychology research, and BHPM's indicator design has to sit deliberately inside that landscape rather than around it. That landscape did not develop as one coherent design; it accumulated in three distinct waves, each answering a different question. Adult clinical medicine produced checklist screeners meant to trigger a diagnosis. School psychology, working a decade or two later and often against those same clinical instruments as a baseline, built its own tools meant to trigger a tiered-support decision rather than a clinical one. And state infrastructure built a third kind of instrument meant to describe a population rather than flag any individual student at all. Treating these as interchangeable rows in one comparison table, rather than as three different tools built to answer three different questions, misreads what each one can defensibly be asked to do. The eight instruments surveyed here also split along a second, orthogonal axis worth naming up front: who is asked to report. Most are self-report or teacher-report checklists resting on a single informant's judgment; BASC-3 BESS and the SDQ each gather more than one informant's perspective on the same student rather than relying on just one; the Columbia-Suicide Severity Rating Scale substitutes a clinician's structured interview for either; and CHKS is answered by the student but never resolves to an individual score at all. None relies on a performance task or an observed-behavior sample independent of someone's rating — a gap that matters later in this section, because it means every accuracy figure quoted below is an accuracy figure about how well one kind of report predicts another, not about how well any of them capture behavior directly.

Two of the most widely used instruments in that landscape were never built for schools at all. The PHQ-9 and the GAD-7 are adult-primary-care screeners, validated against structured clinical interviews and later adapted downward into adolescent and school-based use because nothing purpose-built displaced them. Both are self-report checklists mapped directly onto DSM criteria, both return a single cutoff score with a published sensitivity/specificity profile, and both were designed to trigger one specific downstream action: a clinical referral decision. The PHQ-9's original validation, run on 6,000 primary-care and ob-gyn patients with a criterion sample of 580 assessed by structured interview, found that a cutoff of 10 or above identified major depression with 88 percent sensitivity and 88 percent specificity, and the same study found severity climbing in step with worse functional status across every SF-20 subscale and with higher sick-day and healthcare-utilization rates — evidence that the instrument works as both a categorical diagnostic algorithm and a continuous severity scale, not only the former. The GAD-7 was validated the same way for generalized anxiety, with a cutoff of 10 yielding 89 percent sensitivity and 82 percent specificity, and its validation work established the now-standard "past two weeks" self-report window and confirmed good internal consistency alongside factorial, criterion, and procedural validity across a large primary-care cohort. Their durability across two decades of behavioral health practice is real, but so is the fact that they were built to answer a narrower question than a school system asks: not "how is this student doing across the domains a district's behavioral health framework cares about," but "does this specific respondent meet criteria for this specific disorder today."

The instruments built for schools take a different shape. BASC-3's Behavioral and Emotional Screening System (BESS) is the most commercially entrenched of them, with a Teacher Form whose higher-order factor structure holds up under confirmatory analysis on a K-5 sample of 1,472 students and shows configural, metric, and scalar invariance across student age — invariance solid enough that the same study found BESS scores predicting academic achievement, behavioral engagement, and school climate outcomes, not just concurrent risk status. A Student Form extends that validity evidence to self-report: in a sample of 737 urban students in grades 4 through 8, its bifactor structure held and measurement equivalence was demonstrated specifically across White, Black, and Latinx students and across gender , a fairness finding that matters directly to whether a commercial screener's psychometrics travel across a district's own student population rather than only across the sample it was validated on. SAEBRS took a similar structural-validity path from a different starting point. It grew out of an initial 12-item Social and Academic Behavior Risk Screener built through a content-validation process plus reliability and exploratory factor analysis on 243 students rated by 54 teachers, with an item pool drawn deliberately from developmental research on behavior-problem trajectories and social and academic competence models rather than from an existing clinical instrument . From there it added a multiple-gating procedure that combines teacher nomination with screener scores, and a two-study investigation anchored by an 864-student sample demonstrated acceptable-to-optimal internal consistency, concurrent validity, and diagnostic accuracy across elementary and middle school settings . It was later subjected to an argument-based validation using bifactor item response theory and structural equation modeling — a bifactor model that outperformed unidimensional and correlated-factors alternatives, discriminated moderate- from high-risk students, held up under a hierarchical-omega reliability check on its Total Behavior score, and was tested for convergent and discriminant relations directly against the BASC BESS itself , a maturity benchmark few school screeners reach and one that effectively cross-validates two of this landscape's most entrenched instruments against each other rather than asking either to stand alone. Both instruments carry that commercial, gated-procedure weight forward into the cost-and-time comparison below — a dimension the psychometric literature above does not measure at all, but one that shapes which districts can realistically field either instrument at scale.

The Strengths and Difficulties Questionnaire took the opposite path to the same destination. Rather than commercial packaging or a gated procedure, its 25-item, five-scale design was tested in a clinical sample of 403 children against the established Rutter questionnaires, and a five-factor structure — four difficulty scales plus a prosocial-behavior scale — replicated across parent and teacher raters, not just one informant type . That multi-informant replication, combined with free and noncommercial distribution, made the SDQ the most internationally adopted screener in this landscape. SRSS-IE occupies similar territory domestically: a free, brief, dual-scale teacher rating whose internalizing and externalizing risk classifications predicted subsequent academic and behavioral outcomes across 4,465 elementary students in 14 schools , with the dual structure itself shown to carry independent predictive value over a single-dimension screener — evidence that splitting risk into two channels rather than blending it into one score is not just conceptually tidier, it captures something a single score misses.

Two more instruments sit at the edges of this landscape rather than inside its center, and both matter to how BHPM should think about scope. The Columbia-Suicide Severity Rating Scale is not a universal screener at all; it is a semi-structured, clinician-administered interview built for one purpose, distinguishing degrees of suicide risk severity, and its validation across three multisite adolescent and adult clinical trials established convergent and divergent validity against existing suicidality measures and found that lifetime worst-point ideation with intent or a plan predicted roughly four times the future attempt risk of current ideation alone — a finding robust enough across three independent samples that the instrument became the standard reference not only for clinical trials but for school crisis-response protocols specifically. It belongs at the individual crisis-response tier, not the universal-screening tier, and a metrics framework that conflates the two risks either over-medicalizing a screening program or under-resourcing a crisis one. At the opposite end sits California's own state-sponsored infrastructure, the California Healthy Kids Survey within the CalSCHLS system: a standardized, anonymous survey administered by WestEd for the California Department of Education since 1997, developed together with Duerr Evaluation Resources and expert advisory committees, with a Core Module required in grades 7 and 9 and current participation above 69 percent of California districts. CHKS is deliberately built to produce no individual-level score at all . That design choice, anonymity purchased at the cost of individual actionability, is not a limitation the instrument's designers overlooked; it is the tradeoff the rest of this section turns on.

Four structural gaps run underneath this entire landscape, and each one bears directly on what a BHPM indicator set has to be built to withstand. The first is a method-effect problem inherent to self-report and teacher-report measurement itself. Duckworth and Yeager's review of how personal qualities beyond cognitive ability get measured, using self-control as its worked case study, found that no single method — self-report, other-report, or performance task — is free of bias, and that terminology debates over what to call these qualities obscure far more agreement than disagreement about which specific qualities actually matter . Their conclusion was not to identify the least-biased method and standardize on it; every method they examined carried some version of the same limitation, which is why they recommended triangulating across methods rather than trusting any single one at high-stakes scale. Every instrument named above inherits some version of this limitation regardless of how well its factor structure replicates, and the informant split named at the top of this section is exactly where that limitation lives: a teacher-report instrument like SAEBRS or the BESS Teacher Form is vulnerable to whatever a teacher can and cannot observe, a self-report instrument like the PHQ-9 or GAD-7 is vulnerable to whatever a student is willing to disclose on a checklist, and even the two instruments that gather more than one informant's view, BASC-3 BESS and the SDQ, report each informant's rating separately rather than combining them into a single score that could offset either one's blind spot.

The second is the anonymity-actionability tradeoff CHKS makes explicit but does not resolve. A screening instrument can be built to protect a respondent's identity or built to trigger an individual follow-up action; the instruments surveyed here sit at different points along that line — CHKS at one pole, C-SSRS at the other, and the six school-based universal screeners in between at varying degrees of individual actionability — and a district cannot simply average across them without deciding, deliberately, where its own program needs to sit. That decision is entangled with the informant-type split described above rather than separate from it: an individually actionable score requires knowing which student produced it, which in turn requires an identified informant, so a district cannot adopt CHKS's anonymity and still expect PHQ-9-style individual follow-up out of the same instrument. The tradeoff is a design constraint on the whole indicator set, not a property of any one screener that a district could simply screen for and avoid.

The third gap is acceptability, and it has its own growing evidence base, now with a resourcing dimension attached to it. Two systematic reviews from 2019 and 2020 found the evidence on school-based mental health identification mixed at best: feasibility and acceptability findings split along stakeholder lines, with parents, students, and health professionals generally supportive and school staff more skeptical, largely over role-scope concerns, though the same review noted that near-universal reach and proximity to a school's existing support resources are exactly what make screening worth the friction in the first place , while a companion review found the effectiveness and cost-effectiveness evidence base for these programs still limited and methodologically uneven, and identified specific gaps in how that research has so far been designed . A 2025 systematic review applying the Theoretical Framework of Acceptability specifically to universal mental health screening brought that evidence current, using the Mixed Methods Appraisal Tool to grade study quality and a narrative synthesis to separate acceptability across students, caregivers, and school staff as three distinct stakeholder groups with distinct drivers and barriers, rather than compressing all three into a single undifferentiated "acceptability" score . A district-level implementation study adds a piece none of the three reviews measures directly: a CDC-funded district screening effort found that fewer than 15 percent of U.S. schools had, at that point, implemented any systematic behavioral-health screening procedure at all, and that even a funded, university-supported rollout reached only 68.6 and 68.5 percent student completion across two successive years, dropping to 56 and 57 percent at the high school level specifically . Acceptability, in other words, is not only a matter of whether stakeholders approve of screening in principle; it is also a matter of whether a district without grant-funded technology and data-management support can actually get the screener into most students' hands in the first place.

The fourth gap is the one the psychometric literature states most concretely, and it is a screening-to-service gap hiding inside an aggregate accuracy number. The CDC's multi-site Project to Learn About Youth-Mental Health dataset tested how well teacher-completed BASC-2-BESS and SDQ scores, gathered across four districts, predicted parent-report-confirmed diagnoses assessed through a structured clinical interview, in a sample of 1,054 K-12 students: both instruments reached an AUC of .73 for externalizing disorders but only .58 for internalizing ones, leading the study's own authors to conclude that both screeners carry only modest predictive utility against a rigorous diagnostic criterion . A single reported accuracy figure for either screener would flatten that gap invisibly. The instruments in this landscape are, on average, reasonably good at flagging the students whose distress is visible to a teacher and considerably weaker at flagging the students whose distress is not, which is precisely the population a behavioral health performance management system exists to find. Siceloff and colleagues' feasibility work reaches a similar conclusion from the implementation side rather than the psychometric one: the universal screeners in their district-wide sample showed stronger sensitivity for externalizing problems than for internalizing ones, which is part of why newer dual-scale instruments like SRSS-IE were built to separate the two risk channels rather than blend them into a single score . Two independent evidence bases, one measuring diagnostic accuracy in a controlled comparison and the other measuring what happens in an actual district rollout, converge on the same finding: internalizing risk is the harder detection problem here, not a minor addendum to externalizing risk.

None of the four structural gaps above tracks cleanly onto cost, and that is itself worth stating plainly before the comparison table does the rest of the work. The free instruments in this landscape, the SDQ, SRSS-IE, and CHKS, are not free of the method-effect, acceptability, or screening-to-service problems that burden the commercial and clinical instruments beside them; they simply carry those same evidentiary limitations at a lower price and a lower staff-time cost. A district weighing BASC-3 BESS's stronger cross-validated structural evidence against SRSS-IE's lower cost is not trading rigor for savings in any simple sense — it is trading one already-imperfect method for another, at a different price point, inside the same set of structural constraints. The table below lays these eight instruments side by side across the dimensions that matter for indicator design, not just psychometric quality in isolation: who reports, in what format, with what accuracy evidence, whether that accuracy translates into an individually actionable score, and what it costs a district in dollars and staff time to run. Selecting each row opens its full source profile.

Screener landscape comparison: instrument x informant x format x psychometrics x individual actionability x cost/time

InstrumentInformantFormatPsychometricsIndividual actionabilityCost & time
PHQ-9Student self-report9-item Likert, past 2 weeks, DSM-IV-mapped depression severity≥10 cutoff: 88% sensitivity / 88% specificity against a structured clinical interview (n=580)Individual score plus diagnostic-algorithm cutoff; built for direct clinical follow-upPublic domain; roughly 5 minutes to administer
GAD-7Student self-report7-item Likert, past 2 weeks, generalized anxiety severityCutoff of 10: 89% sensitivity / 82% specificity in a large primary-care validation sampleIndividual score and cutoff; companion instrument to PHQ-9 for clinical triagePublic domain; roughly 3-5 minutes to administer
BASC-3 BESS (Teacher and Student Forms)Teacher rating (Teacher Form); student self-report (Student Form)Brief multi-item behavioral/emotional risk rating; higher-order factor structure (Teacher Form), bifactor structure (Student Form)Confirmed factor structure with measurement invariance across student age (Teacher Form) and across race/ethnicity and gender (Student Form)Widely adopted for individual risk tiering, but the predecessor BASC-2-BESS teacher form shows meaningfully weaker predictive accuracy for internalizing disorders (AUC .58) than externalizing (AUC .73) against clinician diagnosisProprietary, publisher-licensed instrument; brief teacher or student rating
SAEBRS (Social, Academic, and Emotional Behavior Risk Screener)Teacher rating12-item multi-domain (social, academic, emotional) risk screener, developed from the earlier SABRSAcceptable-to-optimal internal consistency and concurrent validity; bifactor IRT/SEM modeling supports score interpretation and discriminates moderate- from high-risk studentsA multiple-gating procedure combining teacher nomination with SAEBRS scores supports tiered individual identificationBrief teacher rating; the multiple-gating procedure adds a second screening step
Strengths and Difficulties Questionnaire (SDQ)Multi-informant: student, parent, and teacher versions25-item, five-scale instrument (emotional symptoms, conduct problems, hyperactivity/inattention, peer problems, prosocial behavior)Strong correspondence with the established Rutter questionnaires; five-factor structure replicated across ratersTeacher-rated SDQ shows the same predictive-accuracy pattern as BASC-2-BESS against clinical diagnosis: AUC .73 for externalizing disorders vs. .58 for internalizing disordersFree, noncommercial distribution; a major driver of its international adoption
SRSS-IE (Student Risk Screening Scale — Internalizing and Externalizing)Teacher ratingBrief dual-scale risk screener covering both internalizing and externalizing riskRisk classifications predicted subsequent academic and behavioral outcomes in a sample of 4,465 elementary studentsDual-dimension score supports differentiated tiering; predictive validity evidence is strongest at the elementary levelFree; brief teacher rating designed for large-scale universal screening
Columbia-Suicide Severity Rating Scale (C-SSRS)Student, via trained-clinician-administered interviewSemi-structured interview assessing suicidal ideation and behavior severityLifetime worst-point ideation with intent or plan predicts roughly 4x the future attempt risk of current ideation aloneAn individual, crisis-tier instrument built for direct safety-planning response, not universal classroom screeningRequires a trained administrator; not a self-administered classroom tool
California Healthy Kids Survey (CHKS), within the CalSCHLS systemStudent self-reportAnonymous statewide survey with a standardized Core Module (required in grades 7 and 9)State-sponsored, standardized administration in place since 1997, currently used by more than 69% of California districtsNone at the individual level by design: anonymized, aggregate-only reportingState-sponsored infrastructure; no per-student licensing cost to districts

Bottom Line

BHPM's indicator set holds a deliberate position inside this landscape: a complement to the eight instruments surveyed above, not a replacement for any of them. That position starts with the anonymity-actionability line running through the landscape — CHKS at one pole, the C-SSRS at the other, and the six school-based universal screeners falling at varying degrees of individual actionability between them. Naming that line explicitly, and showing where an indicator sits on it, is what lets a reader see which tradeoff is being made rather than discover it by implication. The PLAY-MH internalizing/externalizing accuracy gap (AUC .73 vs .58), reinforced by Siceloff and colleagues' district-level finding that universal screeners under-detect internalizing risk in practice as well as in controlled comparison, is the evidence that grounds the case for language-based measurement in Section 2: a concrete detection failure, not a general appeal to innovation.

IMPACTER PathwaySanta Clara County Office of Education

This site publishes no student-level data. All figures are aggregate, de-identified, and reviewed under FERPA, COPPA, and California student privacy law (AB 1584 / SOPIPA).

Prepared in support of SCCOE under BHSOAC Contract No. 25BHSOAC019.