← D2 overviewSection 4 of 7

Process & Outcome Indicators in MTSS-Aligned Behavioral Health

A four-year study tracking twenty-nine schools implementing Multi-Tiered Systems of Support found something specific enough to build a measurement architecture on: schools with higher fidelity in the behavior-support domain of their MTSS implementation logged significantly fewer suspension events than matched comparison schools, while schools with higher fidelity in the academic domain, reading and math instruction delivered through the same tiered structure, showed no such reliable relationship to achievement gains . Scott, Gage, Hirn, Lingo, and Burt reached that finding using the Academic and Behavior Response to Intervention School Assessment, an instrument built to score fidelity separately on behavior-domain and academic-domain subscales rather than producing one blended implementation score, and they tracked those subscales against suspension counts and state achievement data across the same twenty-nine schools each year for four years. Their own conclusion is blunt: school-wide strategies that do not specifically involve effective instruction in academic areas are unlikely to raise academic achievement on their own, no matter how faithfully the behavior-support scaffolding around that instruction is run. That is not a caveat buried in a footnote. It is the finding that decides how BHPM's indicator set has to be built: fidelity and outcome are not the same measurement, one domain's fidelity does not transfer its predictive power to a different outcome domain, and an indicator system that collapsed "the program ran as designed" into "the program worked" would have been wrong about exactly the distinction this study was built to test.

That separation is easier to state than to operationalize, because school-wide positive behavior support is not a single discrete practice with a clean efficacy-trial literature behind it — it is a systems-level, multi-component framework layering universal screening, tiered response, data-based decision rules, and ongoing coaching across three tiers of intensity. Horner, Sugai, and Anderson confronted that complexity directly by proposing six criteria for defining any educational practice clearly enough to assess its evidence base at all: that the practice be operationally defined, that the settings where it is expected to work be specified, that its target population be defined, that the qualifications required to use it successfully be defined, that its expected outcomes be defined, and that the conceptual theory explaining why it should work be stated . Applied to SWPBS, those criteria forced the reviewers to treat the practice as a constellation — three tiers of intervention plus the organizational systems that sustain them — rather than pretend it reduces to one procedure with one efficacy trial behind it. Their evidence review covered 46 peer-reviewed studies published between 2000 and 2009, split across the three tiers, twenty at Tier 1, thirteen each at Tiers 2 and 3, and graded against a second, five-part standard: whether the practice and participants were defined with enough precision to replicate, whether the measures used were valid and reliable, whether the designs were rigorous, whether documented effects came without harmful side effects, and whether those effects held up over time. Their conclusion, tier by tier, is where the caution BHPM needs actually lives. Tier 1 implementation fidelity was found to be achievable at meaningful scale — elementary school teams receiving four to six days of training reached eighty percent or better fidelity on a standardized evaluation tool in a randomized waitlist-controlled trial — and associated with real reductions in problem behavior and improved organizational health. On academic outcomes specifically, the review's own language is deliberately hedged: SWPBS implementation is "promisingly, but not definitively" associated with increased academic outcomes, and the authors state outright that it is premature to claim SWPBS is causally associated with improved academic achievement, because the underlying mechanism runs through improved instructional engagement rather than through behavior support directly producing learning. That caution, published independently of Scott and Gage's later fidelity study, arrives at close to the same place from a different kind of evidence — a synthesis of dozens of studies rather than one four-year panel — converging on the conclusion that behavior-domain fidelity and academic outcome are not the same claim and should not be reported as though they were.

The same review flags two further limitations worth carrying into BHPM's own feasibility notes rather than smoothing over. Tier 2 and Tier 3 evidence rests overwhelmingly on single-case designs rather than the randomized and quasi-experimental group trials available for Tier 1, a real difference in methodological weight that a performance system reporting all three tiers on one dashboard should not obscure. And implementation at meaningful fidelity is measurably harder in high schools than in the elementary settings most of this evidence base was built on, a pattern the reviewers describe as consistent across the roughly 1,000 high schools that had adopted SWPBS by the time of their review. Neither limitation is a reason to abandon a tiered indicator design; both are reasons a site-level fidelity report should say which tier and which school level it is describing rather than publishing one undifferentiated score.

Fidelity itself needs a validated instrument before any of this can be scored, and the field-standard measure is the Tiered Fidelity Inventory, developed and maintained under the OSEP-funded Center on PBIS specifically to assess fidelity across Tier 1 through Tier 3 in one linked tool rather than three unrelated instruments. McIntosh and colleagues established its technical adequacy across three separate linked studies: a content-validity study confirming the instrument's items actually represent the practices each tier is supposed to measure, a usability and reliability study, and a large-scale validation drawing on a broad multi-state sample . Across those three studies the TFI showed strong construct validity at all three tiers, strong interrater agreement, stable two-week test-retest reliability, and strong correspondence with the SWPBIS fidelity measures that predated it. That last property matters for a system trying to connect to existing district practice: a site that already has fidelity data collected under an older SWPBIS measure is not starting from zero when it adopts TFI-aligned scoring, because the two are validated against each other rather than measuring unrelated constructs.

The outcome side of the ledger carries its own, sharper caution, drawn from the one independent randomized controlled trial in this space large enough to trust. An IES-funded evaluation of MTSS-B, the behavior-focused variant of Multi-Tiered Systems of Support, randomized roughly eighty-nine elementary schools across nine districts that had not previously implemented PBIS, and its results are not the clean success story a vendor slide would prefer . Averaged across all students, the program did not improve disruptive behavior, other behavior outcomes, or achievement. Where effects appeared, they were concentrated entirely among the roughly fifteen percent of students identified at baseline as struggling most with behavior: that group showed improved reading scores and reduced disruptive behavior, but only by the second year of full implementation, not the first. Some broader improvements in classroom management and school climate measures did appear at the school level. And even the one clear win did not fully hold — the reading effect for the highest-need students did not persist into the follow-up year after program support ended. An indicator system built only to detect aggregate, school-wide movement would miss this trial's entire finding, because the finding lives inside a subgroup that a school-wide average erases, and it would also miss the timing: an outcome measured only in year one of implementation would have reported a null result the trial's own second-year data contradicts. BHPM's outcome and contextual indicators are therefore designed to be disaggregated by subgroup and need tier from the start, not retrofitted for equity reporting after the fact, and to be reported on a timeline long enough to distinguish a first-year null from a second-year effect rather than closing the book early.

Read against each other, these three sources do not fully agree, and the disagreement is itself informative. Scott and Gage's finding comes from a quasi-experimental panel comparing schools against matched comparisons, not random assignment; Horner, Sugai, and Anderson's caution about academic outcomes is a synthesis judgment drawn across dozens of underlying studies of uneven design quality; and the MTSS-B brief is the one true randomized trial in the group, the design least vulnerable to the selection effects that can inflate fidelity-outcome associations in a matched-comparison study. That the RCT's own aggregate result is null while the quasi-experimental and synthesized evidence report real associations is not a contradiction to paper over — it is exactly the pattern a well-built indicator system should expect and survive. A system that reported only the matched-comparison finding would overstate what fidelity data alone can promise; a system that reported only the RCT's null aggregate result would erase the subgroup effect that same trial documents. BHPM's three-family indicator structure — process measured on its own terms, outcome measured and disaggregated on its own terms — is a direct response to that pattern, not an attempt to resolve it into one number.

Three indicator families follow from that literature, and the boundaries between them track the design tension above rather than arbitrary bookkeeping. Process indicators measure whether the system is running as designed: screening completion, implementation fidelity scored against a validated instrument, training participation, and data-review cadence — the domain where Scott and Gage's finding and the Horner-Sugai-Anderson framework criteria are strongest. Outcome indicators measure what the system is meant to produce: tier distribution and movement over time, domain-level growth on measured competencies, and covariates like attendance and exclusionary discipline that the fidelity and RCT literature both link, in different ways, to behavior-domain implementation. Contextual indicators describe who is being measured and who is receiving services: demographic composition of the screened population, baseline tier and need-level distribution, and service-receipt cross-tabs by subgroup — the domain the MTSS-B trial's subgroup-concentrated effect makes non-optional rather than a nice-to-have. None of the three families is optional. A process-only report cannot answer whether the work mattered; an outcome-only report, especially one that only ever reports the school-wide average, cannot answer for whom it mattered or explain why results moved.

Those three families map directly onto the four reporting domains the Commission has asked for: services provided, student demographics and characteristics, service utilization, and outcomes and impact. Process indicators populate the services-provided domain. Contextual indicators populate the demographics domain and travel alongside outcome indicators wherever disaggregation is required — the direct design response to the MTSS-B trial's subgroup finding above. Service utilization sits partly inside BHPM's own platform data and partly inside the state and federal reporting infrastructure already built around school-based behavioral health billing, which is why two comparable initiative reporting frames, not academic studies, anchor that domain below rather than a fifth research citation. Outcome indicators close the loop back to the Commission's central accountability question, disaggregated on the timeline the RCT evidence above argues for rather than collapsed into a single annual figure.

The diagram below re-sorts the same working indicator set by the Commission's own four domain names and order, rather than by the three literature-driven families above. In this candidate set the two groupings line up one-to-one — process indicators are the services-provided domain, contextual indicators are the demographics domain, and the outcome and outcome-covariate indicators together are the outcomes domain — but the families and the domains answer different questions and won't necessarily stay aligned as the indicator set grows: the families explain why an indicator belongs in the literature base, the domains are how the Commission has asked to receive it.

Domain alignment — candidate indicators sorted into the Commission’s four reporting domains

Services Provided
Students Served
Service Utilization
Outcomes & Impact
1Services Provided

4 indicators

  • Screening administration completion rate
    Process
  • Implementation fidelity score (TFI-aligned)
    Process
  • Staff and educator training participation rate
    Process
  • Data-review cycle completion
    Process
2Students Served

Student demographics & characteristics

3 indicators

  • Screened population profile (grade band, language, program status)
    Contextual
  • Baseline tier and need-level distribution
    Contextual
  • English learner and IEP/504 status cross-tab on screening participation
    Contextual
3Service Utilization

3 indicators

  • Referral-to-service rate
    Utilization
  • Billable screening encounters (CPT 96127)
    Utilization
  • School-linked provider network utilization
    Utilization
4Outcomes & Impact

3 indicators

  • Tier movement over time
    Outcome
  • Exclusionary discipline rate (suspensions and expulsions)
    Outcome covariate
  • Domain-level growth on measured competencies
    Outcome

The table below is the working candidate indicator set, and it is the skeleton BHPM's indicator reporting is built on rather than an illustrative sample: each row states the literature basis for including the indicator, the collection specification a site would actually implement, the Commission domain it reports into, and a feasibility note distinguishing what BHPM can measure today from what depends on a data-sharing agreement or a district's own billing infrastructure. Clicking a row opens the full profile.

Candidate indicator set — S4, mapped to the Commission's four reporting domains

IndicatorIndicator familyCommission domainLiterature basisCollection specFeasibility note
Screening administration completion rateProcessServices providedFidelity-to-outcome logic: implementation fidelity is a distinct, prior condition to outcome movement, not a proxy for it.Platform-logged screening completions divided by enrolled eligible students, per school, per reporting window.High — captured automatically by the platform; no manual entry.
Implementation fidelity score (TFI-aligned)ProcessServices providedTiered Fidelity Inventory is the field-standard, psychometrically validated tiered fidelity-of-implementation measure.Periodic TFI or TFI-aligned local fidelity check scored across Tier 1-3 domains by a trained site coordinator.Moderate — requires a trained local scorer and a fixed administration cadence.
Staff and educator training participation rateProcessServices providedReporting category drawn from a comparable federal school behavioral health program's own public accounting of training reach.Count of staff completing BHPM onboarding and technical-assistance modules divided by the site staff roster.High — logged through existing TTA delivery tracking.
Data-review cycle completionProcessServices providedFidelity measurement frameworks for tiered systems treat a regular data-review routine as a scoreable implementation component, not an assumed byproduct of adoption.Count of scheduled site data-review meetings held, against the cadence set in the site's implementation plan.High once a site's review calendar is established in D1.
Screened population profile (grade band, language, program status)ContextualStudent demographics and characteristicsIndependent RCT evidence that effects concentrate among highest-need students argues for reporting the screened denominator's composition, not just its size.Enrollment-derived demographic fields cross-tabulated against screening completion, aggregated to school level.High — fields already reside in roster feeds.
Baseline tier and need-level distributionContextualStudent demographics and characteristicsTiered-system framing for evaluating school-wide behavior support treats tier placement as the organizing unit of analysis.Count of students at each MTSS tier at first screening administration, by school.High.
English learner and IEP/504 status cross-tab on screening participationContextualStudent demographics and characteristicsSubgroup-concentrated effect sizes in the IES MTSS-B trial motivate disaggregating participation, not only outcomes, by subgroup.Screening participation rate disaggregated by English learner and IEP/504 status.Moderate — depends on the site's data-sharing agreement covering those fields.
Referral-to-service rateUtilizationService utilizationComparable federal program reporting category: referral volume as a standing public accountability figure for a school behavioral health initiative.Count of students referred for behavioral health service following screening, tracked to service entry where data-sharing allows.Moderate — depends on district-provider data-sharing agreements.
Billable screening encounters (CPT 96127)UtilizationService utilizationCPT 96127 is the AMA-maintained billing code for a brief, standardized-instrument emotional/behavioral assessment with scoring and documentation.Count of standardized-instrument screening encounters coded and documented under CPT 96127.High where a site's billing infrastructure already supports the code; low where none exists yet.
School-linked provider network utilizationUtilizationService utilizationCalifornia's statewide CYBHI Fee Schedule Program, authorized under Welfare and Institutions Code Section 5961.4, is the in-state precedent for reporting reimbursed school-linked encounters.Count and rate of reimbursed encounters under the DHCS fee schedule at CYBHI-enrolled sites.Depends on the district's CYBHI cohort enrollment status.
Tier movement over timeOutcomeOutcomes and impactBehavior-domain implementation fidelity is empirically associated with downstream movement in student-level indicators across a multi-year, multi-school sample.Change in tier placement between baseline and each subsequent screening window, aggregated to school level with small-cell suppression.High once two screening administrations exist.
Exclusionary discipline rate (suspensions and expulsions)Outcome covariateOutcomes and impactBehavior-fidelity gains associate with reduced exclusionary discipline; an independent randomized trial separately found reduced disruptive behavior concentrated among the highest-need students.District-reported suspension and expulsion counts, cross-referenced against BHPM tier and fidelity data.High — already collected for state reporting.
Domain-level growth on measured competenciesOutcomeOutcomes and impactApplies the same process/outcome separation established in the MTSS fidelity literature to IMPACTER's own longitudinal attribute scores.Within-student change in scored attribute levels across administrations.High — computed automatically by the scoring pipeline.

Service utilization is the domain where BHPM's own platform data runs out fastest, because a referral or a billed encounter is only visible to a school-based measurement system when the receiving service sits inside the same reporting relationship. Two comparable initiative reporting frames were chosen precisely because they already report at this level in public, not because either one evaluates a program's causal effect. SAMHSA's Project AWARE, a federal grant program that builds school-based behavioral health infrastructure through state, county, and district partnerships rather than delivering services directly, reports 2018-2024 aggregate outcomes that read like a preview of what BHPM's own service-utilization and outcomes domains will eventually publish: 3,492 organizations entered formal partnership agreements, 1,509 state or local policy changes were attributed to the program, 1,384,460 school staff, families, and community members participated in related trainings, and 327,332 students were referred for mental health or related services . Those figures describe a federal grantee's own program accounting — the same category of reporting a comparable BHPM rollup would eventually produce — not a peer-reviewed evaluation of what the funding caused; SAMHSA's own page presents them as cumulative reach, not as an effect size against a comparison condition. BHPM's indicator set borrows only the reporting categories, partnerships formed, staff trained, students referred, not the causal claim, and the distinction is worth stating plainly because a partnership-agreement count and a randomized trial's effect size answer different questions even when they appear on the same kind of program dashboard.

California offers a state-level precedent with more legal specificity than a federal grant program's aggregate reach. The Children and Youth Behavioral Health Initiative's Fee Schedule Program, authorized under Welfare and Institutions Code Section 5961.4, requires commercial health plans and Medi-Cal to reimburse school-linked behavioral health providers at published rates regardless of network status, effective January 1, 2024 . DHCS has rolled the program out in cohorts — 484 local education agencies and five institutions of higher education across Cohorts 1 through 4, with Cohort 6 participants announced February 27, 2026 following an operational-readiness review covering each applicant's Medi-Cal enrollment status, service-delivery capacity, data-collection practices, and billing infrastructure . That readiness review is itself a useful precedent for BHPM's own feasibility notes: DHCS does not treat "the district joined the program" as equivalent to "the district can bill against it," and a BHPM service-utilization indicator should not either. The billing-code detail that ties screening activity to a documented, reimbursable service touchpoint is CPT 96127, the standardized code for a brief emotional or behavioral assessment with scoring and documentation, billed per standardized instrument administered, up to four units per encounter rather than by time spent, and commonly billed alongside instruments like the PHQ-9 and GAD-7 . Wherever a BHPM site's own billing infrastructure already supports that code, screening volume and billed-encounter volume become the same countable event — the strongest possible link between a process indicator and a utilization indicator, and one that will only exist at sites whose CYBHI enrollment and billing systems have cleared the same operational bar DHCS itself applies.

None of this is a wish list dressed up as a table. Every row in the indicator set above states what BHPM can measure today under current site infrastructure and what still depends on a data-sharing agreement, a district's CYBHI enrollment status, or a billing system that does not yet exist at every site. That candor is deliberate, and it follows directly from the pattern the literature above establishes: fidelity data, synthesis judgments, and randomized evidence do not always agree with each other, subgroup effects can hide inside school-wide averages, and a performance management system that overpromises data it cannot yet produce, or certainty a single study cannot yet support, fails the same accountability test it exists to serve.

Bottom Line

Process and outcome are two separate, independently reportable indicator families, not one blended score. The Scott-and-Gage fidelity study, the Horner-Sugai-Anderson evidence-quality review, and the MTSS-B randomized trial converge on that distinction, even though they disagree on how strong any single fidelity-outcome association actually is.

The indicator set above is the working skeleton, already mapped to all four Commission domains. Service-utilization indicators stay flagged as feasibility-dependent rather than promised outright, pending confirmation of a site's CYBHI enrollment and billing infrastructure.

Because the MTSS-B trial's real effects were concentrated in a highest-need subgroup, disaggregation by subgroup and need tier is a priority for the outcomes domain, not an optional add-on.

The Project AWARE and CYBHI Fee Schedule figures are comparable reporting precedent for what the service-utilization domain can look like, not evidence that BHPM itself has produced those results.

This literature informs the indicator set's structure and weighting; it does not by itself finalize which indicators ship.

IMPACTER PathwaySanta Clara County Office of Education

This site publishes no student-level data. All figures are aggregate, de-identified, and reviewed under FERPA, COPPA, and California student privacy law (AB 1584 / SOPIPA).

Prepared in support of SCCOE under BHSOAC Contract No. 25BHSOAC019.