What does a single Behavioral Health Performance Index score actually tell a board member, and what does it hide? Picture two LEAs landing on the same composite number at the end of a reporting cycle. One runs high participation, strong process fidelity, and fast crisis response, but its outcome measures are barely moving. The other has thinner service utilization and a slower crisis-response track record, but real outcome gains wherever it is measuring at all. Averaged and compensated against each other inside a single index, those two different behavioral health systems can land on the same number, and a board reading only the top-line score has no way to tell them apart. That is not a flaw unique to behavioral health measurement; it is the problem composite-indicator methodology exists to solve, and why the Commission's four reporting domains — services provided, student demographics and characteristics, service utilization, and outcomes and impact — cannot simply be averaged into a single number without an explicit, published construction method behind the arithmetic.
IMPACTER's approach to the Behavioral Health Performance Index (BHPI) follows the ten-step composite-indicator framework set out by the OECD and the European Commission's Joint Research Centre . The sequence begins well before any arithmetic runs: a theoretical framework tying each indicator to a construct the Commission actually asked for, deliberate variable selection rather than "whatever the platform already logs," and a documented, defensible treatment of missing data, so a partial-year deployment doesn't simply drop out of a domain sub-index and understate the sites furthest along the readiness curve. The framework also calls for multivariate checks — correlation structure and, where the indicator count supports it, a principal-components read — confirming indicators grouped into a domain actually move together; a domain that doesn't covary is measuring more than one construct under a single name. From there the sequence states a normalization method, a weighting method, and an aggregation method; requires sensitivity analysis on every prior choice; and closes with a presentation layer built for the audience that will actually read it — superintendents and county office leaders, not psychometricians. Each of the four Commission domains becomes a sub-index built through that sequence; the four sub-indices then aggregate into the BHPI itself. Nothing about the sequence is exotic. What is unusual, at least in K-12 behavioral health reporting, is publishing it in full rather than handing districts a single color and a shrug.
Three of those ten steps carry the most consequence and the most disagreement in the literature, which is why the full Index Methodology page treats each at length rather than summarizing it away here. Normalization has no neutral option. Z-scores center every indicator on the sample's own mean and spread, min-max rescaling stretches each indicator across a fixed range anchored to the observed minimum and maximum, and distance-to-target methods measure a site's distance from a policy-set benchmark instead of how its peers happened to perform that cycle — three different implicit questions about what "average" means, and with a program built on six LEAs and eighteen schools, peer-relative normalization is also the most exposed to noise from a small comparison set. The choice changes which schools look strong before a single weight is assigned.
Weighting is the step most exposed to hidden values, and the framework catalogs several families rather than endorsing one. Equal weighting looks neutral but quietly asserts every indicator in a domain matters the same amount — itself a judgment, just an unexamined one. Data-driven weighting through principal components assigns weight by how much variance an indicator shares with its domain, surfacing structure equal weighting would flatten, but it can swing sharply with a handful of added or dropped sites in a program this size. Budget-allocation or expert-elicitation weighting — the family associated with structured stakeholder input, and the natural fit for the crosswalk's Activity 2.6 LEA-team consultations — asks reviewers to directly allocate a fixed number of importance points across indicators, making the value judgment visible and arguable instead of laundering it through statistics. Visible and arguable is preferable even when less tidy: a weighting scheme nobody can see is one nobody can contest.
This is where the sequence stops being a formality and starts deciding things on the Commission's behalf. Aggregation decides whether a strength in one domain can offset a weakness in another. Linear, compensatory aggregation — a straight weighted average — lets exactly the scenario that opened this section happen: strong process fidelity and fast crisis response mathematically canceling out flat outcomes, or the reverse, with the resulting number giving no indication which trade occurred. Geometric or otherwise non-compensatory aggregation requires every domain to clear its own floor before it can contribute fully to the composite — real stakes for a behavioral health index specifically, where a district should not be able to average away a crisis-response gap with strong attendance numbers, and a non-compensatory rule is the mechanism that prevents it. The Index Methodology page works through both families with the sensitivity analysis requires — Monte Carlo-style resampling across plausible combinations of normalization, weighting, and aggregation choices, testing whether a site's ranking holds or swings depending on which defensible specification was picked — and lays out BHPI's actual, Commission-reviewed construction under Activity 2.4: a linear weighted sum of five components (Voice 0.40, Behavior 0.20, Engagement 0.16, Academic 0.12, Service Receipt 0.12), reported as Domain 4's flagship outcome metric rather than as a hypothetical.
Two further commitments round out the construction and are non-negotiable regardless of which weighting or aggregation choice the Activity 2.6 structured review lands on. First, sub-index and composite reliability will be reported using coefficient omega rather than Cronbach's alpha. Alpha assumes tau-equivalence — that every indicator feeding a scale carries the same relationship to the underlying construct — and each BHPI domain's indicator set violates that assumption by design, mixing process counts, participation rates, and outcome measures with different units and different loadings on whatever the domain is measuring. Omega is derived instead from a factor-analytic model that lets each indicator's contribution to reliability track its actual loading rather than assuming a uniform one, and it stays defensible under exactly the heterogeneous-indicator condition where alpha is known to mislead ; both a domain sub-index and the composite it feeds can be estimated this way, not only the top level.
Second, the public-facing dashboard applies small-cell suppression by default, following the disclosure-avoidance practice federal longitudinal-data guidance recommends for any subgroup report where re-identification risk rises as group size falls . That guidance distinguishes primary suppression — masking any cell below the minimum group size on its own — from complementary suppression, which masks additional cells in the same table that would otherwise let a reader back-calculate the suppressed value by subtraction from a published total. It documents states setting that threshold anywhere from five to thirty students, ten the most common, and recommends one stated rule rather than deciding case by case; the working default here is ten, pending SCCOE data-governance sign-off. That risk isn't abstract in an eighteen-school program: the smallest cells usually sit at the intersection of a small school and a small demographic subgroup, exactly what complementary suppression exists to catch.
For performance banding — translating a continuous index score into something a board can act on — the Index Methodology page proposes adapting the in-state precedent every California LEA already knows: the School Dashboard's five-by-five status-and-change grid . Status reflects a percentile-ranked performance level in the current year; Change reflects year-over-year movement; crossing the two five-level scales produces twenty-five cells mapped onto a five-color scale. That design already generalizes across dissimilar indicator types on the Dashboard — chronic absenteeism, suspension rate, graduation rate, and academic indicators all run through the same logic — evidence the grid was not built for one narrow use case, and its technical guide is archived as a new, separately versioned document every year rather than silently revised in place, a discipline the BHPI's own "draft methodology, v0.9" posture is modeled on. Reusing the grid means districts read the BHPI the way they already read academic and climate indicators, rather than learning a second accountability grammar from scratch — and it lets the two hypothetical districts from the opening actually separate: one lands in a cell showing high status but flat change, the other lower status but strong positive movement, giving a board two distinguishable pictures instead of one blurred average.
Bottom Line
The BHPI is a constructed, published methodology, not a black-box score. Activity 2.4 fixes that construction as an actual composite rather than a framework left abstract: BHPI = 0.40(Voice) + 0.20(Behavior) + 0.16(Engagement) + 0.12(Academic) + 0.12(Service Receipt), reported inside Domain 4 (Outcomes & Impact) as that domain's flagship metric, not a fifth index competing with the Commission's four-domain taxonomy. Voice carries the heaviest weight because authentic student-voice evidence is the differentiating signal no other state instrument collects; Behavior is second because suspension, expulsion, and referral data is the highest-stakes signal already sitting in a district's SIS; Engagement (attendance) and Academic follow, with Academic weighted lightest by design since it is context for behavioral health rather than a driver of it; Service Receipt closes the loop and is optional and renormalizing. If a dimension is missing, the remaining weights scale up proportionally rather than penalizing a site for a data gap — the same 4-of-6 renormalization logic the student-level Wellness Index (.250–1.000, six BHPM-specific competencies: Self-Insight, Emotional Resilience, Relational Awareness, Conflict Resolution, Effective Help-Seeking, Reflective Growth) applies underneath BHPI's Voice dimension. Bands are Leading (≥.75), Establishing (≥.60), Developing (≥.45), and Emerging (<.45). The Index Methodology page carries the full diagram, the omega commitment, and the small-cell suppression default. Every dashboard number stays provisional until Activity 2.6's structured review with all six LEA teams finalizes it: the formula above is a working draft under Commission and partner review, not a closed specification, stated as such rather than left to silence.
