Draft status
Draft methodology for Commission and partner review · v0.9 · weighting and banding finalize through Activity 2.6 structured review with all six LEA teams.
A single number invites a single question: how was it built? A second question follows close behind: could someone who was not in the room reconstruct it from what's published, line by line, without having to take anyone's word for it? The Behavioral Health Performance Index (BHPI) exists to give the Commission, county office leadership, and the six participating LEAs one defensible read on behavioral health performance across a school year, and answering both questions in the open is the actual work of this page. Not a summary for a board slide, but the working method an outside reviewer, an auditor, or a skeptical superintendent could rebuild from scratch, decision by decision, including the decisions that are not yet final. It applies the ten-step construction framework set out jointly by the OECD and the European Commission's Joint Research Centre to the four reporting domains the Commission specified for BHPM: services provided, student demographics and characteristics, service utilization, and outcomes and impact. Every section below names which of those ten steps it is answering, states the choice BHPI's working draft currently makes, and flags plainly where that choice is still provisional pending Activity 2.6.
From four domains to one index
Step one of the OECD/JRC framework is a theoretical framework: naming, before any data collection, what each indicator is supposed to represent and how it ties back to a construct someone actually asked for . For BHPI, that construct is set by the Commission's four domains rather than by whatever the platform happens to already log. Each domain becomes its own sub-index — a Services Sub-Index, a Demographics/Characteristics Sub-Index, a Utilization Sub-Index, and an Outcomes/Impact Sub-Index — built independently through the full ten-step sequence before the four sub-indices are combined into the single BHPI score a district sees on the public dashboard. Building sub-indices first and combining second, rather than blending every indicator into one flat weighted average, keeps the domain structure legible: a reader can see which of the four Commission domains is driving a district's overall standing, not just the standing itself. It also means a data quality problem or a contested weighting choice in one domain stays contained to that domain's sub-index instead of silently diffusing across the whole score in a way no one downstream could trace back to its source.
Within each domain, the candidate indicator set comes from the crosswalk developed under Activity 2.3 — the same indicator families documented on the S4 page (process, outcome, and contextual indicators) — rather than from a separate index-specific indicator hunt. That continuity matters for defensibility: a reviewer checking the index methodology should find the same indicators, defined the same way, as a reviewer checking the Metrics Brief itself. It also means the two documents cannot drift apart over time without the drift being visible; if a future revision adds an indicator to the crosswalk without a matching update here, or vice versa, that mismatch is itself a signal that one of the two documents needs to catch up to the other, not a permissible state to leave standing.
Building the sub-indices: selection, imputation, coherence
Steps two through four of the OECD/JRC sequence are the unglamorous groundwork that most published composites skip past, and BHPI does not get to skip them because the Commission's audience will ask . Variable selection has to be documented well enough that a reader can see why an indicator is in the index and not simply available in the platform; an indicator earns its place in a sub-index because it maps to a construct the Commission's domain language actually names, not because a report happened to already compute it. Missing-data treatment has to be stated in advance — which imputation rule applies when a school has not yet reported a given quarter's data, and when a cell is instead suppressed rather than imputed (the suppression rule is covered separately below, since it is a privacy control as much as a statistical one). A quarter with no data and a quarter with data below the suppression threshold are two different situations that call for two different labels on the dashboard, and conflating them, treating an unreported quarter as if it were a reported-but-suppressed one, would misstate what the platform actually knows about a district in that period.
Before any indicator is normalized or weighted, the indicators inside a domain need a coherence check: do they actually move together, in the way that justifies treating them as facets of one underlying sub-index, or is a domain quietly bundling two unrelated constructs under one label? The OECD/JRC handbook recommends multivariate methods — correlation structure and, where the indicator count supports it, factor analysis — as that check, and flags a domain that fails it as a candidate for splitting rather than forcing it into a single sub-index . This is not a formality that gets run once and filed away. A domain whose indicators pass the coherence check in year one of BHPM can fail it in year two, once six LEAs of different size, staffing model, and program maturity start producing genuinely different indicator profiles from one another; a services-domain sub-index that looked like one coherent construct when every site was still ramping up implementation could split into a training-completion facet and a service-delivery facet once sites diverge in how far along they are. BHPM applies that check per domain before finalizing sub-index composition, and will report the result of that check, and re-run it at each subsequent Activity 2.6-style review, alongside the sub-index definitions once the Activity 2.6 review closes.
Normalization: choosing what "comparable" means
Indicators arrive in incompatible units — a completion rate, a days-to-service count, a tier-movement percentage, a per-pupil dollar figure — and none of them can be added together until they are put on a common scale. The OECD/JRC handbook treats normalization as a substantive methodological choice rather than a mechanical step, because each common option answers a different implicit question . Z-scoring asks "how many standard deviations from the six-LEA mean," which is comparable and interpretable but sensitive to outliers and to the specific comparison group in view; with only six LEAs contributing to that mean, one site's unusual quarter has more leverage over everyone else's z-score than it would in a larger reference population, a distortion worth naming plainly rather than treating as a footnote. Min-max rescaling asks "where does this fall between the observed floor and ceiling," which is bounded and intuitive but means a single unusually strong or weak district shifts everyone else's scale, and with six LEAs that ceiling or floor can be set by a single site's single quarter. Distance-to-target normalization asks a different question entirely: "how far from a stated goal," which anchors the index to policy targets rather than to relative performance but requires the Commission to set defensible targets for every indicator before the index can run, and a target set carelessly, too easy or too aspirational relative to what the six sites can realistically reach in the contract period, would distort the resulting score just as badly as an unstable comparison group would.
For BHPI's working draft, the sub-indices use distance-to-target normalization where the Commission's crosswalk already implies a target (for example, indicators tied to a completion or fidelity threshold) and z-scoring elsewhere, pending the full indicator-by-indicator normalization table the Activity 2.6 review will finalize. The rationale is that a behavioral health performance measure serving six LEAs of different size and baseline capacity should not let a single high- or low-performing site distort everyone else's normalized scale the way min-max rescaling would, and should use fixed policy targets wherever the underlying program (a fidelity checklist, a service-completion standard) already defines one. That rationale is a working position, not a settled one: the specific list of which indicators get a distance-to-target treatment versus a z-score treatment is exactly the kind of indicator-by-indicator call the Activity 2.6 review exists to confirm or revise, and this page will publish the finalized table, not just the general rule, once that review closes.
Weighting: making the value judgments visible
Every weighting scheme embeds a value judgment about which indicators matter more; the only real choice is whether that judgment stays implicit or gets stated and defended . Equal weighting is the most common default in published composites and the easiest to explain, but it quietly asserts that every indicator in a domain deserves identical influence, an assertion nobody actually made on purpose; it is a default chosen by not choosing. Statistical weighting via principal components analysis lets the data itself set weights based on shared variance, which removes the appearance of an arbitrary human choice but replaces it with an equally arbitrary statistical one, and can produce weights that are difficult to explain to the board member reading the dashboard, or weights that shift meaningfully from one reporting period to the next simply because the underlying correlation structure across six LEAs moved, with no change in anyone's actual judgment about what matters. Budget-allocation and expert-elicitation weighting make the underlying value judgment visible by asking a defined group of people to distribute a fixed number of points across indicators and defend the distribution, trading the appearance of statistical neutrality for a record of who decided what and why.
BHPI's working plan uses expert elicitation, conducted through the same structured consultation format as the Activity 2.3 crosswalk sessions, with the weighting exercise run separately by domain across all six LEA teams during Activity 2.6. Running the exercise separately by domain, rather than asking each team to weight all four domains' indicators in one sitting, keeps the elicitation task cognitively manageable and keeps a team's judgment about, say, outcome indicators from bleeding into their judgment about demographic indicators simply because both were asked back to back. That keeps the weighting judgment traceable to named participants and a documented process rather than to an unexplained statistical artifact, and it means the published methodology can say who set the weights and how, which equal weighting and PCA weighting cannot say. Mandatory sensitivity analysis on the resulting weights is not optional under the OECD/JRC framework; a composite whose ranking of districts flips under a modest reweighting has not demonstrated anything robust , and the Activity 2.6 deliverable will include a Monte Carlo perturbation of both the weights and the normalization choice, reporting how much a district's BHPI score and rank move under plausible alternative specifications.
Aggregation and the compensability question
Once indicators are normalized and weighted, they still have to be combined, and how they are combined decides whether strength in one place can cover for weakness in another. Linear (additive) aggregation is fully compensatory: a high score on one indicator can offset a low score on another indicator at whatever exchange rate the weights imply. Geometric aggregation is partially non-compensatory: because it multiplies rather than sums normalized components, a very low score on any one component drags the whole product down disproportionately, so no amount of strength elsewhere fully covers for it , and the OECD/JRC framework is explicit that this is not a stylistic preference but a policy-relevant fork with genuinely different implications for a floor-sensitive construct.
That distinction is not academic for a behavioral health index specifically. A linear, fully compensatory BHPI would let a district average away a real gap, for instance a crisis-response or service-completion shortfall, against strong attendance or participation numbers elsewhere in the same domain, producing a headline score that looks fine while masking exactly the kind of gap the Commission most needs to see. The working draft therefore leans toward geometric aggregation within domains where a floor genuinely matters (service utilization and outcomes/impact), reserving linear aggregation for domains where indicators are closer to genuinely substitutable (demographics/characteristics), on the reasoning that a demographic profile's component figures describe who a district serves rather than how well it is performing, and substitutability among descriptive figures does not carry the same stakes as substitutability among performance figures. The final aggregation rule per domain is one of the specific items the Activity 2.6 structured review will confirm or revise with all six LEA teams before publication, and the worked example further down this page shows concretely why the choice cannot be settled by intuition alone: geometric aggregation is undefined for a negative normalized value, which means adopting it for a domain also commits BHPI to a rescaling step the current draft has not yet fully specified.
From score to color: performance banding
A continuous index score is precise, but precision is not the same thing as actionable. For a board member deciding where to direct attention, a single decimal value is harder to act on than a small number of clearly defined performance bands. BHPI proposes adapting the design California school districts already use for the state's own accountability reporting: the California School Dashboard's five-by-five status-and-change grid, which crosses a percentile-based status level against a year-over-year change level to place each indicator into one of twenty-five color-coded cells, spanning five performance colors from blue at the top of the grid through green, yellow, and orange to red at the bottom, with cut points published in the accompanying technical tables . Reusing that grid logic for the BHPI means every LEA reads behavioral health performance in the same visual grammar they already use for academic and climate indicators on the state Dashboard, rather than learning a second, unrelated color system, and it means a superintendent's cabinet does not need a new legend the first time BHPI appears next to their existing Dashboard tiles.
The status axis places a district's current BHPI (and each sub-index) relative to the six-LEA distribution for the current reporting period. The change axis places the year-over-year movement in that same score. A district can land, for example, in a band that reads "solid current performance, improving" or "below-median current performance, declining," a framing that separates the two questions a superintendent actually has, where do we stand and are we getting better, rather than collapsing both into one number. Cut points for both axes, and the exact banding table, are among the deliverables the Activity 2.6 review will finalize; this page will publish the resulting table once confirmed rather than pre-committing to provisional cut scores here. It is also worth being plain that CDE's own cut points were calibrated against the full population of California's roughly one thousand school districts, and a direct import of those specific numeric thresholds onto a six-LEA reference distribution would not be defensible; what BHPI borrows from the Dashboard is the grid's structure and color grammar, not its specific percentile cut scores, which have to be recalculated against BHPM's own six-LEA distribution.
The change axis raises a measurement question the OECD/JRC handbook does not fully resolve, because computing "improvement" is a harder problem than computing "status": a naive year-over-year subtraction treats every district's prior-year score as equally reliable, which is not true when reporting periods, cohort sizes, or administration timing differ across the six LEAs. The literature on relative-rating systems offers a more rigorous alternative worth flagging for the Activity 2.6 review: the paired-comparison rating framework originally developed for competitive ranking, which estimates relative skill from win-loss outcomes against opponents of known rating , and refined into a dynamic setting that models a competitor's ability as a value that changes over time via a non-linear state-space process, fit with a computationally simple, non-iterative algorithm that explicitly represents how uncertain each rating estimate is rather than treating every observed value as equally precise . Adapting that logic to a change-level calculation would mean weighting a district's year-over-year movement by how confidently each period's score was measured (a period built on a fuller quarter of data would carry more weight in the change calculation than a period built on a partial or newly onboarded quarter), rather than by simple subtraction, a refinement to consider for the change axis specifically, separate from and in addition to the four-domain composite construction this page otherwise describes. This is flagged as literature worth considering, not a construction decision BHPI has made; whether the added statistical complexity of a state-space change model is worth it for a six-LEA reporting population this size is itself a question the Activity 2.6 review should weigh against the simpler, more transparent alternative of a straightforward year-over-year difference reported alongside its own confidence interval.
Composite reliability: omega, not alpha
A composite index should report how reliably its components hang together, the same way any measurement instrument reports internal consistency. The conventional choice, Cronbach's alpha, assumes the components are tau-equivalent, that each indicator measures the underlying construct with the same precision and the same units of "true score." That assumption does not hold for a domain sub-index that mixes a completion count, a fidelity rating, and a service-utilization rate, each with a different relationship to the construct it is meant to represent, and when the tau-equivalence assumption is violated, alpha is a known-biased estimate of the sub-index's actual reliability rather than merely a slightly-off one. Coefficient omega, derived from a factor-analytic rather than tau-equivalent measurement model, remains a defensible reliability estimate precisely under the heterogeneous-component condition where alpha is not . BHPI will report omega for each domain sub-index and for the composite BHPI itself, alongside the correlational coherence check described above (the two are related but distinct checks: coherence asks whether a domain's indicators belong together at all, omega asks how reliably they measure whatever they belong together as, once that grouping is settled), rather than defaulting to alpha because it is the more familiar figure.
It is worth separating this composite-level reliability question from a different reliability figure that will come up elsewhere on this site: the scoring-model agreement statistics reported on the current Model Card, which describe how consistently IMPACTER's ordinal scoring pipeline (Model Suite v6.0) agrees with trained human raters on an individual student response, an average quadratic-weighted kappa of 0.918 across the eight competency attributes, ranging by attribute from 0.893 to 0.952. That figure answers "does the machine learning model score a single response the way a human would." Composite reliability, by contrast, answers "do the indicators inside a domain sub-index cohere as a group." A high omega does not certify individual scoring accuracy, and a high scoring-model kappa does not certify that a domain's indicators belong together; the two are complementary but not substitutes for each other. They should not be conflated, and neither figure should be cited as if it stood in for the other in a board presentation where only one number fits on a slide.
Uncertainty quantification and the mandatory sensitivity check
Every choice documented above, which normalization, which weights, which aggregation rule, is a specific decision among several defensible alternatives, and the OECD/JRC framework treats reporting a single point estimate without testing its sensitivity to those choices as incomplete methodology, not merely good practice . The specific statistical technique behind that requirement, rather than sensitivity analysis as a vague gesture toward "checking robustness," is the pairing of uncertainty analysis with variance-based sensitivity analysis that a JRC team formalized for composite-indicator construction and demonstrated on international benchmarking indices, using a generalization of Sobol's method to apportion how much of a composite's ranking uncertainty traces to one specific construction choice, such as the weighting scheme, versus another, such as the missing-data imputation rule . That distinction matters for what the Activity 2.6 deliverable should actually report: not one undifferentiated robustness number, but an apportioned answer to the question "if a district's BHPI score or rank is unstable, which specific methodological choice is driving that instability," so a district or a reviewer knows which construction decision to scrutinize rather than being told only that the score is somewhat uncertain overall.
The Activity 2.6 deliverable accompanying this page will include a documented sensitivity analysis along those lines: a Monte Carlo procedure that perturbs the normalization method, the weight set, and the aggregation rule across a defined range of plausible alternatives, and reports how much each district's BHPI score and relative rank move as a result, decomposed by which perturbed input is responsible for how much of that movement. A district whose ranking is stable across that perturbation range has a defensible score; a district whose ranking flips under small, reasonable changes to the construction method has a score that carries real caveats rather than false precision, and the apportioned sensitivity result should say specifically whether that instability traces to the weight set, the normalization choice, or the aggregation rule, since the appropriate fix differs depending on which one it is. Until that sensitivity analysis is complete, every BHPI figure referenced elsewhere on this site should be treated as provisional in exactly the sense the draft-status banner above states.
Protecting students: suppression and the no-student-level-data rule
The public-facing BHPM dashboard reports aggregate, district- and school-level figures only; it does not and will not publish student-level data, and every domain indicator is defined at a level of aggregation designed to prevent indirect re-identification of an individual student through a small reporting cell. Federal guidance for statewide longitudinal data systems documents that most states set a minimum reporting-group size between five and thirty students for exactly this reason, with ten the most common threshold, and recommends that states adopt and publicly state a specific rule rather than leaving suppression to case-by-case judgment . That same guidance also flags that threshold suppression alone can still allow disclosure by subtraction across related cells unless paired with complementary suppression: if a school's total enrollment and every demographic subgroup but one are published, a reader can back out the suppressed subgroup's value by subtracting the published ones from the total, even though that one cell never appeared on the page in numeric form. Complementary suppression closes that gap by suppressing at least one additional cell in the same related set whenever a single cell falls below threshold, so no combination of published figures lets a reader reconstruct the number that was withheld, which is a design detail the BHPI dashboard's suppression logic needs to account for and not just a single minimum-n check.
The working default for BHPI is a minimum reporting-group size of ten students per cell, following the modal practice the federal guidance documents, applied together with complementary suppression across related demographic and outcome breakdowns. That default is not yet final: it requires sign-off from SCCOE data governance before it can be treated as the production rule, and this page will update once that approval is recorded. Until then, any cell falling below the working threshold displays as suppressed rather than as an estimated or rounded figure, and no component of the BHPI construction pipeline retains or exposes student-level records outside the platform's existing FERPA- and COPPA-governed data environment. A suppressed cell and a not-yet-reported cell should read differently to anyone looking at the dashboard, since one reflects a privacy protection working as designed and the other reflects a data-collection gap that the platform should eventually close; collapsing the two into a single blank or dash would erase a distinction the reader actually needs.
Keeping the method honest: versioning and public documentation
The final step of the OECD/JRC sequence is presentation built for the audience that will actually use it, and a version-controlled, publicly accessible methodology page is how BHPI keeps that commitment durable rather than one-time . This page carries an explicit version marker (v0.9, pending the Activity 2.6 review) precisely so that a change to normalization, weighting, aggregation, or banding produces a new version number and a visible changelog rather than a silent edit to a score's definition. The same discipline applies to the underlying scoring pipeline referenced above: model versions evolve over time and are versioned and audited for traceability and reproducible scoring, and any figure quoted from a specific model version on this site is tied to that version rather than presented as a permanent constant. A reader returning to this page in a year should be able to tell exactly what changed, when, and under whose review, the same standard the California Department of Education holds itself to by archiving each year's Dashboard technical guide separately rather than overwriting the prior version . Concretely, that means a v1.0 release of this page at the close of Activity 2.6 will retain v0.9 as an archived prior version rather than replacing it outright, so that any BHPI figure already reported to the Commission under v0.9 stays traceable to the exact methodology that produced it, even after the production methodology has moved on.
Activity 2.4 — the real BHPI, as already reviewed with the Commission
Everything above works through the general OECD/JRC composite-construction sequence in the abstract — the families of options BHPI could draw on for normalization, weighting, and aggregation. Activity 2.4's actual job is narrower and more concrete: define the specific composite BHPM is reporting. That composite already exists in working form, was presented to Pilar and SCCOE alongside the "v2" dashboard and educator-portal prototypes on 2026-08-07, and is reused here exactly rather than re-derived from the general framework above. Unlike the aggregation discussion two sections up, BHPI's realized construction is a deliberate, fully linear, compensatory weighted sum, not a geometric or otherwise non-compensatory aggregation — a specific choice for this specific composite, not a contradiction of the general aggregation-family literature reviewed above, which still applies to how BHPM evaluates aggregation choices domain by domain elsewhere in the framework.
BHPI is Domain 4's flagship outcome metric — the number Outcomes & Impact reports — not a fifth index competing with the Commission's four-domain taxonomy. Domains 1 through 3 (services provided, student demographics and characteristics, service utilization) remain the input and context reporting categories built out under Activity 2.3; Domain 4 is where this composite gets reported.
Activity 2.4 · reported inside Domain 4 — Outcomes & Impact
The Behavioral Health Performance Index
BHPI is Domain 4's flagship outcome metric, not a fifth index competing with the Commission's four reporting domains. Domains 1–3 (services provided, students served, service utilization) are the input and context reporting categories; this is where the outcome gets reported.
BHPI = 0.40(Voice) + 0.20(Behavior) + 0.16(Engagement) + 0.12(Academic) + 0.12(Service Receipt)
Five components, by weight
- Voice · 40% — Weighted heaviest because authentic student-voice evidence — what a student actually says, in their own words — is the differentiating signal no other state behavioral-health instrument collects.
- Behavior · 20% — Second-heaviest because suspension, expulsion, and referral data is the highest-stakes signal already sitting in a district's own student information system.
- Engagement · 16% — Attendance is a strong, well-established wellbeing signal — present, but weighted below Voice and Behavior rather than treated as equally decisive.
- Academic · 12% — Lightest by design: academic performance is context for behavioral health, not a driver of it, so it contributes without dominating the composite.
- Service Receipt · 12% — Closes the loop between what was delivered and what it produced. Optional and renormalizing — a site without service-receipt data yet is not penalized for it.
Missing-data renormalization — If a site is missing data for one of the five dimensions, that dimension is excluded and the remaining weights scale up proportionally so they still sum to 1.0 — a site is never penalized for a data gap, it is scored on what it has actually reported.
Performance bands
The layer beneath Voice
Student-level Wellness Index (.250–1.000)
BHPI's Voice dimension is built from a separate, student-level Wellness Index that routes each student to a support tier (Tier 1 universal · Tier 2 targeted · Tier 3 intensive), scored across six BHPM-specific competencies:
These six competencies were developed specifically for BHPM, co-developed with the San Diego County Office of Education — a distinct, BHPM-specific set, not the general IMPACTER 8 competency model used elsewhere on the platform.
4-of-6 rule — Every student needs at least 4 of the 6 competencies reported to receive a Wellness Index score. If one or two are unreported, the index renormalizes across the remaining reported competencies — the same proportional-scaling logic BHPI itself uses at the site level.
The bands above translate a continuous BHPI score into the same kind of actionable read the general banding discussion above calls for, but they are BHPM's own four-level scale — Leading, Establishing, Developing, Emerging — set against fixed cut points rather than the California Dashboard's percentile-based five-by-five grid described earlier on this page. Both are strengths-based, plain-language framings; which one the production dashboard ships with, or whether the two get reconciled into one grammar, is itself one of the specific questions Activity 2.6's structured review with all six LEA teams will settle, consistent with this page's draft, v0.9 status. Nothing about publishing the real formula here changes that status: weighting, banding, and the domains' relationship to each other still finalize through Activity 2.6, and this composite is presented as the working draft under Commission and partner review, not as a closed specification.
The per-student rubric beneath Voice: six competencies, four levels
The "layer beneath Voice" panel above names the six BHPM-specific Wellness Index competencies — Self-Insight, Emotional Resilience, Relational Awareness, Conflict Resolution, Effective Help-Seeking, Reflective Growth — as a badge list. What follows is the actual rubric behind those six names: the four-level Foundational / Developing / Competent / Advanced scoring guide a rater (human or machine-learning-scored) applies to a single student's single spoken reflection on a single competency.
This is a different scale from the BHPI performance bands shown above, even though both happen to have four levels, and the two should not be conflated. The rubric below scores one student, on one competency, from one piece of authentic student speech — it lives entirely inside the Voice component, well below the level BHPI itself reports at. The Leading / Establishing / Developing / Emerging bands above score an LEA's (or the composite's) aggregate BHPI value on a 0–1 axis, at the site level, once every student's scores across every dimension have already rolled all the way up. Four levels each is a coincidence of scale design, not a shared instrument: a student rated Competent on Self-Insight is not "the same" as a district banded Establishing on BHPI, and this page's two four-level scales should be read as two separate rubrics measuring two separate things, not one rubric reported twice.
Per-student, per-competency scoring rubric
The Wellness Index rubric — six competencies, four levels
This scores one student, on one competency, from a single spoken reflection. It is not the BHPI performance bands shown above — those score an LEA’s aggregate BHPI score on a separate 0–1 axis, at the site level, not the student level. Four levels each, coincidentally; two different scales, measuring two different things, at two different grains.
- FoundationalStruggles to notice or describe their actions, motives, or values; avoids self-reflection or misreads how actions connect to personal identity; little sense of learning from experience.
- DevelopingSometimes recognizes the gap between actions and values but is vague or needs prompting; limited detail about what was realized or how self-insight could lead to growth.
- CompetentClearly describes a moment when actions didn't fit their values; explains what was realized, how self-understanding changed, and gives examples of learning from experience.
- AdvancedOffers deep, thoughtful insight into motives, identity, and growth; connects actions, values, and learning; reflects on how increased self-insight guides future choices and helps others.
- FoundationalHas difficulty describing how they manage or recover from difficult emotions or stressful situations; may focus on the problem without naming feelings or using any coping strategies.
- DevelopingSometimes describes handling tough moments or feelings but gives only general or inconsistent details; may mention strategies or supports but lacks depth or clarity.
- CompetentClearly explains how they responded to and managed a challenging emotion or situation; names specific coping strategies, supports, or resources used to recover or move forward.
- AdvancedShares rich, specific examples of managing adversity and emotions; explains how resilience developed over time; reflects on using and adapting strategies to recover, grow, and help others facing challenges.
- FoundationalRarely notices or describes changes in relationships or their own role; may focus on disconnection, conflict, or confusion; shows little empathy or reflection on impact or learning.
- DevelopingSometimes recognizes changes or shifts in relationships and can mention feelings or outcomes, but responses lack detail or personal insight and often require prompting.
- CompetentClearly describes a relationship change, their role in it, and how it felt; explains what was learned about themselves or others, and how they adapted or grew from the experience.
- AdvancedOffers nuanced, thoughtful insight into relationship dynamics and shifts; reflects on personal and others' perspectives; discusses how relational awareness shapes trust, belonging, and future relationships.
- FoundationalAvoids or minimizes discussion of conflict or making amends; focuses only on the problem or harm; does not mention repair or solutions; little reflection on role or outcomes.
- DevelopingSometimes discusses trying to resolve conflict or repair relationships, but with limited details; needs prompting to explain own actions, impact, or learning from the experience.
- CompetentDescribes actions taken to resolve conflict, repair harm, or make things right; explains steps clearly, what was learned, and reflects on personal growth in conflict situations.
- AdvancedExplains conflict resolution deeply; describes leading or modeling repair for others, using restorative practices, and promoting positive change in relationships and the community.
- FoundationalReluctant or unable to talk about asking for help; may express distrust, shame, or discomfort using supports; rarely explains benefit or value of seeking help.
- DevelopingSometimes describes seeking help, but inconsistently or with uncertainty; benefits of support may be unclear or responses lack detail; may need prompting to discuss using resources.
- CompetentClearly explains when and how they sought help or support; describes how seeking help made a positive difference for themselves or others; reflects on value of support systems.
- AdvancedRegularly seeks help confidently; explains how support benefits self and community; encourages others to seek help and describes effective strategies for using resources in spoken reflection.
- FoundationalStruggles to notice or describe personal change; may feel stuck or unable to reflect on growth, learning, or how experiences have shaped their identity or choices.
- DevelopingSometimes mentions changes in themselves or growth, but descriptions are vague, general, or require prompting to connect change to specific actions or learning.
- CompetentClearly explains a meaningful way they have changed; describes how growth happened, what contributed to it, and what was learned or how perspective shifted through reflection.
- AdvancedGives rich, detailed reflection on their growth over time; connects changes to actions, learning, and experience; explains how reflective growth informs ongoing improvement and helps others believe in their own growth.
Source: "SDCOE BH Rubric - Final," developed with the San Diego County Office of Education for BHPM specifically — the same six-competency set named in the Wellness Index panel above, not the general IMPACTER 8 competency model used elsewhere on the platform.
Scale note — This four-point scale adapts Amy Berry's Engagement Continuum (Disrupting/Avoiding, Withdrawing, Participating/Investing, Driving), combined into four levels reflecting disengagement-to-engagement.
Bottom Line
BHPI never stands alone as a single fixed number. The version marker, not the score by itself, is the citable unit: any reported figure carries its version (v0.9, pending Activity 2.6), the real formula and per-dimension rationale documented above, the missing-data renormalization rule, and the four performance bands together, so a reader can see exactly what produced a district's score, and why Voice carries the heaviest weight in that sum.
Performance bands read in a plain-language, strengths-based grammar: either BHPI's own four-band scale (Leading, Establishing, Developing, Emerging) or the California Dashboard's status-and-change grid described above, so LEA teams read behavioral health performance in a register they already recognize. Which of the two the production dashboard leads with, or whether the two get reconciled into one, is Activity 2.6's call. Any breakdown by demographic or small subgroup pairs with the minimum-cell-size suppression rule described above, once that rule has SCCOE data-governance sign-off.
Composite reliability (omega) and scoring-model agreement (QWK) stay visibly separate figures. They answer different questions, and conflating them overstates what either one actually certifies.
The real formula, the bands, and the renormalization rule above are settled working draft. The sensitivity-analysis and final-banding questions are still open, and Activity 2.6's structured review with all six LEA teams is where they get resolved.
