← D2 overviewSection 5 of 7

Implementation Science & Process Evaluation

Picture a district that does everything a rollout checklist would tell it to do. It buys a validated screener, schedules a single August training day for every counselor across twelve schools, and sends staff back to their buildings with binders and good intentions. Eight months later, three schools are using the tool exactly as designed, five have quietly shortened it into something nobody wrote down, and four have gone back to the paper form they used the year before. Nothing about the screener changed. What changed is that training landed once and was never followed by coaching, and nobody was watching the data closely enough to notice the drift until an outside monitoring visit surfaced it months later. That pattern, strong content, weak delivery system, is close to the modal failure mode implementation science exists to explain, and the National Implementation Research Network's foundational synthesis names the underlying question directly: what actually separates programs that take hold in real service settings from those that stay confined to a manual . This section lays out the frameworks BHPM draws on to make deployment itself measurable, not just the program being deployed, and argues that the project's own activity sequence, readiness assessment, configured deployment, technical assistance, collaborative learning, is a direct application of that literature.

Implementation is a staged process, not an event

The Fixsen synthesis proposes that implementation unfolds across four identifiable stages rather than arriving in one step once a program is adopted. Exploration is the stage where a site assesses whether an innovation actually fits an identified need and whether the resources exist to support it, before anyone commits to it. Installation follows: the structural supports, staffing, funding lines, data systems, policy alignment, get put in place before the first real use, on the theory that a site should not be expected to build its own plane while flying it. Initial implementation is the awkward middle stage, where a new practice meets an old system and needs intensive support to survive first contact with real caseloads and real schedules. Full implementation is reached only when the new practice becomes integrated into a site's normal, ongoing routines, not a special project running alongside business as usual but simply how the work gets done. Programs can and do stall at any of the first three stages indefinitely; the synthesis treats that as the expected risk to manage against, not an edge case.

Within that staged model, the synthesis groups the specific practices that predict whether a site actually reaches full implementation into three interacting categories of "implementation drivers." Competency drivers, staff selection, training, coaching, and performance evaluation, build the capacity of the people doing the work. Organization drivers, systems intervention, facilitative administration, and a decision-support data system, build the capacity of the organization around them. Leadership drivers, technical and adaptive, hold the other two together when conditions get complicated. Two findings matter directly for a multi-district behavioral health system. First, training without ongoing coaching and a data-feedback loop rarely produces full implementation on its own; a single August in-service is a competency-driver input with no organization-driver support behind it, close to a textbook description of why the district above stalled. Second, a decision-support data system is a core organization driver in its own right, not an add-on layered onto the "real" work of service delivery. A performance management platform generating usage and fidelity data as a byproduct of normal operation is, in this framework's own terms, implementation infrastructure rather than a separate reporting obligation.

A complementary, more recent framework fills a gap the 2005 synthesis leaves underspecified: why the identical program, delivered with identical training, succeeds at one site and stalls at another. The Consolidated Framework for Implementation Research organizes the contextual determinants of implementation success into five domains, intervention characteristics, outer setting, inner setting, individual characteristics, and implementation process, consolidating constructs from nineteen earlier implementation theories into one shared taxonomy rather than a new causal model . Its authors describe it as operationalizing the multi-level influence the Fixsen synthesis already gestures at, giving that framework's driver categories more granular constructs to sit under. For a project spanning six LEAs with different staffing, prior behavioral-health infrastructure, and district leadership, the inner-setting domain, implementation climate, leadership engagement, available resources, is close to what Deliverable 1's readiness assessment is built to capture before installation. Naming it explicitly is a useful check on that instrument: it should ask about organizational context, not only technical capacity, since CFIR treats the two as separable predictors that do not move together.

What counts as evidence that implementation succeeded

A parallel line of work addresses a definitional problem that follows directly from the staged model: if implementation is its own process, what does it mean to say that process succeeded, independent of whether the practice ultimately changed outcomes for students? Proctor and colleagues answer by proposing eight conceptually distinct implementation outcomes, defined precisely enough to be measured rather than gestured at. Acceptability is the perception among stakeholders that a practice is agreeable or satisfactory. Adoption is the intention or decision to try a practice in the first place. Appropriateness is its perceived fit for a given setting or population. Feasibility is whether it can be carried out within an agency's real constraints. Fidelity is the degree to which it was delivered as the developer intended. Implementation cost is the resource impact of the effort itself, separate from the underlying service's cost. Penetration is how deeply a practice has spread into a service setting and its subsystems, not just whether it exists somewhere in the building. Sustainability is whether it survives past its initial funding or novelty period and becomes part of ongoing operations .

Those eight sit inside an explicit three-tier logic, and the tiers matter as much as the definitions. Implementation outcomes are necessary but not sufficient evidence of service-system outcomes, which the paper anchors to the Institute of Medicine's quality dimensions, care that is safe, effective, patient-centered, timely, efficient, and equitable. Service-system outcomes are in turn necessary but not sufficient evidence of the ultimate client-level outcomes an initiative is meant to produce. A practice delivered with high fidelity, at real scale, sustained for years, can still fail to improve student well-being if the underlying practice was weak or mismatched to the population; the logic runs the other way too, a genuinely effective practice delivered with low fidelity in a handful of pilot classrooms will not show population-level effects no matter how good it is on paper. Neither tier substitutes for the other, and Proctor and colleagues were candid that the field's measurement toolkit for most of these eight constructs was still underdeveloped, existing instruments largely home-grown, and that measures built for efficacy trials would likely prove too cumbersome for real-world implementation studies, a gap that is part of why a platform-generated usage and fidelity signal, rather than a bespoke survey instrument, is a defensible design choice on its own terms.

That taxonomy gives BHPM a principled way to separate two kinds of indicator easy to conflate in a single reporting table. Fidelity, penetration, and adoption describe whether the system is being used as designed, by the people it was designed for, at meaningful scale, questions the platform can answer directly from its own usage records. Whether students are better off is a downstream, service-system and outcomes-level question the project's other components, the indicator set and the composite index, are built to address. Keeping the two tiers distinct, rather than collapsing "the platform was used" into "the platform worked," is a direct application of the Proctor framework, and it is the reasoning this section's evidence points to for treating process and outcome indicators as structurally separate families rather than variations on one scale.

Continuous improvement over episodic evaluation

A third strand concerns not just what gets measured but how measurement feeds back into practice while a program is still running. Deming's formulation of the Plan-Do-Study-Act cycle, developed for industrial quality management, offered a general method for testing a change on a small scale, studying what happened, and adjusting before committing to system-wide rollout . A second contribution from that same body of work matters as much as PDSA itself: the distinction between special-cause and common-cause variation, whether a swing in a metric reflects something specific and addressable, or ordinary noise the underlying process always produces. Most fidelity- and process-monitoring plans lean on that distinction whether or not they name it, because treating every quarter-to-quarter wobble in a usage indicator as a signal worth acting on is how a monitoring system trains its own users to stop trusting it.

Bryk, Gomez, Grunow, and LeMahieu import that logic directly into K-12 improvement work, arguing that education reform tends to fail less for lack of good ideas than for lack of disciplined methods to adapt those ideas to real, variable local conditions . Their book organizes the argument around six core principles: make the work problem-specific and user-centered rather than solving in the abstract; treat variation in performance across sites as the thing to understand, not an inconvenience to average away; see the whole system that produces current outcomes before changing a piece of it; embed measurement of outcomes and processes so the network can tell a genuine improvement from a coincidence; anchor practice change in disciplined PDSA inquiry rather than a single big-bang rollout; and accelerate improvement through networked communities, on the premise that a network learns faster collectively than any one site learns alone. Their proposed structure, the networked improvement community, organizes practitioners and researchers across sites around one specific, shared problem of practice with common measures, so that what one site learns through a PDSA cycle can inform the next site's cycle rather than staying local, a meaningfully different design than a generic professional learning community convened around a topic; the shared measures are what let the network compare notes and accumulate evidence across cycles rather than exchange anecdotes.

This is the direct literature basis for the Collaborative Learning Group structure specified under Activities 2.4 and 2.5 of the BHPM workplan. The quarterly learning-group cycle across the six participating LEAs reads naturally against the networked-improvement-community model rather than a status-meeting model: the composite index and indicator set are positioned to supply the common measures a network needs to compare notes across sites, and each quarter's summaries can function as the "study" phase closing the loop into the next quarter's "plan." The quarterly cycle earns that comparison only if it runs with the disciplined PDSA structure and shared measures Bryk et al. specify, not as a status-meeting series in new packaging; that is the standard the cycle should be checked against.

What the evidence says works, specifically, in school mental-health implementation

A 2023 systematic review in Prevention Science narrows the general implementation-science literature to the specific setting BHPM operates in: universal, school-based mental-health prevention programming. Reviewing 21 studies published between January 2000 and October 2021, only five of them randomized controlled trials, the rest split across quasi-experimental, single-arm pre-post, cross-sectional, and qualitative or mixed-methods designs, the authors identified 22 distinct implementation strategies with some evidence of improving fidelity or adoption . The strongest evidence concentrated in three places, and the review is specific about which outcome each one moved. Audit-and-feedback strategies, school or project staff observing program delivery and giving teachers direct, positive-reinforcement-oriented feedback, showed the most consistent effect on fidelity, positive results in four of five studies that tested it. Engaging school principals as visible program champions, local opinion leaders who promote the program and steer resources toward it, showed a consistent positive association with adoption specifically, in every qualitative study that examined it. Strategies building buy-in among front-line teachers touched both outcomes, fidelity in one study and adoption in two others.

The same review is candid about limitations worth naming rather than glossing over. Strategies aimed specifically at adoption, as distinct from fidelity, have been evaluated almost entirely through small qualitative or case-study designs, and the authors call explicitly for larger trials before treating those findings as settled. Fidelity, by contrast, was measured in sixteen of the twenty-one studies, adoption in only five, an imbalance any indicator set built from this review should carry forward honestly rather than implying adoption is as well-supported as fidelity is. The authors also flag that two-thirds of the reviewed studies had no explicit theoretical grounding for the strategy being tested, that most evidence comes from English-language, U.S.-based studies, and that staff turnover in schools introduces attrition problems that make longitudinal fidelity tracking harder to interpret over multi-year periods, exactly the kind of period a six-LEA initiative like BHPM runs across.

The practical implication for BHPM is narrower and more useful than a general endorsement of "monitoring." The strand of implementation strategy with the best evidence in this exact setting is continuous monitoring joined to feedback, not a periodic external audit, and that evidence is strongest specifically on fidelity rather than on adoption or the other six Proctor constructs. A system that surfaces usage and fidelity signals on a rolling basis, so a school or district can respond to its own implementation pattern between formal check-ins, is aligned with where this literature's evidence is strongest. A system that instead relies on an annual retrospective survey to reconstruct the prior year is aligned with a design this evidence base does not particularly support, and it compounds a second, well-documented problem: self-report recall about program fidelity degrades over exactly the time horizon an annual instrument requires, the same demand-characteristic risk the Baffsky review flags in its own underlying studies.

How BHPM's own sequence maps onto these frameworks

Layered on top of each other, these strands do not describe four separate literatures so much as one coherent pathway, and BHPM's own activity sequence follows it closely enough to state the mapping plainly. Deliverable 1's site readiness assessment corresponds to the exploration and installation stages in the Fixsen model, and, read through CFIR's inner-setting domain, to a check on organizational context rather than technical capacity alone, establishing before deployment whether a site has the preconditions the driver framework treats as necessary. Deliverable 2's configured deployment, the indicators, index, and scoring infrastructure this Brief documents, is the decision-support data system named directly as an organization driver. Deliverable 3's training and technical-assistance plan supplies the coaching component the Fixsen synthesis found training alone cannot substitute for. Deliverables 4 and 5's Collaborative Learning Group structure is the networked-improvement-community mechanism Bryk et al. specify, running on a PDSA cadence with the composite index as a candidate common measure. And underneath all four sits the platform's continuous analytics layer, which is what lets fidelity and adoption, in Proctor's terms, be observed as they happen rather than reconstructed once a year later, consistent with where Baffsky et al. locate the strongest evidence, specifically for fidelity, in this exact setting.

That correspondence is easier to see laid out stage by stage than carried in a single paragraph, so the diagram below puts the Fixsen model and BHPM's own sequence side by side.

How BHPM’s own sequence maps onto these frameworks

From exploration to full implementation

The stages below are the Fixsen synthesis’s own four-stage model ; the cards beneath each one are the specific BHPM deliverable this section argues occupies it.

The implementation-science model

Exploration
Installation
Initial Implementation
Full Implementation

Assess whether the innovation fits an identified need and whether the resources exist to support it, before anyone commits to it.

Put structural supports in place before first real use: staffing, funding lines, data systems, policy alignment.

The awkward middle stage, where a new practice meets an old system and needs intensive support to survive first contact.

The new practice is integrated into a site's normal, ongoing routines, not a special project running alongside business as usual.

BHPM’s activity sequence

Deliverable 1

Site readiness assessment

Exploration + Installation

Checked against CFIR's inner-setting domain as well as the driver framework, so the instrument asks about organizational context, not only technical capacity, before deployment.

Deliverable 2

Configured deployment

Installation

The indicators, index, and scoring infrastructure this Brief documents: the decision-support data system Fixsen names directly as its own organization driver, not an add-on to service delivery.

Deliverable 3

Training & technical assistance

Initial Implementation

Supplies the ongoing coaching component the synthesis found training alone cannot substitute for.

Deliverables 4-5

Collaborative Learning Groups

Full Implementation

The networked-improvement-community mechanism Bryk et al. specify: a quarterly PDSA cadence with the composite index as a candidate common measure.

Underneath all four

The platform’s continuous analytics layer lets fidelity and adoption be observed as they happen rather than reconstructed once a year later, consistent with where locates the strongest evidence, specifically on fidelity, for continuous monitoring joined to feedback in this exact setting.

None of this substitutes implementation success for the separate, harder question of whether the initiative improves outcomes for students. Proctor's own three-tier logic is explicit that implementation outcomes are a necessary floor, not a proxy for the ceiling, and the Baffsky review's candor about its own evidence gaps, small samples, thin theoretical grounding, adoption limited to qualitative designs, is a reminder that even that floor is measured with real uncertainty. What this literature does establish is that treating implementation as a measured, staged, continuously monitored process, rather than a launch event followed by an annual check-in, is itself an evidence-based design choice, distinct from any claim about student outcomes, and it is the choice BHPM's activity structure was built around.

Key Takeaway

Process and fidelity indicators are organized around Proctor et al.'s eight-domain taxonomy and kept structurally separate from outcome indicators rather than blended into one scale, the same two-tier separation this literature draws throughout. The process-indicator family, completion, fidelity, training participation, data-review cadence, is the Fixsen/CFIR core-driver list: those are the specific levers this literature identifies as differentiating full implementation from stalled rollout, weighed against what the platform can actually observe. The quarterly Collaborative Learning Group cycle (D4/D5) is a PDSA-style networked improvement community in the Bryk et al. sense, with the composite index serving as the shared cross-site measure that carries what one quarter's cycle learns into the next. Because Baffsky et al.'s review finds its strongest evidence specifically for continuous monitoring joined to performance feedback on fidelity, not for periodic recall surveys and not, yet, as strongly for adoption, BHPM's platform analytics is the primary implementation-measurement instrument, with formal review points layered on top of that signal rather than substituting for it. That is the indicator architecture this literature supports.

IMPACTER PathwaySanta Clara County Office of Education

This site publishes no student-level data. All figures are aggregate, de-identified, and reviewed under FERPA, COPPA, and California student privacy law (AB 1584 / SOPIPA).

Prepared in support of SCCOE under BHSOAC Contract No. 25BHSOAC019.