An assessment is only as useful as the decision it informs — and the only way to know whether a tool informs that decision well is to test it against something independent. Validation is the unglamorous work of checking whether a screener predicts what it claims to predict, whether its cut scores sort students correctly, and whether its results hold steady across a school year.
This article covers three things: the review bodies that evaluate assessments and what each of them actually checks, four tools LXD Research has examined directly, and what evidence a company building a new assessment will need to produce. For an overview of assessment types themselves, see our companion piece on the different types of assessments.
What Does “Data for the Teacher” Actually Mean?
The assessments in this article share a defining characteristic: the student completes them, and the data goes to an adult who then decides what happens next. There’s no feedback loop back to the student in the moment. That isn’t a flaw — it’s a design choice that suits a specific purpose. Screeners, progress monitors, and unit assessments are tools for teacher decision-making: identifying who needs intervention, placing students in the right instructional group, and confirming whether a unit of instruction did what it was supposed to do.
What separates well-designed tools in this category from poorly designed ones is usually specificity. A score that tells a teacher a student is “struggling in reading” isn’t very useful. A score that tells a teacher a student has not yet mastered consonant blends but has solid phonemic awareness — and that this pattern predicts difficulty on the state reading assessment — is. The best assessment tools in this space are built around that kind of targeted, actionable signal.
That is also what makes them testable. A tool that promises only a vague sense of who is behind can’t really be wrong. A tool that claims a specific threshold identifies students at risk of missing grade-level benchmarks is making a checkable claim — and validation is the process of checking it.
Who Evaluates Assessment Quality — and What Does Each Reviewer Actually Check?
Several organizations publish evaluations of edtech products and assessments, and they’re often discussed as though they’re interchangeable stamps of approval. They aren’t. Each answers a different question, and knowing which question a given badge answers is the difference between reading it correctly and over-reading it.
| Reviewer | What It Evaluates | The Question It Answers | What You Get |
|---|---|---|---|
| NCII | The instrument itself | Is this screener or progress monitor technically sound? | Tools chart ratings by standard |
| Digital Promise | Product design, and separately, impact | Is it built on research (Tier 4)? Does it show outcomes (Tier 3)? | Product certification, renewable every two years |
| Evidence for ESSA | Program impact | Did students learn more because of this program? | Strong, Moderate, or Promising rating |
| EduEvidence | The overall evidence portfolio | How much verified evidence stands behind this product? | Bronze, Silver, or Gold badge |
How Does NCII Review Assessments?
The National Center on Intensive Intervention, housed at the American Institutes for Research, maintains a set of tools charts covering academic and behavioral screening, progress monitoring, and intervention. For assessments, the review is squarely about measurement quality. External Technical Review Committees of content and methodological experts rate submitted tools against published rubrics covering:
- Classification accuracy — whether the cut score correctly separates students at risk from students who aren’t, evaluated through sensitivity, specificity, and area under the curve, and rated separately for each criterion measure and time of year.
- Reliability — whether scores are stable and consistent across forms, raters, and administrations.
- Validity — whether scores relate as theory predicts to an external measure of the same construct.
- Sample representativeness and bias analysis — who the evidence was gathered from, and whether the tool performs comparably across student groups.
Three features of the process are worth understanding before reading a chart. Submission is voluntary, so absence from a chart means a tool wasn’t submitted, not that it failed. Tools are rated against fixed criteria rather than ranked against one another, so the chart doesn’t produce a winner. And NCII states plainly that appearing on a chart is not an endorsement — the ratings are inputs to a local decision, not a recommendation.
Why Doesn’t an Assessment Usually Get an ESSA Tier Rating?
This is where the distinction gets practically important. ESSA evidence tiers describe whether a program caused improved student outcomes, and the review bodies oriented toward impact — Evidence for ESSA, Digital Promise’s Evidence-Based Edtech certification, the higher EduEvidence levels — are asking that question. But an assessment doesn’t teach anything. Its entire effect on student outcomes is mediated by what a teacher or team does after reading the results.
That’s not a weakness in assessments; it’s a description of their role. It does mean, though, that judging a screener by whether it has an ESSA Tier I rating is close to a category error. The right questions for an assessment are whether it measures accurately, whether its thresholds identify the right students, and whether the results are specific enough to route a student to the right instruction. The intervention on the other end of that routing is what an ESSA tier rating properly describes.
Reading badges correctly: A design certification tells you a product’s rationale is documented and grounded in research. An impact rating tells you students did better. A technical rigor rating tells you the instrument measures what it says it measures. A product can hold one and not the others, and each is worth something different depending on the decision you’re making.
How Do the 95 Percent Group Screeners Work — and What Did LXD Research Find?
95 Percent Group has developed two distinct assessment tools that serve different moments in the intervention cycle — the Phonics Screener for Intervention (PSI) and the Phonological Awareness Screener for Intervention (PASI) — and LXD Research has conducted validity work on both. Understanding how they work, and how they differ from the curriculum’s unit assessments, illustrates how assessment and instruction can function as a coherent system rather than separate products.
The PSI and PASI are brief, individually administered progress monitoring tools given every three weeks — but only to students already identified as needing Tier 2 or Tier 3 intervention. They aren’t whole-class assessments. Students who are on track don’t take them; the students receiving intervention get a frequent, skill-specific signal about whether that intervention is working. The PSI uses pseudowords and sentences specifically so students must decode rather than recall, with mastery set at 90% accuracy, and three alternate forms make it practical for repeated administration.
LXD Research’s validity study examined the PSI across 2,780 students in grades K–3 in two demographically diverse districts — one in California, one in New Mexico — and validated it against four independent external assessments: iReady and Acadience in California, DIBELS 8 and iStation in New Mexico.1 The study addressed four distinct validity questions:
- Does the skill sequence reflect real developmental progression? It does — basic phonics is mastered October through December, syllables not until March through April, the pattern you’d expect from genuine phonics development.
- Does it correlate more strongly with phonics-proximal outcomes than with vocabulary and comprehension? Yes, consistently across grades and assessments.
- Do grade-specific mastery thresholds distinguish on-track from off-track students? Yes, with significant odds ratios at every grade and instrument combination.
- How does its predictive strength compare to the external assessments’ ability to predict each other? In kindergarten, the PSI’s associations with external outcomes exceeded the external assessments’ ability to predict one another — a strong signal that it measures something real, and does so efficiently.
The unit assessments within the Phonics Lesson Library work differently — given at the end of an instructional unit, they function more like a final exam than a check-in. Where the PSI answers “is this student on track during intervention?”, the unit assessments answer “did this unit of instruction produce the mastery it aimed for?” LXD Research validated these as well, examining whether unit assessment mastery predicted outcomes on broader literacy measures — findings that support their use as meaningful summative checkpoints within the intervention curriculum, not just practice exercises.
What Does Validation Look Like for a Whole-School Reading Screener?
MindPlay Signals, the company’s universal screener, takes a wider-angle approach — screening across multiple reading domains including phonics, fluency, listening vocabulary, and reading level, through an adaptive administration that adjusts difficulty based on each student’s responses. LXD Research conducted a concurrent validity study of the screener among 643 students in grades 2–8 in Bridgeport, Connecticut, a high-need district where 87% of students qualified for free or reduced-price lunch and over a third were classified as English language learners.2 That population context matters: the findings are most directly applicable to struggling readers rather than students performing at or above grade level, and the report is transparent about that limitation.
The study examined both internal validity — whether the screener’s metrics hold up consistently over time — and external validity against two established assessments: DIBELS for grades 2–6 and ReadingPlus InSight for grades 7–8. What it found:
- Strong reliability. The primary Reading Level metric held up across three benchmark windows (r = .94–.98).
- Solid external agreement. Correlations with DIBELS composite scores ranged from .59 to .79 across grades, most strengthening from beginning to middle of year.
- Consistent classification in the upper grades. Agreement with ReadingPlus InSight for grades 7–8 showed no statistically significant differences in how the two assessments sorted students into proficiency categories.
- Sensitivity to real progress. Students using MindPlay 60 or more minutes per week gained nearly a full grade level on Reading Level in the first half of the year, against roughly half a grade level for lower-usage peers — evidence the screener captures growth, not just a stable trait.
One finding worth noting for practitioners: MindPlay consistently classified higher proportions of students as below benchmark compared to DIBELS, reflecting more stringent criteria for grade-level proficiency. That’s not a flaw — a screener designed to flag students needing intervention should probably err on the side of caution — but it’s worth understanding when interpreting results alongside other assessments.
A separate LXD Research efficacy study on MindPlay Reading examined outcomes against NWEA MAP, a norm-referenced instrument, to validate that skill-level gains translated into broader achievement. Layering criterion-referenced and norm-referenced evidence this way is increasingly common in well-designed efficacy research, and it’s part of what gives a study credibility across different audiences.
Read the MindPlay Validity Report →How Does the Same Logic Apply in Math?
Forefront Education’s Universal Screener for Number Sense (USNS) brings the same data-for-teacher logic to math. LXD Research’s role here was as expert reviewer rather than study conductor: we advised Forefront on their validation approach, reviewed and provided feedback on their technical documentation, and produced an accessible summary of the evidence they accumulated. The underlying validation work — conducted by Forefront Education across multiple studies from 2020 through 2025 — is substantial: several thousand students per grade level, kindergarten through fifth grade, with a nationally representative sample spanning all U.S. Census regions.3 External validity was established against Renaissance STAR Math, with overall classification accuracy ranging from 73% to 95% across grades and time points, and AUC values in the “very good” to “excellent” range for most grades. An expert panel of 19 education professionals confirmed the screener’s alignment to grade-level number sense content, rating it 3.62 out of 4.0. The USNS is a hybrid instrument — combining one-on-one interviews with group-administered tasks — designed to give teachers diagnostic information specific enough to guide instruction, not just a risk flag.
Read the Forefront USNS Validity Brief →What Pattern Holds Across Validated Tools?
| Tool | Type | LXD Research’s Role | What Was Examined |
|---|---|---|---|
| PSI & PASI | Progress monitor for intervention | Study conductor | Developmental sequence, convergent validity, threshold accuracy |
| PLL Unit Assessments | Unit-level summative | Study conductor | Whether unit mastery predicts broader literacy outcomes |
| MindPlay Signals | Adaptive universal screener | Study conductor | Reliability over time; agreement with DIBELS and ReadingPlus InSight |
| Forefront USNS | Universal screener (math) | Expert reviewer | Classification accuracy against STAR Math; content alignment |
The through-line is narrowness of purpose. Each of these tools was built to answer one question for one audience, and that constraint is what made the validation work tractable in the first place. A screener that tries to be a diagnostic, a progress monitor, and a growth measure at once is difficult to evaluate precisely because it never commits to a claim specific enough to test.
If You’re Building an Assessment, What Evidence Will You Need?
A growing number of edtech companies are entering the assessment market, and the reasoning is usually sound. Established commercial tests are frequently criticized on three fronts: they consume too much instructional time, their results go stale within weeks relative to the pace of teaching, and they report at a grain size that doesn’t correspond to any instructional decision a teacher can act on. A shorter, more frequent, more skill-specific instrument is a legitimate answer to all three.
But the evidence bar for an assessment is different from the one for a curriculum, and companies coming from the curriculum side are often surprised by it. You are not primarily trying to prove that students learned more. You’re trying to prove that your instrument measures accurately and that its thresholds identify the right students. That means measurement evidence, not impact evidence:
- Classification accuracy against an external criterion. Sensitivity, specificity, and area under the curve, computed against a measure outside your own product — and reported separately by grade and time of year, because accuracy in fall says little about accuracy in spring.
- Reliability. Stability across forms and administrations, which matters most if you intend the tool to be given repeatedly.
- Validity. Evidence that scores relate to independent measures of the same construct in the direction and magnitude theory predicts — and that they relate less strongly to constructs you aren’t claiming to measure.
- A defensible sample and a bias analysis. Who the evidence came from, and whether the instrument performs comparably across student groups.
- Growth and decision rules, for progress monitors. Alternate forms, slope estimates, and guidance on what a teacher should actually do when the data says the intervention isn’t working.
In practice, the bottleneck is rarely the statistics. It’s access: a real district sample of adequate size, and an external criterion measure administered to the same students in the same year. That’s a recruitment and data-sharing problem before it’s an analytic one, and it’s the stage where internal validation efforts most often stall.
The recurring mistakes are worth naming plainly. Validating against another assessment you also own establishes internal consistency, not external validity. Setting cut scores after seeing outcome data, rather than defining a risk criterion in advance, produces thresholds that look better on paper than in a new district. And a sample drawn from a single friendly partner district will rarely satisfy a reviewer asking whether the tool works anywhere else. None of these are fatal — but each is far cheaper to avoid at design time than to fix after the fact.
At LXD Research, we’ve validated assessment tools across a range of categories, and the pattern that holds across the strongest performers is consistent: the data is specific, the connection to an instructional decision is clear, and the tool was designed for a defined question rather than for general measurement. That clarity is what separates assessment that changes practice from assessment that fills a dashboard — and it’s also what makes a tool straightforward to validate.
Frequently Asked Questions
What does it mean for an assessment to be validated?
What is NCII and what does it review?
What is the difference between NCII and Evidence for ESSA?
Can an assessment earn an ESSA evidence rating?
What does a Digital Promise certification mean?
What is EduEvidence certification?
How do I validate a new assessment product?
Who conducts independent validity studies for edtech assessment products?
What evidence do districts want before buying a screener?
What mistakes do companies make when validating their own assessments?
Who validated the 95 Percent Group PSI?
What did LXD Research find about MindPlay Signals?
- LXD Research, Phonics Screener for Intervention (PSI) Validity Study. 2,780 students, grades K–3, across districts in California and New Mexico. Full report.
- LXD Research, MindPlay Universal Screener Concurrent Validity Study. 643 students, grades 2–8, Bridgeport, Connecticut. Full report.
- Forefront Education, USNS technical documentation, 2020–2025. LXD Research served as expert reviewer and produced the summary brief. Full brief.
Building an Assessment? Let’s Talk About the Evidence.
LXD Research designs and conducts independent validity and efficacy studies for edtech companies — including school recruitment, IRB submission, and data sharing agreements — and helps districts interpret the evidence behind the tools they’re considering.
Schedule a Free Consultation Browse Published Studies