Skip to content
Home » Validating Assessments

Validating Assessments

Is Your EdTech Tool an Assessment?

What Does Evidence Look Like for Assessments?

Assessment tools — screeners, diagnostics, progress monitors, unit tests — face a different evidence bar than curricula. Research-Validated Measures is an independent validity review that tests classification accuracy, reliability, and validity against external measures, and hands you a technical report, an expert-review badge, and press-kit material you can put in front of districts and reviewers.

Formal reviews of assessment tools — the kind conducted by the National Center on Intensive Intervention (NCII) — are rigorous, competitive, and infrequent. A company can spend years building a genuinely strong instrument and still be waiting for a formal review slot.

Research-Validated Measures is the intermediate step. It’s not a preliminary pass at a formal review — it’s a real, independent expert evaluation of your instrument’s measurement quality, producing evidence you can use immediately while you build toward (or alongside) a formal certification.

Measurement Evidence, Not Impact Evidence

Assessment tools face a different evidence bar than curricula, and companies coming from the curriculum side are often surprised by it. You’re not primarily trying to prove that students learned more — an assessment doesn’t teach anything, and its effect on outcomes runs entirely through what adults do with the data. You’re trying to prove that your instrument measures accurately and that its thresholds identify the right students.

That’s why an ESSA tier rating is usually the wrong target for a standalone assessment. ESSA describes program impact; NCII-style validation describes measurement quality. Our review is built around the latter — the evidence districts and reviewers actually look for in a screener or diagnostic.

Five Things a Real Validation Review Checks

These are the same criteria formal reviewers like NCII apply — we run them early, so you know where you stand before a formal submission.

01

Classification Accuracy

Sensitivity, specificity, and area under the curve, computed against a measure outside your own product — reported separately by grade and time of year, since fall accuracy says little about spring accuracy.

02

Reliability

Stability of scores across forms and administrations — most important if the tool is meant to be given repeatedly over a school year.

03

Validity Against External Measures

Evidence that scores relate to independent measures of the same construct in the direction and magnitude theory predicts — and relate less strongly to constructs you aren’t claiming to measure.

04

Sample & Bias Analysis

Who the validation evidence came from, and whether the instrument performs comparably across student groups.

05

Growth & Decision Rules (Progress Monitors)

Alternate forms, slope estimates, and guidance on what a teacher should actually do when the data says an intervention isn’t working.

Built for Companies With a Tool, Not Yet a Formal Review Slot

You’ve built a screener, diagnostic, or unit assessment with a real design behind it, but you haven’t yet gone through — or aren’t yet eligible for — a formal review like NCII’s tools charts. This service gives you a credible, independent evaluation to use in the meantime.

EdTech Startups

Building an assessment tool and need evidence before pursuing formal certification.

Established EdTech Companies

Adding a new measurement product and need independent validation to support sales conversations.

Curriculum Developers

Building an assessment to accompany a program — like a unit test bank or embedded checkpoint.

Researchers & Product Teams

Exploring a new measurement approach and want a rigor check before broader rollout.

What Validation Looks Like

1

Technical Examination

Detailed design analysis, educational-standards alignment, and statistical-properties evaluation of your instrument as built.

2

Comparative Analysis

Benchmarking against existing measures and establishing preliminary validity evidence against an external criterion.

3

Implementation Insights

Real-world usability assessment and practical recommendations for improving measurement consistency.

4

Dissemination Support

Technical report preparation, expert-review badge, and press-kit material.


Three Stages of Research Recognition

Not every assessment needs to clear every criterion on a formal NCII review to be valid, reliable, and actionable in the meantime. The Research-Validated Measure badge marks the rigor of the validity evidence itself in three tiers, so you have credible, usable recognition at each stage — not just at the finish line.

Research-Validated Measure badge

Preliminary Evidence

Exploratory-level evidence: a defined sample and an initial look at reliability and face validity, often from a single setting. A credible first step while a fuller evidence base is being built.

Research-Validated Measure badge with checkmark

Product-Level Validity

Moderate rigor: a larger, more representative sample, consistency demonstrated across contexts, and a moderate-to-strong correlation with an established external measure with real predictive power.

Research-Validated Measure badge with checkmark plus

Cross-Product Level Validity ✓+

Our highest tier: a large, representative sample tested across multiple contexts, with multiple forms of reliability and validity evidence and low classification error — the level of rigor closest to what a formal, published review requires.

Exact technical thresholds for each tier vary somewhat by assessment type (survey measure, skills assessment, or embedded assessment) and are set by our research team on a study-by-study basis.


Real Validation Work, With Real Numbers

These are actual LXD Research validation studies and expert reviews — not hypothetical case studies. Full reports are available on our Published Studies page or on ResearchGate.

Concurrent Validity StudyGrades 2–8

MindPlay — Signals Universal Screener

In a concurrent validity study of 643 students in grades 2–8 in Bridgeport, Connecticut, the screener’s Reading Level metric showed strong reliability across three benchmark windows (r = .94–.98) and correlated with DIBELS composite scores between .59 and .79. Classification agreement with ReadingPlus InSight in grades 7–8 showed no statistically significant differences.

Sample643 students
Reliabilityr = .94–.98
DIBELS Correlation.59–.79
View on ResearchGate →
Expert Review

Forefront Education — Universal Screener for Numeracy Skills (USNS)

LXD Research served as expert reviewer for Forefront’s USNS, examining classification accuracy against STAR Math and content alignment, and producing a summary brief of the technical documentation (2020–2025).

View on ResearchGate →
Validity Brief

95 Percent Group — 95 Phonics Core Program Unit Assessments

A recent client’s reading assessment tool was built to accompany their curriculum; with a few adjustments, it was validated as a standalone diagnostic. Results showed 95 PCP Unit Assessments were positively correlated with iReady Reading and STAAR, with strong internal correlations across checkpoints supporting convergent validity.

View on ResearchGate →
Efficacy Brief

Edpuzzle — Teacher-Created & Product-Created Quizzes

Nineteen pairs of scores comparing Edpuzzle quiz performance to Star Math and Reading assessments were analyzed; 95% of correlations were of large or medium strength, with particularly strong relationships for grades 5–6 on Edpuzzle Originals.

View the Efficacy Brief →

A Bridge, Not a Destination

These expert reviews aren’t a replacement for formal certification — they’re a strategic stepping stone that builds momentum and evidence while you prepare for a more comprehensive review, and that you can use immediately for sales and district conversations.

Technical Report

A full write-up of the statistical and design analysis, usable in RFPs and procurement conversations.

Expert Review Badge

A credibility marker for your website and marketing while you build toward formal certification.

Press Kit

Ready-to-use language and materials for announcing the review to press, partners, and prospective districts.

Improvement Roadmap

Specific, actionable next steps for strengthening measurement consistency ahead of a formal submission.

Who benefits: assessment developers get expert guidance before formal review cycles and a way to build a robust evidence base incrementally. Educational institutions gain confidence in emerging assessment tools and access to more innovative measurement approaches, backed by independent review rather than marketing claims.

Frequently Asked Questions

What does it mean for an assessment to be validated?
Validation means testing whether an instrument does what it claims. For a screener, that usually means checking classification accuracy (does the cut score correctly separate at-risk students from others), reliability (do scores hold steady across administrations), and validity against an external measure (does the score relate to an independent assessment of the same construct as expected).
How is this different from a formal NCII review?
NCII reviews are conducted by external Technical Review Committees against published rubrics, and submission is voluntary and infrequent. Our Research-Validated Measures review applies the same categories of evidence — classification accuracy, reliability, validity, sample and bias — as an earlier, faster intermediate step, not a substitute for the formal review itself.
Can our assessment earn an ESSA tier rating instead?
Rarely on its own. ESSA tiers describe evidence that a program caused better student outcomes, and an assessment’s effect on outcomes runs entirely through what adults do with the data. An assessment alone is better evaluated on measurement quality — which is what this review, and NCII-style reviews, are built to assess.
What’s the most common reason a validation study gets flagged by reviewers?
Validating against another instrument your company also owns shows internal consistency, not external validity. Setting cut scores after seeing outcome data, rather than defining a risk criterion in advance, is another common issue — as is drawing the validation sample from a single friendly partner district. Each is far cheaper to avoid at study-design time than to correct afterward.
Do you help with data access and school recruitment?
Yes. In practice, the bottleneck in validating an assessment is rarely the statistics — it’s getting a real district sample of adequate size with an external criterion measure administered to the same students. We help with recruitment, IRB submission, and data-sharing agreements as part of the process.

Ready to Know Where Your Assessment Actually Stands?

Schedule a free consultation to talk through your instrument, what a validation review would look like for your specific case, and a realistic timeline and investment.

Services are customized based on your instrument, your data, and your goals.