What Does Evidence Look Like for Assessments?
Assessment tools — screeners, diagnostics, progress monitors, unit tests — face a different evidence bar than curricula. Research-Validated Measures is an independent validity review that tests classification accuracy, reliability, and validity against external measures, and hands you a technical report, an expert-review badge, and press-kit material you can put in front of districts and reviewers.
Formal reviews of assessment tools — the kind conducted by the National Center on Intensive Intervention (NCII) — are rigorous, competitive, and infrequent. A company can spend years building a genuinely strong instrument and still be waiting for a formal review slot.
Research-Validated Measures is the intermediate step. It’s not a preliminary pass at a formal review — it’s a real, independent expert evaluation of your instrument’s measurement quality, producing evidence you can use immediately while you build toward (or alongside) a formal certification.
Measurement Evidence, Not Impact Evidence
Assessment tools face a different evidence bar than curricula, and companies coming from the curriculum side are often surprised by it. You’re not primarily trying to prove that students learned more — an assessment doesn’t teach anything, and its effect on outcomes runs entirely through what adults do with the data. You’re trying to prove that your instrument measures accurately and that its thresholds identify the right students.
That’s why an ESSA tier rating is usually the wrong target for a standalone assessment. ESSA describes program impact; NCII-style validation describes measurement quality. Our review is built around the latter — the evidence districts and reviewers actually look for in a screener or diagnostic.
Five Things a Real Validation Review Checks
These are the same criteria formal reviewers like NCII apply — we run them early, so you know where you stand before a formal submission.
Classification Accuracy
Sensitivity, specificity, and area under the curve, computed against a measure outside your own product — reported separately by grade and time of year, since fall accuracy says little about spring accuracy.
Reliability
Stability of scores across forms and administrations — most important if the tool is meant to be given repeatedly over a school year.
Validity Against External Measures
Evidence that scores relate to independent measures of the same construct in the direction and magnitude theory predicts — and relate less strongly to constructs you aren’t claiming to measure.
Sample & Bias Analysis
Who the validation evidence came from, and whether the instrument performs comparably across student groups.
Growth & Decision Rules (Progress Monitors)
Alternate forms, slope estimates, and guidance on what a teacher should actually do when the data says an intervention isn’t working.
Built for Companies With a Tool, Not Yet a Formal Review Slot
You’ve built a screener, diagnostic, or unit assessment with a real design behind it, but you haven’t yet gone through — or aren’t yet eligible for — a formal review like NCII’s tools charts. This service gives you a credible, independent evaluation to use in the meantime.
EdTech Startups
Building an assessment tool and need evidence before pursuing formal certification.
Established EdTech Companies
Adding a new measurement product and need independent validation to support sales conversations.
Curriculum Developers
Building an assessment to accompany a program — like a unit test bank or embedded checkpoint.
Researchers & Product Teams
Exploring a new measurement approach and want a rigor check before broader rollout.
What Validation Looks Like
Technical Examination
Detailed design analysis, educational-standards alignment, and statistical-properties evaluation of your instrument as built.
Comparative Analysis
Benchmarking against existing measures and establishing preliminary validity evidence against an external criterion.
Implementation Insights
Real-world usability assessment and practical recommendations for improving measurement consistency.
Dissemination Support
Technical report preparation, expert-review badge, and press-kit material.
Three Stages of Research Recognition
Not every assessment needs to clear every criterion on a formal NCII review to be valid, reliable, and actionable in the meantime. The Research-Validated Measure badge marks the rigor of the validity evidence itself in three tiers, so you have credible, usable recognition at each stage — not just at the finish line.
Preliminary Evidence
Exploratory-level evidence: a defined sample and an initial look at reliability and face validity, often from a single setting. A credible first step while a fuller evidence base is being built.
Product-Level Validity ✓
Moderate rigor: a larger, more representative sample, consistency demonstrated across contexts, and a moderate-to-strong correlation with an established external measure with real predictive power.
Cross-Product Level Validity ✓+
Our highest tier: a large, representative sample tested across multiple contexts, with multiple forms of reliability and validity evidence and low classification error — the level of rigor closest to what a formal, published review requires.
Exact technical thresholds for each tier vary somewhat by assessment type (survey measure, skills assessment, or embedded assessment) and are set by our research team on a study-by-study basis.
Real Validation Work, With Real Numbers
These are actual LXD Research validation studies and expert reviews — not hypothetical case studies. Full reports are available on our Published Studies page or on ResearchGate.
MindPlay — Signals Universal Screener
In a concurrent validity study of 643 students in grades 2–8 in Bridgeport, Connecticut, the screener’s Reading Level metric showed strong reliability across three benchmark windows (r = .94–.98) and correlated with DIBELS composite scores between .59 and .79. Classification agreement with ReadingPlus InSight in grades 7–8 showed no statistically significant differences.
Forefront Education — Universal Screener for Numeracy Skills (USNS)
LXD Research served as expert reviewer for Forefront’s USNS, examining classification accuracy against STAR Math and content alignment, and producing a summary brief of the technical documentation (2020–2025).
View on ResearchGate →95 Percent Group — 95 Phonics Core Program Unit Assessments
A recent client’s reading assessment tool was built to accompany their curriculum; with a few adjustments, it was validated as a standalone diagnostic. Results showed 95 PCP Unit Assessments were positively correlated with iReady Reading and STAAR, with strong internal correlations across checkpoints supporting convergent validity.
View on ResearchGate →Edpuzzle — Teacher-Created & Product-Created Quizzes
Nineteen pairs of scores comparing Edpuzzle quiz performance to Star Math and Reading assessments were analyzed; 95% of correlations were of large or medium strength, with particularly strong relationships for grades 5–6 on Edpuzzle Originals.
View the Efficacy Brief →A Bridge, Not a Destination
These expert reviews aren’t a replacement for formal certification — they’re a strategic stepping stone that builds momentum and evidence while you prepare for a more comprehensive review, and that you can use immediately for sales and district conversations.
Technical Report
A full write-up of the statistical and design analysis, usable in RFPs and procurement conversations.
Expert Review Badge
A credibility marker for your website and marketing while you build toward formal certification.
Press Kit
Ready-to-use language and materials for announcing the review to press, partners, and prospective districts.
Improvement Roadmap
Specific, actionable next steps for strengthening measurement consistency ahead of a formal submission.
Who benefits: assessment developers get expert guidance before formal review cycles and a way to build a robust evidence base incrementally. Educational institutions gain confidence in emerging assessment tools and access to more innovative measurement approaches, backed by independent review rather than marketing claims.
Frequently Asked Questions
What does it mean for an assessment to be validated?
How is this different from a formal NCII review?
Can our assessment earn an ESSA tier rating instead?
What’s the most common reason a validation study gets flagged by reviewers?
Do you help with data access and school recruitment?
Ready to Know Where Your Assessment Actually Stands?
Schedule a free consultation to talk through your instrument, what a validation review would look like for your specific case, and a realistic timeline and investment.