ESSA evidence requirements aren’t one-size-fits-all — they scale with the tier you’re targeting, and each tier adds to the one below it rather than replacing it. Knowing exactly what your target tier demands before you start is what keeps a study from being redesigned halfway through.
How Do the Four ESSA Tiers Compare?
Each tier adds a research design requirement on top of the one below it: Tier IV asks for a rationale, Tier III for a correlational analysis, Tier II for a matched comparison group, and Tier I for random assignment.
| Tier | Design Required | Comparison Group? | Minimum Sample |
|---|---|---|---|
| Tier IV Demonstrates a Rationale | Logic model + literature review | Not required | 5+ users (user testing/co-design) |
| Tier III Promising Evidence | Correlational study with statistical controls for bias | Not formally required | 50+ students minimum |
| Tier II Moderate Evidence | Quasi-experimental design | Matched, with documented baseline equivalence | 350+ participants, multiple sites |
| Tier I Strong Evidence | Randomized controlled trial | Randomly assigned | 350+ participants, multiple sites |
Sample-size figures reflect how Digital Promise and Evidence for ESSA — the only two U.S. nonprofit certifiers in this space — operationalize each tier for review. The ESSA statute itself doesn’t specify exact participant counts, and other evaluators or state rubrics may apply different thresholds.
What Does Every Tier Require, at Minimum?
Every ESSA tier — even Tier I — rests on the same baseline documentation that Tier IV requires on its own:
- A logic model connecting the product’s design to research: the problem it addresses, its inputs and activities, its outputs, and the short- and long-term outcomes it should move, each grounded in cited research rather than asserted.
- A literature review or annotated bibliography establishing that the model’s assumptions are supported by existing research on similar interventions.
- A FERPA-compliant data privacy letter, since any study touching student records has to document how that data is protected.
- Public accessibility of the research — a report a reviewer, a district administrator, and eventually a buyer can all actually find and read.
Think of this as the Tier IV floor. Every tier above it adds a research design requirement on top — none of them subtract from this baseline.
What Does a Complete Tier IV Package Require?
A qualifying logic model has to do more than list features — it needs to specify, for each strategy or activity the product carries out, which research literature supports it and which outcome it’s meant to move. Reviewers are checking for a genuine two-way connection: every claim in the literature review should map to something the product actually does, and every feature the model highlights should be backed by a citation, not asserted on its own authority. There’s no fixed magic number of citations that guarantees approval — what reviewers are actually checking is whether the model’s core claims are each traceable to real, cited research, not whether a citation count hits a threshold.
Digital Promise is the U.S. nonprofit certifier for Tier IV, through its Research-Based Design certification — it’s one of only two U.S. nonprofit certifiers in this space (Evidence for ESSA is the other, though it covers Tiers I through III rather than Tier IV). EduEvidence, which issues Bronze, Silver, and Gold badges depending on evidence strength, also accepts Tier IV documentation, but it’s an international nonprofit rather than a U.S.-based one. The most common reason a Tier IV application gets sent back isn’t a weak idea — it’s a literature review that wanders into research the product doesn’t actually implement, or a logic model written separately from the narrative that quietly contradicts it.
Tier IV doesn’t require outcome data, but it isn’t a pure paperwork exercise either — under Digital Promise‘s review criteria, a logic model is generally expected to be informed by at least a small amount of direct engagement with real users, typically a minimum of five participants through user testing or co-design sessions. That’s a low bar compared to any tier above it, but a rationale built without ever putting the product in front of a student or teacher reads as untested rather than research-based.
What Does Tier III Add on Top of the Floor?
Tier III — Promising Evidence — requires correlational research with statistical controls for selection bias: put simply, the analysis has to account for the fact that the students or schools using the product may differ systematically from those that don’t, rather than crediting the product for a difference that was already there. This doesn’t require a control group in the experimental sense, but it does require an outcome measure that’s independent of the product itself — your own internal assessment doesn’t qualify as evidence of your own effectiveness, no matter how well it’s designed.
Sample size and site requirements at Tier III are meaningfully more flexible than Tiers I and II, which is exactly why it’s accessible to smaller studies and existing data — Digital Promise generally looks for a minimum of about 50 students. But “more flexible” isn’t “no requirement” — a single classroom’s worth of data rarely clears the bar. Tier III is also sometimes where a study designed for a higher tier lands after review, for reasons that go beyond sample size alone — this diagnostic walks through what determines your actual tier in more detail.
What Does Tier II Require Beyond Tier III?
Tier II — Moderate Evidence — requires a matched comparison group: students or schools using the product studied alongside a comparable group that isn’t, with documented baseline equivalence between the two showing they were similar before the intervention started. This is also the most commonly overclaimed tier — a study with a comparison group that wasn’t actually matched on relevant characteristics, or wasn’t validated as equivalent at baseline, doesn’t meet the bar no matter how positive the outcome looks.
Tier II also calls for multi-site validation — under Evidence for ESSA‘s criteria (Digital Promise doesn’t certify Tiers I or II), the same 350+ participants across multiple school sites threshold that Tier I requires, evidence that the effect holds across more than one district, not just a single favorable implementation. LXD Research has supported more than 25 studies to Evidence for ESSA approval. What separates the tiers isn’t the sample size — it’s whether students were randomly assigned (Tier I) or placed in a matched comparison group (Tier II).
The line that trips people up most: a positive result from one classroom, without a matched comparison group and without baseline equivalence documented, is compelling marketing but not Tier II evidence.
What Does Tier I Require?
Tier I — Strong Evidence — requires a randomized controlled trial: students or schools randomly assigned to use the product or not, with at least 350 participants across multiple school sites under Evidence for ESSA’s review criteria. This always requires a new, prospective study. Random assignment can’t be added retroactively — it’s a design decision that has to be built in from the first day of the study, which is why Tier I is never a candidate for a Look Back Study using historical data.
The Mistakes That Get Studies Rejected or Downgraded
Almost every rejected or downgraded application traces back to one of two moments: a decision locked in before the study ever starts, or something that only becomes visible once results are in.
Avoidable at recruitment and design — these are still fixable before data collection begins:
- Using an internal outcome measure. An assessment your own product is built around can’t also serve as independent proof that the product works.
- Using a researcher-created measure. An instrument built specifically for the study, without its own established validity record, invites the same skepticism as an internal one — reviewers want a measure with a track record of its own, not one designed to fit the study.
- Single-site samples presented as evidence for Tier I or Tier II, both of which call for validation across more than one school or district.
- Skipping the plan for baseline equivalence. A comparison group has to be established as similar to the treatment group before the study begins — this isn’t something that can be reconstructed convincingly after the fact.
Only visible once results come in — no amount of design care prevents these:
- Attrition. Enough students or schools dropping out over the course of a study can shrink the sample below what a tier requires, or bias the result if the students who leave aren’t representative of the ones who stay.
- Ignoring clustering in the analysis. When students are nested within classrooms and classrooms within schools, the statistics need to account for that structure — treating clustered data as if every student were an independent observation inflates the apparent significance of the result.
- Negative or null results in part of the range. A study spanning multiple grades that shows a negative or non-significant effect in even one grade band can undercut or disqualify the claim for the entire range, regardless of how strong the results look elsewhere.
The practical distinction: the first set of mistakes can be corrected simply by making a different choice before the study begins. The second set can’t be fixed after the fact — which is exactly why measure selection, site count, and baseline documentation deserve more scrutiny upfront than they usually get.
Do Requirements Vary by Evaluator and State?
Yes — and this is where a company can do everything right on paper and still land somewhere unexpected. Evidence for ESSA (the Johns Hopkins database), Digital Promise, and individual state review boards don’t all apply the requirements identically. Arizona’s Move On When Reading program is a useful example: the state requires a qualifying product to reach one of the top three ESSA tiers and demonstrate a statistically significant positive effect on a relevant outcome — a materially stricter reading than a general Tier III self-designation might assume.
Submitting to the wrong evaluator first can be costly in a way that’s easy to miss: a study that lands as Tier III with one reviewer might be read more conservatively by a state board, and a rating that’s already been published is hard to walk back upward later. Understanding who’s actually going to be reading the evidence — before submission, not after — is what keeps a strong study from being undersold by the wrong first audience.
Frequently Asked Questions
What are the ESSA Tier IV requirements?
What’s the difference between ESSA Tier II and Tier III?
What does ESSA Tier I require?
Can a product’s own assessment count as ESSA evidence?
Do all states apply the same ESSA evidence requirements?
What are the most common mistakes that get ESSA applications rejected?
Not Sure If Your Research Meets the Requirements?
LXD Research reviews what you have and tells you exactly where it lands — and what it would take to go higher.