Skip to content
Home » I Have Research on My Program. What ESSA Level Is It?

I Have Research on My Program. What ESSA Level Is It?

The question usually arrives in an RFP, or from a state reviewer, or from a district curriculum director who has been told to document evidence for everything she buys: what ESSA level is your research? And you do have research — a pilot report, a white paper, a conference paper, a set of growth numbers from a district that loves you. What you don’t have is a confident answer.

This article gives you a way to work it out yourself: the diagnostic sequence that places almost any study, the details that quietly cost a tier, who actually decides, and why the second question — where should this evidence be submitted? — often matters more than the first.

The short answer: Your study’s design sets its ceiling — randomized for Tier I, matched comparison group for Tier II, correlational with statistical controls for Tier III, logic model alone for Tier IV. Sample size, statistical significance, and the independence of the outcome measure then determine whether the study reaches that ceiling or settles a tier below it. Which tier a reviewer actually awards depends on whose rubric is being applied, and those differ by state.

What Actually Determines a Study’s ESSA Tier?

A study’s ESSA tier is determined by its research design, sample, and statistical results — not by how large or impressive the findings are. The tiers rank how confidently a study can attribute an outcome to a program, which is a statement about methodology rather than product quality.

That distinction matters more than almost anything else in this conversation. A rigorously designed randomized trial that finds a small effect can be rated at the highest tier. A program with several studies and much larger reported gains can land at the lowest, because the designs behind those gains can’t rule out other explanations. So you aren’t asking “how good is my program?” You’re asking “what does the structure of this particular study allow anyone to claim?” The second question has a far more determinate answer.

Tier Design Required What It Establishes
Tier I
Strong
Randomized controlled trial, large and multi-site. The program caused the difference. Selection bias is eliminated by design.
Tier II
Moderate
Quasi-experimental study with a matched comparison group, large and multi-site. The program very likely caused the difference, with baseline equivalence standing in for randomization.
Tier III
Promising
Correlational study with statistical controls for selection bias. Program use is associated with better outcomes. Causation isn’t established.
Tier IV
Demonstrates a Rationale
A well-defined logic model grounded in research, plus an evaluation underway. The program is designed on principles that research supports.

Notice the last column of the Tier IV row. The statutory definition asks not only for a logic model but for an ongoing effort to examine the program’s effects. A rationale document filed and forgotten is weaker than the tier suggests; a rationale document paired with a study in the field is exactly what the tier describes.

How Do I Tell What ESSA Level My Study Is?

Work through three questions in order: is there a comparison beyond the students’ own starting scores, how was that comparison formed, and did the study meet the sample and significance requirements. Most research places itself within the first two.

1
Are student outcomes compared against anything beyond the same students’ own starting scores?
No Pre/post growth, usage data, and educator feedback cap you at Tier IV.
Yes Continue. The comparison can be non-using students, a prior-year cohort, lower-usage students, or published growth norms.
2
How was that comparison formed?
Random assignment Ceiling: Tier I.
Matched groups Ceiling: Tier II, if baseline equivalence is documented.
Statistical controls Ceiling: Tier III.
3
Did the study clear the execution bar for that ceiling?
Yes 350+ students, two or more sites, a statistically significant positive effect on an independent measure.
No Missing sample size or clustering drops you to Tier III. No significant effect means no tier at all.

Retrospective studies are a special case. A well-executed look-back analysis of data a district already collected can support a Tier II claim under many state rubrics, but Evidence for ESSA generally classifies post-hoc studies as Promising rather than Moderate. Where you intend to submit matters as much as how the study was run.

Not sure which branch your study falls on? LXD Research reviews the research you already have, confirms the tier your design supports, and gives you the documentation to back that claim up with district and state reviewers.

Request a Study Review

What Disqualifies a Study From an ESSA Tier?

Six issues account for most downgraded ratings: a non-independent outcome measure, a single-site sample, an analysis that ignores clustering, undocumented baseline equivalence, results that aren’t statistically significant, and research that’s too old to describe the current product.

  • The outcome measure is your own. If the only evidence of growth comes from an assessment embedded in your product, reviewers will discount it. The measure needs to be independent — a state test, a widely used benchmark assessment, or a standardized instrument the district administers for its own purposes.
  • The sample is one district. Tiers I and II require a multi-site sample, generally read as at least two schools or districts and at least 350 students. A single friendly partner site, however clean the data, caps you at Tier III.
  • The analysis ignores clustering. Students within a classroom aren’t independent observations. An analysis that treats them as though they are will overstate precision, and reviewers who catch it will move an otherwise qualifying study down a tier.
  • Baseline equivalence was never documented. A quasi-experimental design without evidence that the groups started in comparable places isn’t really a quasi-experimental design. This is the single most common reason a study that calls itself Tier II gets rated lower.
  • The findings weren’t statistically significant. Encouraging trends and directionally positive results don’t qualify for any tier. A study that finds no significant positive effect remains valuable for product development, but it isn’t evidence in the ESSA sense.
  • The research is dated. Several states apply a recency window — Arizona’s guidance points toward research conducted within roughly the last fifteen years. A strong study of a product version that no longer resembles what you sell can be technically sound and practically unpersuasive.

Design Sets the Ceiling. Execution Decides Whether You Reach It.

Nearly every disappointing tier conversation traces back to one of two failures, and they behave very differently. Design decisions are made before data collection and cannot be repaired afterward. Execution requirements are about thresholds and documentation, and some of them can still be met after the fact if the underlying data exists.

Locked at the start
The Ceiling

Whether students were randomly assigned. Whether a comparison group exists at all. Whether the outcome measure was given to both groups at the same points in time. These are structural, and no later analysis creates them.

Determined at the end
The Reach

Sample size, baseline equivalence documentation, clustering-appropriate analysis, statistical significance, and measure independence. Several of these can be strengthened by expanding the sample or reanalyzing existing district data.

This is why the timing of the conversation matters. A company that asks “what tier will this be?” before data collection can usually build toward the tier it wants at modest additional cost. A company that asks after the report is written is often choosing between accepting a lower tier and running a second study.

The practical takeaway: if you have an implementation in the field right now with no research plan attached to it, that’s the cheapest tier you will ever buy — and the window closes when the school year does.

Have a live implementation this year? We’ll tell you what evidence it could produce and what it would take to get there.

Talk to a Researcher

Who Decides What ESSA Level My Research Is?

Three different parties make the call, and they don’t always agree: the vendor self-designates, an independent clearinghouse may rate the study against published standards, and state review committees apply their own rubrics and can adjust a claimed tier.

  • You do, first. Vendors self-designate when responding to an RFP or submitting to a state list. This is legitimate as long as it’s honest and documented. Michigan’s Levels of Evidence Worksheet works exactly this way — vendors select a tier, submit the study, explain the design and outcomes in plain language, link to the research without a paywall, and document significance and sample details. The state’s Committee for Literacy Achievement then verifies the claim and can adjust it downward.
  • Independent clearinghouses do, second. Evidence for ESSA, run out of Johns Hopkins, and the What Works Clearinghouse apply published standards and publish the result. Their ratings carry weight precisely because they aren’t yours. They are also stricter than most state rubrics on clustering, sample size, and retrospective designs.
  • States do, third, and on their own terms. State requirements are not copies of the federal definitions. The same study can be Moderate evidence in one submission, Promising in another, and unrated by a clearinghouse.

Do All States Apply the Same ESSA Requirements?

No. States set their own evidence rules on top of the federal definitions, and the differences are consequential — Arizona requires a control group even at Tier III, which the federal correlational standard does not.

State What It Adds Beyond the Federal Definition
Arizona Requires a control group and a large sample at every qualifying tier, including Tier III. Core, supplemental, and intervention programs must each independently qualify. Applies a recency window of roughly fifteen years.
Massachusetts The CURATE bid splits its scored evaluation evenly between standards alignment and whether a curriculum is “efficacious in implementation” — 50 points each. The efficacy half has no defined pathway in the bid documents, and quality reviews like EdReports satisfy alignment rather than efficacy.
Michigan Vendors self-select a tier on an evidence worksheet; a state committee verifies and may adjust it. Ranking points scale with tier, and studies must be linked without a paywall. Evaluates study design, results, related studies, sample size and setting, and match to Michigan students.
Mississippi Senate Bill 2487 extends evidence requirements into the middle grades, with an eighth-grade reading gate phasing in from 2027.
States With Requirements 34 Now encourage or require evidence-based programs — each with its own rubric and submission window
Large District RFPs 25–35% Now require ESSA evidence, growing 10–15% annually since 2022
Tools With Evidence 25% Only a quarter of top edtech tools have ESSA-aligned evidence of any kind

Why Should Someone Else Review My Study Before I Submit It?

Because knowing your tier is only half the problem. The other half is knowing which states and clearinghouses your evidence will actually satisfy, and in what order to approach them — and that requires someone tracking rules across dozens of jurisdictions that change every legislative session.

A vendor can read the federal definitions and reach a defensible conclusion about a study’s design. What’s much harder to do from inside a company is answer the questions that follow, each of which turns on knowledge of the review landscape rather than knowledge of the study:

  • Which lists is this evidence actually good enough for? A study that falls just short of Evidence for ESSA’s threshold may comfortably clear several state rubrics. Knowing which ones saves you from writing off usable evidence.
  • Where should it go first? A published low rating is difficult to undo. Submitting a borderline study to the strictest reviewer before strengthening it can lock in a designation you’d rather have avoided.
  • When do the windows open? State adoption cycles have deadlines. Missing one can mean waiting a full procurement cycle, regardless of how good the evidence is.
  • Does the sample match the state’s students? Michigan explicitly evaluates the match between your study population and Michigan students. A strong study conducted in the wrong demographic context scores lower than a modest one conducted in the right one.
  • What does the portfolio need next? Most companies are better served by a cross-section of studies at varying levels than by one study pursued to perfection — so that a search on ERIC or Evidence for ESSA returns a substantive profile rather than a single entry.
  • What makes the claim land? A tier designation carries more weight when it arrives with the details reviewers look for — sample size and setting, the outcome measure, significance, baseline equivalence. Documenting those turns a self-designation into something a procurement committee can verify, which is what separates an evidence claim that opens doors from one that invites questions.

This is the work LXD Research does before designing anything: read what you already have, place it accurately, and map it against the states and clearinghouses where it will do the most good. Since 2020 we’ve had more than 25 studies approved by Evidence for ESSA — 64% at the two most rigorous tiers — and supported state approval processes in Arizona, Michigan, Utah, Arkansas, and Wyoming. That experience is mostly useful to clients as pattern recognition: we’ve seen which arguments reviewers accept and which they send back.

What Do You Get From a Study Review?

Every LXD Research review produces four things you can use immediately: an educator-friendly summary written for district audiences, a named researcher’s authorship on that summary, an LXD Research ESSA badge reflecting the tier the evidence supports, and publication to our ResearchGate profile so the work is discoverable outside your own website.

The reason this matters is timing. Third-party review is slow by design — clearinghouse queues run months, and state adoption cycles open once a year. Meanwhile your sales team is answering evidence questions this quarter. The deliverables below are meant to close that gap without overstating what you have.

  • An educator-friendly summary, co-branded with your product. Effect sizes translated into the terms practitioners actually use, educator voices included, benchmark context provided. It’s written to stand on its own so a rep can hand it to a curriculum director or attach it to a conference session, without requiring anyone to read a full technical report.
  • Named authors on the work. Our researchers put their names and credentials on what they write. A summary signed by an outside researcher who can be looked up carries a different weight with district reviewers than an unsigned internal white paper — and 75% of educators report trusting tools more when they’re independently validated, while trusting company marketing considerably less.
  • An LXD Research ESSA badge you can use right away. The badge reflects the tier your evidence supports under an independent researcher’s assessment, and it’s available as soon as the review is complete. It isn’t a substitute for an Evidence for ESSA or state designation, and we’re explicit about that distinction — it’s what you use while those reviews are pending. Our badge criteria and definitions are published, so anyone can check what a badge means.
  • Publication to ResearchGate and ERIC. Completed studies are indexed where researchers, state reviewers, and procurement teams already search — more than 100 LXD Research reports now sit in ERIC and academic databases. Evidence hosted only on your own site invites the question of who produced it; evidence in an independent index doesn’t.

Taken together, that’s the difference between saying your program has evidence and being able to show someone where it lives, who wrote it, and what standard it was held to.

Send us the study you have. Our Evidence Scorecard covers up to three studies per product — an expert review, a written scorecard placing each study against the criteria that matter, and a 60-minute strategy call on where to submit and what to strengthen first.

Get Your Evidence Scorecard

What If My Study Lands Lower Than I Hoped?

A lower tier is a real evidence claim, not a failure — and with roughly three quarters of widely used edtech tools carrying no ESSA-aligned evidence at all, a documented Tier III or Tier IV still puts you ahead of most of the market. There are usually three paths forward.

  • Reanalyze what you already have. If the partner district has records for students who didn’t use the program, a matched comparison group may already exist inside data in hand. This is the least expensive route to a higher tier and is available more often than companies assume.
  • Run a look-back study. Existing implementations and district-administered assessments can support a comparison design built retrospectively — no new data collection, no imposition on partner schools, and a workable path to Tier III or Tier II depending on how the comparison is formed.
  • Treat the current study as a floor. Document the logic model so the Tier IV claim is solid, launch a prospective study on this year’s implementations, and let the evidence base build in layers.
  • What won’t work: adding random assignment retroactively. Tier I always requires a new prospective study, so a Tier I ambition is a planning decision rather than an analysis decision.

That last sequence — rationale documented, promising evidence published, moderate evidence in the field — is what a mature research portfolio actually looks like. It also reads to reviewers as a company that takes evidence seriously, rather than one that produced a single study and stopped.

Frequently Asked Questions

How do I know what ESSA level my research is?
Answer three questions in order:
  • Are outcomes compared against anything beyond the students’ own starting scores? If not, your ceiling is Tier IV.
  • How was the comparison formed? Random assignment points to Tier I, a matched group with documented baseline equivalence to Tier II, statistical controls for selection bias to Tier III.
  • Did it clear the execution bar — 350+ students across two or more sites, a statistically significant positive effect on an independent measure?
Design sets the ceiling; execution determines whether the study reaches it.
Can I claim an ESSA tier for my own study without outside review?
Yes. Self-designation is standard practice and is how most state submission processes are built — the vendor selects a tier and provides documentation, and a review committee verifies it. What you can’t do is claim a tier the design doesn’t support. Michigan’s worksheet allows the state’s review committee to adjust a vendor’s self-selected tier, so an overstated claim tends to surface at the least convenient moment.
Does a pre/post study with no comparison group count as ESSA evidence?
Not for Tiers I through III. Students grow over a school year regardless of what program they use, so a single-group pre/post design can’t separate the program’s contribution from ordinary development and instruction. Growth data of this kind is useful for product development and district conversations, and it can support a Tier IV rationale, but it doesn’t establish impact in the way the higher tiers require.
How many students does an ESSA Tier I or Tier II study need?
The statute calls for a large, multi-site sample; reviewers generally operationalize that as at least 350 students across two or more schools or districts. Studies that meet every other requirement but fall short on sample size, or that fail to account for clustering, are typically rated at Tier III instead.
Can a peer-reviewed journal article be a low ESSA tier?
Yes, and it surprises people. Peer review evaluates whether a study was conducted and reported competently; ESSA tiers evaluate whether the design supports a causal claim about student outcomes. A published article using a single-group design, a small sample, or an outcome measure internal to the product can be sound scholarship and still sit at Tier III or Tier IV.
Why does the same study get rated differently by different reviewers?
Because different rubrics answer different questions. Clearinghouses apply published methodological standards covering baseline equivalence, attrition, clustering, and sample size. State panels weigh additional factors — how closely the study population resembles their students, how recent the research is, whether the program is currently available. Arizona requires a control group even at Tier III, which the federal correlational standard does not. This is why placement advice matters as much as the tier itself.
Who can review my research and tell me what ESSA tier it qualifies for?
Independent research firms that work across multiple state submission processes are best positioned to do this, because the answer depends on both methodology and jurisdiction. LXD Research conducts study reviews for edtech companies: we read the existing research, identify the tier the design supports, flag what would disqualify it under specific state rubrics, and recommend where to submit and in what order. Since 2020 we’ve had 25+ studies approved by Evidence for ESSA and supported state approval processes in Arizona, Michigan, Utah, Arkansas, and Wyoming.
What do I get from an LXD Research study review?
Four things you can use immediately:
  • An educator-friendly summary, co-branded with your product, with effect sizes translated into terms practitioners use.
  • Named researcher authorship on that summary, with credentials anyone can verify.
  • An LXD Research ESSA badge reflecting the tier your evidence supports, available as soon as the review is complete.
  • Publication to our ResearchGate profile and submission to ERIC, so the work is discoverable outside your own website.
Can I use an evidence badge while waiting for Evidence for ESSA or state review?
Yes. Clearinghouse queues run months and state adoption cycles open annually, which leaves a long gap between finishing a study and holding a third-party designation. An LXD Research ESSA badge reflects the tier your evidence supports under an independent researcher’s assessment and is available as soon as the review is complete. It isn’t a substitute for an Evidence for ESSA rating or a state approval, and the distinction should be stated plainly wherever the badge appears. The badge criteria and definitions are published, so a district reviewer can check what a given badge means.
Can I upgrade an existing study to a higher ESSA tier?
Sometimes. If the partner district has outcome data for students who didn’t use the program, a matched comparison group may be constructible from records you already have, which can move a Tier IV or Tier III claim to Tier II. Random assignment can’t be added retroactively, so Tier I always requires a new prospective study. The determining factor is usually whether the comparison data exists and whether the district will share it.
Is Tier IV worth documenting if I’m planning a real study anyway?
Generally yes. A well-built logic model is the analytic spine of the study that follows — it forces you to specify which outcomes the program should move, for which students, and through what mechanism, which are exactly the decisions a study design has to make. It also gives sales and marketing something defensible to point to during the two or three years an efficacy study takes to produce results. The statutory definition also expects an evaluation to be underway, so pairing the two is the intended pattern rather than a workaround.

Not Sure What You Have? Send It to Us.

LXD Research reviews existing studies, identifies which tier the design supports, flags what would disqualify it under specific state rubrics, and maps the shortest path to a higher one — whether that’s a reanalysis of data you already have, a look-back study, or a prospective design built to clear review from the outset.

Schedule a Free Consultation Get Your Evidence Scorecard