Assessment is one of those words in education that can mean almost anything — a quick exit ticket at the end of class, a state standardized test, a reading screener given three times a year, an algorithm-driven platform that adjusts difficulty in real time. They’re all assessments, but they serve very different purposes. Treating them as interchangeable is one of the most common ways schools end up with data they can’t use.
This article maps the main assessment types educators encounter — formative and summative, criterion-referenced and norm-referenced, and the adaptive tools that cut across both — and looks at what each one is actually built to tell you. A companion piece, What Independent Validation Actually Shows, examines four specific tools and the evidence behind them.
What Is the Purpose of Assessment in Education?
At their core, assessments exist to answer questions. What does this student know right now? Are they on track to meet grade-level expectations by end of year? How does their skill level compare to peers nationally? Has the instructional approach worked? The type of assessment you choose determines which of those questions you can actually answer — and choosing the wrong tool for the question leads to misinterpretation, not insight.
Assessments guide instruction, support early intervention, inform school-level decisions, and — for better or worse — shape how schools are held accountable. When thoughtfully applied, they create the feedback loop that lets educators respond to what’s actually happening in their classrooms rather than what they assume is happening. The challenge is that “thoughtfully applied” requires understanding what each type of assessment was designed to do.
A working principle: No single assessment tells the full story of a student’s learning. The schools that use assessment data most effectively tend to combine multiple types — each chosen because it answers a distinct question — rather than relying on one platform or one annual test to do everything.
What Are the Main Types of Assessments — and How Do They Differ?
It helps to think about assessments along two axes: when they happen in the learning process, and what they compare performance against. The first axis gives us formative versus summative. The second gives us criterion-referenced versus norm-referenced. Adaptive assessments sit across both, adjusting the experience based on how a student responds in the moment.
Most educators encounter all of these types in the course of a school year — sometimes without realizing that the same platform is doing double duty as both a formative progress monitor and a norm-referenced diagnostic. Understanding the distinctions doesn’t require a research background. It requires knowing what question you’re trying to answer.
Formative vs. Summative: A Question of Timing
Conducted throughout the learning process to provide real-time feedback. The goal is to adjust instruction while there’s still time to adjust. Low-stakes, often informal, and designed to inform the next step — not to render a final judgment.
Given at the end of a unit, semester, or year to measure whether students have met learning objectives. Higher stakes. Used to evaluate achievement, hold schools accountable, and make decisions about placement or program effectiveness.
Criterion-Referenced vs. Norm-Referenced: A Question of Comparison
Measures performance against a fixed set of learning objectives or grade-level standards. The question is: has this student mastered this skill? Results are expressed as proficiency levels — Below, Approaching, Meets, or Exceeds expectations.
Compares a student’s performance to a broader population. The question is: how does this student rank relative to peers? Results are expressed as percentile ranks or standard scores, which communicate relative standing rather than mastery.
Most assessment tools built for literacy and math intervention are primarily criterion-referenced by design. They measure whether a student has acquired a specific skill — which phonics patterns are solid, which comprehension strategies need work — and flag students who haven’t met expected benchmarks. Results tie back to a scope and sequence, so a score tells a teacher not just that a student is struggling, but which skills to address and in what order.
Norm-referenced results — the percentile ranks and grade equivalents that come from nationally normed assessments — are useful for understanding how a school or district compares to national trends and for flagging students who may need further evaluation. But for the day-to-day work of instruction and intervention, criterion-referenced data is usually more directly useful. Knowing that a student scores at the 22nd percentile tells you less about what to do on Monday morning than knowing they haven’t yet mastered consonant blends.
Where Do Adaptive Assessments Fit?
Adaptive assessment is not a third category alongside criterion-referenced and norm-referenced — it’s a method of delivery that can serve either one. An adaptive test adjusts item difficulty in response to how the student is doing: answer correctly and the next item gets harder, miss one and it eases back. The test converges on the student’s actual level rather than marching every student through an identical set of questions. What that buys, and what it costs:
- Shorter tests. The assessment stops spending items on material a student has clearly mastered or clearly hasn’t reached, so it can reach a reliable estimate in less time.
- No floor or ceiling. Fixed-form tests become uninformative for students far above or below grade level. An adaptive test keeps measuring in both directions.
- A large item bank is required. Calibrating enough items to support adaptive routing is a serious engineering investment, and not every tool marketed as adaptive is doing the same thing underneath.
- Less transparency for teachers. Because students see different items, a teacher can’t look at the test and see what a student missed. The score report has to carry that information instead.
That last tradeoff is the one worth weighing carefully. Adaptive delivery shifts the burden onto the reporting layer — which makes how clearly a tool communicates skill-level results just as important as its psychometrics.
What About the Platforms That Organize Assessment Data?
There’s a category of product that doesn’t administer assessments at all but has become central to how schools use them. These platforms pull results from multiple sources — a reading screener, a math benchmark, state test scores, attendance and behavior records — into a single view, and wrap workflow around it. Tools in this space include Panorama Education, Branching Minds, Renaissance’s eduCLIMBER, and the analytics products within PowerSchool, among others. What they typically offer:
- Aggregation across instruments. One dashboard showing a student’s results from assessments that otherwise live in separate systems and don’t speak to each other.
- MTSS workflow. Documenting which students are in which tier, what intervention they’re receiving, who is responsible, and what happened at the last team meeting.
- Early warning indicators. Composite flags that combine academic, attendance, and behavior signals to surface students before a single measure would.
- Reporting for adults who aren’t in the classroom. Building- and district-level views for principals, coaches, and central office staff.
The important thing to understand about this category is that these platforms generally don’t generate assessment data — they organize it. That has a direct consequence for evidence: a data platform inherits the validity of whatever feeds it. Aggregating three weak measures into an attractive dashboard does not produce a strong one, and a composite risk flag is only as trustworthy as its inputs. If a screener’s cut scores were never validated, routing them through a nicer interface doesn’t change that.
It also means the evidence question is a different one. For an assessment, the question is whether the instrument measures what it claims to measure. For a data platform, the claim is about decisions and workflow: that teams identify students earlier, meet more efficiently, or make better placement calls. That’s a harder thing to study, and it’s studied far less often. If you’re evaluating a tool in this category, these are the questions worth asking:
- Which specific decision does this platform change, and who makes that decision today without it?
- Does it surface a recommendation, or does it display data and leave the interpretation to the team?
- Whose assessments does it ingest, and does it preserve those instruments’ own cut scores and proficiency definitions — or apply its own?
- Is there any independent evidence that teams using it identify students earlier or more accurately than teams that don’t?
A note on our own limits: LXD Research has not conducted independent studies of any assessment data management or MTSS platform. The companies named above are offered as examples of the category, not as recommendations, and their inclusion here is not an endorsement. We think the category is genuinely useful and genuinely under-researched — which is a fair description of where the evidence base stands rather than a criticism of any particular product.
Are Assessments Too Long — or Just the Wrong Shape?
One tension worth naming directly: as diagnostic tools have become more sophisticated, many assessments have grown longer — because teachers genuinely want to know where to intervene, and more granular data seems to promise better answers. The problem is that the students who don’t need intervention end up sitting through the same lengthy battery as those who do, simply so that everyone has a complete data profile. That’s a real cost, in instructional time and in student experience.
But length is only the most visible complaint, and probably not the most serious one. Three distinct problems get bundled together under “our assessments take too long”:
- Duration. A screener that has quietly become a diagnostic battery, administered to everyone, including the majority of students it was never going to flag.
- Shelf life. A benchmark given three times a year describes a student as of the testing window. Two weeks later, instruction has moved and the picture is stale. The data becomes a photograph being used as a map.
- Grain size. Results arriving at a resolution nobody can act on — either too coarse (“below benchmark in reading” names no skill) or too fine, where a sub-skill measured with four items produces a number too noisy to trust.
Grain size is the subtlest of the three and the one most often gotten wrong in both directions. The useful resolution is the level at which a teacher can do something different tomorrow: specific enough to name an instructional target, stable enough that the number means something. That’s a design decision, not an accident, and it’s worth asking a vendor directly what grain size their reporting was built for.
It’s also worth granting the incumbents their reasons. Established assessments are long partly because measuring a broad construct with defensible reliability genuinely takes items. A short, sharply targeted instrument trades coverage for actionability — a legitimate trade, but one that has to be made deliberately and then demonstrated, not just asserted. Screeners are supposed to be brief by design, sorting students into “likely fine” and “look more closely.” The students who need a closer look should get one. The students who don’t have already told you what you need to know.
A trusted resource: The Florida Center for Reading Research maintains a well-organized overview of assessment types and their appropriate uses — a useful reference for educators navigating the landscape without a research background.
What Are the Persistent Challenges in Assessment — With or Without Technology?
Technology has solved some assessment problems and created new ones. Assessment fatigue — students sitting through multiple screeners, benchmarks, and unit checks across different platforms — is a real risk in districts that have added edtech tools on top of existing assessment calendars rather than replacing them. More data is not always better data, particularly when teachers lack the time or support to act on what they’re seeing.
The tools that tend to work best in practice are those designed with a clear question in mind and a clear audience for the answer. When the question is specific and the data lands with someone who knows what to do with it, even a brief screener can meaningfully shift instructional decisions. When neither of those conditions holds, even a sophisticated platform generates noise rather than signal.
Interpretation remains the most consistent gap. Whether a platform is validated at ESSA Tier I or Tier IV, results only matter if the person looking at them knows what instructional decision to make next. Professional development focused on data interpretation — not just platform training — is one of the highest-leverage investments a school or district can make alongside any assessment adoption.
What Does a Thoughtful Assessment Strategy Actually Look Like?
Effective assessment practice isn’t about maximizing the number of tools in use — it’s about making sure each tool is doing a distinct job and that the data flows toward decisions. A reasonable framework starts with asking what question each assessment is meant to answer, and the table below lays out how the common types divide that work.
| Type | When | Who Uses the Data | The Question It Answers |
|---|---|---|---|
| Universal Screener | Two or three times a year; all students | Teacher / intervention team | Who may be at risk and needs a closer look? |
| Diagnostic | After a screener flags a student | Teacher / specialist | Which specific skills are missing? |
| Progress Monitor | Every one to three weeks; intervention students only | Intervention teacher | Is this intervention working for this student? |
| Unit Assessment | End of an instructional unit | Teacher | Did this unit produce the mastery it aimed for? |
| Summative / State | End of year; all students | Administrators, state, families | Did students meet grade-level standards? |
The question worth asking of any assessment tool — validated or not — is simple: what decision will this data inform, and who will make it? If the answer is clear and the person making the decision has what they need to act, the tool is earning its place. If the answer is vague, it probably isn’t.
Which tools hold up? Understanding the categories is the first half of the problem; knowing whether a particular tool does what it claims is the second. Our companion article walks through the review bodies that evaluate assessments — and four tools LXD Research has examined — in what independent validation actually shows.
Frequently Asked Questions
What are the main types of assessments in education?
What is the difference between formative and summative assessment?
What is the difference between criterion-referenced and norm-referenced assessment?
Are adaptive assessments criterion-referenced or norm-referenced?
What is a universal screener and how often should it be given?
What software organizes assessment data for MTSS?
Why are so many new assessment tools entering the market?
How many assessments are too many?
Want to Know If Your Assessment Data Is Actually Working for You?
LXD Research helps edtech companies design and evaluate assessments with independent rigor — and helps districts make sense of the data they’re already collecting. Start a conversation with our team.
Schedule a Free Consultation View Our Services