Skip to content
Home » Types of Assessments in Education and Edtech

Types of Assessments in Education and Edtech

Assessment is one of those words in education that can mean almost anything — a quick exit ticket at the end of class, a state standardized test, a reading screener given three times a year, an algorithm-driven platform that adjusts difficulty in real time. They’re all assessments, but they serve very different purposes. Treating them as interchangeable is one of the most common ways schools end up with data they can’t use.

This article maps the main assessment types educators encounter — formative and summative, criterion-referenced and norm-referenced, and the adaptive tools that cut across both — and looks at what each one is actually built to tell you. A companion piece, What Independent Validation Actually Shows, examines four specific tools and the evidence behind them.

Short answer: Assessments divide along two axes — when they happen (formative during learning, summative after) and what they compare against (criterion-referenced against a standard, norm-referenced against peers). Adaptive testing is a delivery method that can serve either. In practice these combine into five working roles: universal screener, diagnostic, progress monitor, unit assessment, and end-of-year summative. The right question isn’t which type is best, but which question you need answered and who will act on the answer.

What Is the Purpose of Assessment in Education?

At their core, assessments exist to answer questions. What does this student know right now? Are they on track to meet grade-level expectations by end of year? How does their skill level compare to peers nationally? Has the instructional approach worked? The type of assessment you choose determines which of those questions you can actually answer — and choosing the wrong tool for the question leads to misinterpretation, not insight.

Assessments guide instruction, support early intervention, inform school-level decisions, and — for better or worse — shape how schools are held accountable. When thoughtfully applied, they create the feedback loop that lets educators respond to what’s actually happening in their classrooms rather than what they assume is happening. The challenge is that “thoughtfully applied” requires understanding what each type of assessment was designed to do.

A working principle: No single assessment tells the full story of a student’s learning. The schools that use assessment data most effectively tend to combine multiple types — each chosen because it answers a distinct question — rather than relying on one platform or one annual test to do everything.

What Are the Main Types of Assessments — and How Do They Differ?

It helps to think about assessments along two axes: when they happen in the learning process, and what they compare performance against. The first axis gives us formative versus summative. The second gives us criterion-referenced versus norm-referenced. Adaptive assessments sit across both, adjusting the experience based on how a student responds in the moment.

Most educators encounter all of these types in the course of a school year — sometimes without realizing that the same platform is doing double duty as both a formative progress monitor and a norm-referenced diagnostic. Understanding the distinctions doesn’t require a research background. It requires knowing what question you’re trying to answer.

Formative vs. Summative: A Question of Timing

During Learning
Formative Assessment

Conducted throughout the learning process to provide real-time feedback. The goal is to adjust instruction while there’s still time to adjust. Low-stakes, often informal, and designed to inform the next step — not to render a final judgment.

After Learning
Summative Assessment

Given at the end of a unit, semester, or year to measure whether students have met learning objectives. Higher stakes. Used to evaluate achievement, hold schools accountable, and make decisions about placement or program effectiveness.

Criterion-Referenced vs. Norm-Referenced: A Question of Comparison

Against a Standard
Criterion-Referenced

Measures performance against a fixed set of learning objectives or grade-level standards. The question is: has this student mastered this skill? Results are expressed as proficiency levels — Below, Approaching, Meets, or Exceeds expectations.

Against a Peer Group
Norm-Referenced

Compares a student’s performance to a broader population. The question is: how does this student rank relative to peers? Results are expressed as percentile ranks or standard scores, which communicate relative standing rather than mastery.

Most assessment tools built for literacy and math intervention are primarily criterion-referenced by design. They measure whether a student has acquired a specific skill — which phonics patterns are solid, which comprehension strategies need work — and flag students who haven’t met expected benchmarks. Results tie back to a scope and sequence, so a score tells a teacher not just that a student is struggling, but which skills to address and in what order.

Norm-referenced results — the percentile ranks and grade equivalents that come from nationally normed assessments — are useful for understanding how a school or district compares to national trends and for flagging students who may need further evaluation. But for the day-to-day work of instruction and intervention, criterion-referenced data is usually more directly useful. Knowing that a student scores at the 22nd percentile tells you less about what to do on Monday morning than knowing they haven’t yet mastered consonant blends.

Where Do Adaptive Assessments Fit?

Adaptive assessment is not a third category alongside criterion-referenced and norm-referenced — it’s a method of delivery that can serve either one. An adaptive test adjusts item difficulty in response to how the student is doing: answer correctly and the next item gets harder, miss one and it eases back. The test converges on the student’s actual level rather than marching every student through an identical set of questions. What that buys, and what it costs:

  • Shorter tests. The assessment stops spending items on material a student has clearly mastered or clearly hasn’t reached, so it can reach a reliable estimate in less time.
  • No floor or ceiling. Fixed-form tests become uninformative for students far above or below grade level. An adaptive test keeps measuring in both directions.
  • A large item bank is required. Calibrating enough items to support adaptive routing is a serious engineering investment, and not every tool marketed as adaptive is doing the same thing underneath.
  • Less transparency for teachers. Because students see different items, a teacher can’t look at the test and see what a student missed. The score report has to carry that information instead.

That last tradeoff is the one worth weighing carefully. Adaptive delivery shifts the burden onto the reporting layer — which makes how clearly a tool communicates skill-level results just as important as its psychometrics.

What About the Platforms That Organize Assessment Data?

There’s a category of product that doesn’t administer assessments at all but has become central to how schools use them. These platforms pull results from multiple sources — a reading screener, a math benchmark, state test scores, attendance and behavior records — into a single view, and wrap workflow around it. Tools in this space include Panorama Education, Branching Minds, Renaissance’s eduCLIMBER, and the analytics products within PowerSchool, among others. What they typically offer:

  • Aggregation across instruments. One dashboard showing a student’s results from assessments that otherwise live in separate systems and don’t speak to each other.
  • MTSS workflow. Documenting which students are in which tier, what intervention they’re receiving, who is responsible, and what happened at the last team meeting.
  • Early warning indicators. Composite flags that combine academic, attendance, and behavior signals to surface students before a single measure would.
  • Reporting for adults who aren’t in the classroom. Building- and district-level views for principals, coaches, and central office staff.

The important thing to understand about this category is that these platforms generally don’t generate assessment data — they organize it. That has a direct consequence for evidence: a data platform inherits the validity of whatever feeds it. Aggregating three weak measures into an attractive dashboard does not produce a strong one, and a composite risk flag is only as trustworthy as its inputs. If a screener’s cut scores were never validated, routing them through a nicer interface doesn’t change that.

It also means the evidence question is a different one. For an assessment, the question is whether the instrument measures what it claims to measure. For a data platform, the claim is about decisions and workflow: that teams identify students earlier, meet more efficiently, or make better placement calls. That’s a harder thing to study, and it’s studied far less often. If you’re evaluating a tool in this category, these are the questions worth asking:

  • Which specific decision does this platform change, and who makes that decision today without it?
  • Does it surface a recommendation, or does it display data and leave the interpretation to the team?
  • Whose assessments does it ingest, and does it preserve those instruments’ own cut scores and proficiency definitions — or apply its own?
  • Is there any independent evidence that teams using it identify students earlier or more accurately than teams that don’t?

A note on our own limits: LXD Research has not conducted independent studies of any assessment data management or MTSS platform. The companies named above are offered as examples of the category, not as recommendations, and their inclusion here is not an endorsement. We think the category is genuinely useful and genuinely under-researched — which is a fair description of where the evidence base stands rather than a criticism of any particular product.

Are Assessments Too Long — or Just the Wrong Shape?

One tension worth naming directly: as diagnostic tools have become more sophisticated, many assessments have grown longer — because teachers genuinely want to know where to intervene, and more granular data seems to promise better answers. The problem is that the students who don’t need intervention end up sitting through the same lengthy battery as those who do, simply so that everyone has a complete data profile. That’s a real cost, in instructional time and in student experience.

But length is only the most visible complaint, and probably not the most serious one. Three distinct problems get bundled together under “our assessments take too long”:

  • Duration. A screener that has quietly become a diagnostic battery, administered to everyone, including the majority of students it was never going to flag.
  • Shelf life. A benchmark given three times a year describes a student as of the testing window. Two weeks later, instruction has moved and the picture is stale. The data becomes a photograph being used as a map.
  • Grain size. Results arriving at a resolution nobody can act on — either too coarse (“below benchmark in reading” names no skill) or too fine, where a sub-skill measured with four items produces a number too noisy to trust.

Grain size is the subtlest of the three and the one most often gotten wrong in both directions. The useful resolution is the level at which a teacher can do something different tomorrow: specific enough to name an instructional target, stable enough that the number means something. That’s a design decision, not an accident, and it’s worth asking a vendor directly what grain size their reporting was built for.

It’s also worth granting the incumbents their reasons. Established assessments are long partly because measuring a broad construct with defensible reliability genuinely takes items. A short, sharply targeted instrument trades coverage for actionability — a legitimate trade, but one that has to be made deliberately and then demonstrated, not just asserted. Screeners are supposed to be brief by design, sorting students into “likely fine” and “look more closely.” The students who need a closer look should get one. The students who don’t have already told you what you need to know.

A trusted resource: The Florida Center for Reading Research maintains a well-organized overview of assessment types and their appropriate uses — a useful reference for educators navigating the landscape without a research background.

What Are the Persistent Challenges in Assessment — With or Without Technology?

Technology has solved some assessment problems and created new ones. Assessment fatigue — students sitting through multiple screeners, benchmarks, and unit checks across different platforms — is a real risk in districts that have added edtech tools on top of existing assessment calendars rather than replacing them. More data is not always better data, particularly when teachers lack the time or support to act on what they’re seeing.

The tools that tend to work best in practice are those designed with a clear question in mind and a clear audience for the answer. When the question is specific and the data lands with someone who knows what to do with it, even a brief screener can meaningfully shift instructional decisions. When neither of those conditions holds, even a sophisticated platform generates noise rather than signal.

Interpretation remains the most consistent gap. Whether a platform is validated at ESSA Tier I or Tier IV, results only matter if the person looking at them knows what instructional decision to make next. Professional development focused on data interpretation — not just platform training — is one of the highest-leverage investments a school or district can make alongside any assessment adoption.

What Does a Thoughtful Assessment Strategy Actually Look Like?

Effective assessment practice isn’t about maximizing the number of tools in use — it’s about making sure each tool is doing a distinct job and that the data flows toward decisions. A reasonable framework starts with asking what question each assessment is meant to answer, and the table below lays out how the common types divide that work.

Type When Who Uses the Data The Question It Answers
Universal Screener Two or three times a year; all students Teacher / intervention team Who may be at risk and needs a closer look?
Diagnostic After a screener flags a student Teacher / specialist Which specific skills are missing?
Progress Monitor Every one to three weeks; intervention students only Intervention teacher Is this intervention working for this student?
Unit Assessment End of an instructional unit Teacher Did this unit produce the mastery it aimed for?
Summative / State End of year; all students Administrators, state, families Did students meet grade-level standards?

The question worth asking of any assessment tool — validated or not — is simple: what decision will this data inform, and who will make it? If the answer is clear and the person making the decision has what they need to act, the tool is earning its place. If the answer is vague, it probably isn’t.

Which tools hold up? Understanding the categories is the first half of the problem; knowing whether a particular tool does what it claims is the second. Our companion article walks through the review bodies that evaluate assessments — and four tools LXD Research has examined — in what independent validation actually shows.

Frequently Asked Questions

What are the main types of assessments in education?
Assessments are usually grouped two ways. By timing: formative assessment happens during learning to adjust instruction, while summative assessment happens after learning to measure whether objectives were met. By comparison: criterion-referenced assessment measures a student against a fixed standard, while norm-referenced assessment compares a student to a peer group. In everyday school use these combine into recognizable roles — universal screeners, diagnostics, progress monitors, unit assessments, and end-of-year state tests.
What is the difference between formative and summative assessment?
Timing and stakes. Formative assessment is given during instruction, is typically low-stakes, and exists to change what happens next. Summative assessment is given at the end of a unit, semester, or year, carries higher stakes, and exists to record what was achieved. The same instrument can serve either purpose depending on when it is given and what is done with the result.
What is the difference between criterion-referenced and norm-referenced assessment?
A criterion-referenced assessment asks whether a student has mastered a defined skill or standard, and reports proficiency levels such as Below, Approaching, Meets, or Exceeds. A norm-referenced assessment asks how a student compares to a broader population, and reports percentile ranks or standard scores. Criterion-referenced results are generally more useful for planning instruction; norm-referenced results are more useful for understanding relative standing.
Are adaptive assessments criterion-referenced or norm-referenced?
Either. Adaptive testing is a method of delivery, not a third category. An adaptive assessment adjusts item difficulty based on how the student is responding, which can produce a more precise estimate in less testing time. It can be scored against a standard or against a norm group depending on how it was designed.
What is a universal screener and how often should it be given?
A universal screener is a brief assessment given to all students, typically two or three times a year, to identify who may be at risk and needs a closer look. Because it is given to everyone, brevity matters: a screener that grows into a full diagnostic battery costs instructional time for the majority of students who did not need one. Students flagged by a screener are the ones who should receive fuller diagnostic assessment.
What software organizes assessment data for MTSS?
A category of platforms — including tools such as Panorama Education, Branching Minds, Renaissance’s eduCLIMBER, and PowerSchool’s analytics products — aggregates results from multiple assessments into one view to support MTSS decision-making, intervention tracking, and early warning indicators. These platforms generally do not create assessment data; they organize data produced by other instruments, which means they inherit the strengths and weaknesses of whatever feeds them. LXD Research has not conducted independent studies of any platform in this category, and listing them here is not an endorsement.
Why are so many new assessment tools entering the market?
Largely because of dissatisfaction with three features of established commercial tests: they take a long time to administer, their results go stale quickly relative to the pace of instruction, and they often report at a grain size that doesn’t map to an instructional decision. Newer entrants tend to compete by being shorter, more frequent, and more specific about skills. Whether a given tool delivers on that depends on evidence — a short instrument still has to demonstrate that its cut scores identify the right students.
How many assessments are too many?
There is no fixed number, but a practical test is whether each assessment in the calendar answers a question no other assessment answers, and whether someone acts on the result. Assessment fatigue usually comes from layering new tools on top of an existing calendar rather than replacing anything. If two instruments are answering the same question for the same students, one of them is costing instructional time without adding information.

Want to Know If Your Assessment Data Is Actually Working for You?

LXD Research helps edtech companies design and evaluate assessments with independent rigor — and helps districts make sense of the data they’re already collecting. Start a conversation with our team.

Schedule a Free Consultation View Our Services