As AI becomes increasingly common in education spaces, so too are the questions about its effectiveness. Educators, researchers, families, and school leaders all want to understand whether AI tools are truly improving teaching and learning, yet we are presented with a unique challenge: the technology is evolving much faster than the systems traditionally used to study it.
These questions were at the center of a recent EdTech Insiders webinar, hosted by Ben Kornell, featuring Rebecca Winthrop (Director, Center for Universal Education at Brookings)1, Libby Hills (CEO & Co-founder, EdTechnical)2, Meghan McCormick (Senior Research & Impact Officer, Overdeck)3, Will Thornton (EdTech GTM, ElevenLabs), Bibi Groot (Chief Impact Officer, Eedi)4, and John Whitmer, Ed.D. (Founder & Senior Research Scientist, Learning Data Insights)5.
Evidence Has to Move as Fast as the Technology It’s Evaluating
Traditional efficacy research was built for a world where a product’s core design held steady for a year or more at a time — long enough for a study to enroll schools, collect a full year of outcome data, and publish before the tool itself moved on. AI products don’t hold still that way. A model update, a new guardrail, or a redesigned prompt can change what a tool actually does in a classroom within a single semester.
That doesn’t mean the standard for evidence should drop — it means the methods for gathering it need to modernize. The panel’s answer wasn’t to trust marketing claims in the meantime; it was to develop research approaches, like short, structured sprints, that can generate meaningful evidence on a timeline that matches how the technology actually ships.
Can Educational Research Move as Fast as AI Product Cycles?
One of the biggest themes in the webinar was the gap between the speed of AI development and the pace of educational research. Traditional research studies can take years to complete, while AI tools may change significantly in just a few months. Rather than abandoning rigorous research, participants discussed the need for additional approaches that can generate useful evidence more quickly. Ideas such as research sprints may help researchers gather timely information while still maintaining meaningful standards.
The challenge is no longer simply determining whether a tool works — it is figuring out how to evaluate tools that are rapidly evolving and improving. That’s the same question we walk through in What ESSA Level Is My Research?, where the right evidence standard depends on how mature a product is, not just how promising its early results look. Libby Hills framed this well from her vantage point as an investor: evidence is a journey rather than a single destination, and while not every company needs a randomized controlled trial on day one, every company does need to be building toward one. EdTechnical’s research sprints model is her attempt to match research pace to product pace — the same instinct behind our own view that an AI tool six months into development and one with three years of classroom deployment behind it aren’t asking for the same kind of proof.
The practical version of this problem: by the time a two-year efficacy study on an AI tool is published, the version of the product that was studied may no longer exist. Faster, staged evidence — usability data, then implementation data, then outcome data — lets a research portfolio grow alongside the product instead of trailing years behind it.
What Actually Counts as Evidence That an AI Tool Works?
The conversation also emphasized that not all data is equally meaningful. Metrics such as user satisfaction or product popularity may be useful for product marketing strategies, but they do not necessarily tell us whether students are learning more or if teachers are being better supported.
- Weaker signals: user satisfaction scores, download or sign-up counts, anecdotal teacher enthusiasm, engagement metrics alone.
- Stronger signals: measures of student achievement, teacher effectiveness, and equitable access to learning opportunities — evidence rooted in the educational science of how kids learn and how teachers teach.
Meghan McCormick put a fine point on this from the funder’s side: the Overdeck Family Foundation isn’t looking for “teachers love it” testimonials when it evaluates a grant. It’s looking for evidence that a product is grounded in learning science and that it delivers for the students who are furthest behind, not just for the average student. That distinction — average impact versus impact where it’s needed most — is easy to lose in an evidence summary that only reports overall effects.
Another important takeaway was that evidence should be viewed as a journey rather than a final destination; different types of evidence may be appropriate at different stages of a product’s development. We’ve made this same point in You Already Have Evidence. It’s Not Telling Your Story, and it applies with extra force to AI tools that are still finding their footing in real classrooms: early-stage implementation data and later-stage outcome data are both legitimate evidence, as long as each is labeled honestly for what it is.
What Are AI’s Real Promises and Risks in the Classroom?
The discussion highlighted several areas where AI shows promise, alongside real concerns about how students use these tools outside of structured, supervised settings.
Where AI Shows Promise
- Personalized learning pathways
- Expanded access to tutoring, especially for underserved children
- Reduced administrative burden, freeing teacher time for instruction and student relationships
Where the Risks Sit
- Unscaffolded, open-ended chatbot use outside school
- Reduced opportunities for independent thinking and productive struggle
- Unclear effects on deeper learning and social-emotional development
Rebecca Winthrop pointed to Brookings’ own research here: a recent Center for Universal Education report found that wide, unscaffolded student access to AI is undermining cognitive development. She drew a direct comparison to social media — only this time, she argued, what’s at risk is children’s individual capacity to think independently.
John Whitmer of Learning Data Insights offered a way to hold both sides of that split at once: AI deployed in constrained, pedagogically grounded ways can genuinely support learning, while the opposite — open-ended chatbot use — risks eroding the very thing education is supposed to build, a student’s capacity to sit with difficulty, take feedback, and learn from being wrong. That distinction, constrained and pedagogically grounded versus open-ended, may be a more useful lens than treating “AI” as a single category, since the two design choices point toward opposite outcomes.
The speakers recognized that when developing AI edtech tools, it is crucial that we put students’ learning and development first — finding ways to support them that will work, at scale, both in and out of school.
So What Should Guide the Next Phase of AI in Schools?
One of the strongest messages throughout the webinar was that technology alone is not enough. Learning will not magically transform because of AI, but AI can change what opportunities we have for supporting the learning process and building products for what works. The success of AI in education will depend on how thoughtfully it is designed, implemented, and evaluated.
While the discussion answered many tough questions about AI in education, it also sparked many more that remain unanswered. There was broad agreement, though, that the future of AI in education should be guided by evidence, educational research, and a focus on student learning. As AI continues to evolve, the goal should not simply be to rapidly produce and adopt new technology, but to ensure that it meaningfully supports students and educators.
LXD Research specializes in designing and conducting research studies on edtech programs, including AI-enabled tools, with a focus on providing meaningful evidence that helps educators know how programs work and under what contexts. Browse our published studies to see that evidence in practice.
- Rebecca Winthrop directs the Center for Universal Education at the Brookings Institution. See the full EdTech Insiders webinar recording for her remarks on the Center’s AI-in-education research.
- Libby Hills is CEO & Co-founder of EdTechnical, a research-grounded AI-in-education think tank and investor, home of the research sprints model referenced above.
- Meghan McCormick is Senior Research & Impact Officer at the Overdeck Family Foundation.
- Bibi Groot is Chief Impact Officer at Eedi, an education impact lab focused on AI math tutoring.
- John Whitmer, Ed.D., is Founder & Senior Research Scientist at Learning Data Insights and co-chairs the AI in Measurement Subcommittee of the National Council of Measurement in Education.
Not Sure What Level of Evidence Your Product Has?
Get a research-path recommendation for your product, timeline, and procurement goals — whether it’s an AI tool or a traditional edtech program.
Book an ESSA Evidence Review Browse Published Studies