Data and Competitions

Eedi has spent over a decade collecting the novel and granular data needed to diagnose student misconceptions. We are committed to open-sourcing many of these data sets and sharing them with our colleagues in the Learning Engineering field.


Data

By open sourcing these data sets, Eedi is providing the learning community with access to diagnostic infrastructure to support a new and flourishing ecosystem of rigorous and coherent learning experiences.

Our 7 individual sets of Eedi’s misconception data are regularly downloaded and used in research, and in competitions. Download the data that will support your research below.

Eedi - Mining Misconceptions in Mathematics [Competition 🏆]
In this competition hosted by The Learning Agency, teams developed an NLP model driven by ML to accurately predict the affinity between misconceptions and incorrect answers (distractors) in multiple-choice questions. This solution will suggest candidate misconceptions for distractors, making it easier for expert human teachers to tag distractors with misconceptions.

‍MAP - Charting Student Math Misunderstandings [Competition 🏆]
In this competition hosted by The Learning Agency, teams developed an NLP model driven by ML to accurately predict students’ potential math misconceptions based on student explanations in open-ended responses. This solution will suggest candidate misconceptions for these explanations, making it easier for teachers to identify and address students’ incorrect thinking, which is critical to improving student math learning.

‍Question-Anchored-Tutoring-Dialogues-2k [Public 🆓]
This dataset contains dialogues from math tutoring interventions recorded on Eedi alongside information about the questions they relate to.

‍Trace the Ace [Competition 🏆]This competition focuses on evaluating tutoring effectiveness using only the student–tutor conversation. Participants built models that use tutoring session transcripts to predict whether a student goes on to answer a follow-up question correctly - a practical proxy for whether learning actually took place. Developing better ways of identifying effective tutoring could inform how we train educators, guide real-time support, and expand access to high-quality tutoring through AI-powered tools.

Misconception Graph data (offered in partnership with Learning Commons) [Public 🆓]
The Eedi Misconceptions Graph dataset represents common math misconceptions and the
constructs they get in the way of. A Misconception is a flawed conceptual structure, or a gap in conceptual understanding, that shows up as a systematic, predictable error pattern across problems involving the same mathematical concept. A Construct is granular learning goals that target a specific element of mathematics (e.g. Order fractions with the same denominators)

‍NeurIPS 2022 CausalML Challenge: Causal Insights for Learning Paths in Education [Competition 🏆]
Our second competition with Microsoft Research, building on the first competition to consider causal challenges using time-series data. The tasks were to identify the causal relationships between different constructs and to predict the impact of learning one construct on the ability to answer questions on other constructs. A/B tests were conducted on Eedi to collect this novel dataset.

‍The NeurIPS 2020 Education Challenge (CodaLab) [Competition 🏆]
Our first competition, which we launched in collaboration with Microsoft Research. There were four tasks covering the prediction of students’ responses to future items based on past responses, the estimation of question quality, and the selection of the most informative questions. The dataset includes student responses to multiple-choice questions, and question, student and answer metadata.

We aim to measurably improve learning outcomes for a billion students by 2030.