ASSIST 2009-10 is one of the most widely used benchmark datasets in knowledge tracing and educational data mining and has supported generations of models for predicting student knowledge and performance. It has been used in more than 200 papers and remains a common benchmark for comparing knowledge tracing models.
About the dataset: This dataset contains student interaction data collected through ASSISTments during the 2009-2010 school year. The original release separates Skill Builder, or mastery-learning, interactions from non-Skill Builder data and includes information such as student responses, attempts, skills, and problem-level performance.
Download: The dataset can be downloaded from the ASSISTments 2009-10 data repository. Download ASSIST 2009-10
Citation: When publishing research using this dataset, please cite:
Feng, M., Heffernan, N. T., & Koedinger, K. R. (2009). Addressing the assessment challenge in an Intelligent Tutoring System that tutors as it assesses. User Modeling and User-Adapted Interaction, 19, 243–266.
More information: Additional information about the dataset is available on the ASSISTments data website. Learn more about ASSIST 2009-10
FoundationalASSIST extends traditional educational datasets beyond problem IDs and binary correctness by providing the actual natural-language problem body and students’ responses. Its combination of large-scale learning trajectories and rich textual information makes it especially useful for knowledge tracing, student modeling, and research with large language models.
About the dataset: FoundationalASSIST contains approximately 1.7 million interactions from 5,000 students working on Illustrative Mathematics problems for Grades 6-8. It includes problem text, students’ first submitted answers, skill information, and Common Core-aligned metadata, allowing researchers to study not only whether students were correct but also what they actually answered.
Download: FoundationalASSIST is available through Hugging Face under a CC BY-NC 4.0 license and requires researchers to agree to the dataset’s responsible-use conditions before accessing the data. Download FoundationalASSIST
Citation: When publishing with this dataset, please cite: Worden, E., Heffernan, C., Heffernan, N., & Sonkar, S. (2026). FoundationalASSIST: Dataset for Foundational Knowledge Tracing & Pedagogical Grounding of Large Language Models. In International Conference on Artificial Intelligence in Education (pp. 489-498).
More information: Additional information about the structure, research applications, and benchmarking experiments associated with FoundationalASSIST is available through the ASSISTments project page and dataset documentation. Learn more about FoundationalASSIST
This dataset provides an opportunity to study educational technology at scale by connecting student learning outcomes, ASSISTments usage, teacher practices, and professional development participation across schools in 24 U.S. states.
About the dataset: The dataset comes from a large-scale study of middle school mathematics teachers and students participating in ASSISTments alongside a virtual professional learning community. Data were collected across two cohorts during the 2022-23 and 2023-24 school years and include study outcomes, implementation measures, and detailed ASSISTments usage logs.
Download: The study dataset, problem logs, codebook, and supporting files are publicly available through OpenICPSR. Download Scaling Teachers’ Professional Development for ASSISTments
Citation: When publishing with this dataset, please cite: Feng, M., Li, L., Huang, C., Brezack, N., & Luttgen, K. (2025). Empowering teachers with technology: a national study on a formative assessment platform. In International Conference on Artificial Intelligence in Education (pp. 119-126).
More information: More details about the study design, research questions, sample, intervention, and available data files can be found on the OpenICPSR project page. Learn more about the study
MathNet57789 provides a large-scale collection of authentic student handwritten mathematics work, opening opportunities for research on multimodal learning analytics, handwriting recognition, automated scoring, and vision-language models.
About the dataset: MathNet57789 contains 57,789 handwritten images from 21,212 students responding to 192 mathematics problems in ASSISTments, spanning Grade 2 through high school and collected between 2019 and 2023. The dataset also includes problem information, curriculum and skill metadata, and teacher-assigned scores, with student images processed to remove personally identifiable information.
Download: MathNet57789 is available through Hugging Face under a CC BY-NC 4.0 license and requires users to agree to the responsible-use guidelines before accessing the dataset. Download MathNet57789
Citation: When publishing with this dataset, please cite: Baral, S., Lucy, L., Knight, R., Ng, A., Soldaini, L., Heffernan, N., & Lo, K. (2025). DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students’ Hand-Drawn Math Images. In Proceedings of NAACL 2025, 6902–6920.
More information: Additional documentation, including the data dictionary, dataset statistics, and responsible-use requirements, is available on the MathNet57789 dataset page. Learn more about MathNet57789
DrawEduMath combines authentic student handwritten mathematics with unusually rich teacher-generated annotations, making it a powerful benchmark for evaluating whether vision-language models can meaningfully interpret students’ mathematical work.
About the dataset: DrawEduMath contains 2,030 images of students’ handwritten responses to 188 mathematics problems, along with teacher-written descriptions and 11,661 teacher-written question-answer pairs.
Download: DrawEduMath is available through Hugging Face under a CC BY-NC 4.0 license for research and educational use. Download DrawEduMath
Citation: When publishing with this dataset, please cite: Baral, S., Lucy, L., Knight, R., Ng, A., Soldaini, L., Heffernan, N., & Lo, K. (2025). DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students’ Hand-Drawn Math Images. In Proceedings of NAACL 2025, 6902–6920.
More information: The DrawEduMath project website contains the paper, dataset information, benchmark results, examples, code, and leaderboard for evaluating vision-language models on student mathematical work. Learn more about DrawEduMath