LittleLearner Research Team
Studying how language models acquire knowledge by training them on a controlled U.S. elementary-school curriculum (K-5) to distinguish between skill acquisition and elicitation.
At a Glance
- AI Researchers
- Educational Scientists
- Cognitive Scientists
AI Tools by LittleLearner Research Team
(1)LittleLearner
LM Trained on K5 Curriculum
Discussions
No discussions yet
Be the first to start a discussion about LittleLearner Research Team
Latest News
MPI and ETH release LittleLearner: Language Models Under Pedagogically-Controlled Knowledge Exposure
Release of LittleCurriculum: An 88B-token K-5 educational corpus
LittleLearner model weights made available on Hugging Face
Products & Services
A series of language models (0.6B, 1.3B, 5B parameters) trained from scratch on a pedagogically-controlled curriculum (K-5).
An 88B-token corpus distilled from FineWeb-Edu, filtered to U.S. elementary-school standards (K-5) to study model knowledge boundaries.
Math-specialized variants of LittleLearner post-trained on MathCAMPS using Group Relative Policy Optimization.
Variants of LittleLearner fine-tuned for general conversational behavior while maintaining the K-5 knowledge boundary.
Market Position
Differentiated by 'Controlled Knowledge Exposure,' offering a unique experimental setup compared to 'black-box' models trained on massive, unfiltered datasets.
Leadership
Founders
Wieland Brendel
Independent Group Leader at the Max Planck Institute for Intelligent Systems (MPI-IS) and PI at the ELLIS Institute Tübingen. Previously a researcher focusing on robust machine learning.
Ryan Cotterell
Tenure-track Assistant Professor of Computer Science at ETH Zürich. Renowned for work in computational linguistics and natural language processing.
Executive Team
Fanfei Li
Lead Researcher
PhD student at MPI-IS. Former Data Scientist at TikTok; graduate of UCLA Anderson.
Wieland Brendel
Principal Investigator
Group Leader at MPI-IS / PI at ELLIS Institute Tübingen.
Board of Directors
Founding Story
Modern LMs are trained on such vast datasets that it's impossible to tell if a skill was learned during training or simply elicited. The LittleLearner team created a controlled sandbox to establish clean experimental boundaries for knowledge attribution.
Business Model
Revenue Model
Academic research project / Open-source software. No commercial revenue model reported.
Pricing Tiers
Models and datasets are available for academic research purposes under open-source licenses.
Target Markets
- AI Researchers
- Educational Scientists
- Cognitive Scientists
- RL & Discovery research
- Continual learning studies
- Educational science and human-model comparison
- Interpretability research
- Academic AI community