A Bandit-Based Approach to Educational Recommender Systems: Contextual Thompson Sampling for Learner Skill Gain Optimization

Lukas De Kerpel et al.

INFORMS Transactions on Education2026https://doi.org/10.1287/ited.2025.0174article
AJG 2ABDC C
Weight
0.50

What the paper says

In recent years, instructional practices in operations research, management science, and analytics have increasingly shifted toward digital environments, where large and diverse groups of learners make it difficult to provide practice that adapts to individual needs. This paper introduces a method that generates personalized sequences of exercises by selecting, at each step, the exercise most likely to advance a learner’s understanding of a targeted skill. The method uses information about the learner and their past performance to guide these choices, and learning progress is measured as the change in estimated skill level before and after each exercise. Using data from an online mathematics tutoring platform, we find that the approach recommends exercises associated with greater skill improvement and adapts effectively to differences across learners. From an instructional perspective, the framework enables personalized practice at scale, highlights exercises with consistently strong learning value, and helps instructors identify learners who may benefit from additional support.

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.1287/ited.2025.0174

Or copy a formatted citation

@article{lukas2026,
  title        = {{A Bandit-Based Approach to Educational Recommender Systems: Contextual Thompson Sampling for Learner Skill Gain Optimization}},
  author       = {Lukas De Kerpel et al.},
  journal      = {INFORMS Transactions on Education},
  year         = {2026},
  doi          = {https://doi.org/https://doi.org/10.1287/ited.2025.0174},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

A Bandit-Based Approach to Educational Recommender Systems: Contextual Thompson Sampling for Learner Skill Gain Optimization

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.50

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.50 × 0.4 = 0.20
M · momentum0.50 × 0.15 = 0.07
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.