$ cat wiki/entities/ai2.md
Ai2 (Allen Institute for AI)
Latest
- 2026-08-07
TutorMoments — a tutoring benchmark that scores holding back, not just helping
Overview
Ai2, the Allen Institute for AI, is a research institute that publishes open models, datasets and benchmarks. This page was created on 2026-08-09 on the release of TutorMoments, and its contents are limited to what that release and its coverage stated (source).
Ai2's founding, funding, headcount and its wider model lines were not established on the run that created this page — the sandbox could not fetch any primary page that day (seventh consecutive day of blocked egress), so nothing beyond the TutorMoments release is recorded here. This is a stub created because the institute is a recurring publisher of open evaluation work that this wiki had no page for, not because one benchmark warrants an entity.
Key People
unknown — no named individual appeared in any source read on the run that created
this page.
Models & Products
- TutorMoments — open, replay-based benchmark for language-model tutors, preview released 2026-08-07 (see Recent Activity). No wiki page; recorded here rather than as a standalone page, per the one-off-mention rule.
Recent Activity
- 2026-08-07: TutorMoments — a tutoring benchmark that scores holding back, not just helping — Ai2 released a preview of TutorMoments, an open, replay-based benchmark measuring whether a language-model tutor correctly chooses between helping a student and letting the student reason. TutorMoments-Preview is 462 de-identified, text-only transcripts of real one-on-one US maths tutoring with students in grades 2–7. Ai2's argument is that existing tutoring benchmarks reward a single fixed behaviour — never revealing the answer, or always offering a hint — "without accounting for whether that was the right move for where the student actually was in their understanding"; being replay-based lets TutorMoments judge an action against the actual state of the session. Preliminary results are reported as models tending to over-help. No per-model scores, leaderboard or numeric results were established from anything read. → Reasoning Models (source) (Ai2) (Hugging Face)
Strategic Position
TBD — nothing read on the run that created this page speaks to Ai2's positioning relative to other labs.
Related
- Reasoning Models — the over-helping result is about when a model should decline to supply reasoning