Teach2Eval
Teach2Eval is an interaction-driven protocol for evaluating a language model by how much its guidance improves weaker language models. Student models answer multiple-choice tasks, the candidate teacher inspects each question and a student's response without seeing the choices or gold label, and the student answers again after guidance. The average change in student accuracy is reported as Comprehensive Ability.
Origin and context
The method first appeared as a May 2025 preprint and was substantially revised for ICLR 2026. The published study evaluates 33 LLMs over 60 datasets spanning knowledge, reasoning, understanding and multilingual tasks. Its metrics separate direct Application from Judgment, first-round Guidance and multi-round Reflection. The project repository provides code, test data and result artifacts, although its short README does not constitute a complete reproducibility report.
Why it matters
Static answer accuracy can reward exposure to test items and says little about whether a model can diagnose another model's error. Teach2Eval changes the observable: guidance must cause a weaker model to repair answers. In the authors' runs, its ranking correlated more strongly than direct evaluation with Chatbot Arena and LiveBench. Independent TeachBench research also adopts student improvement as a measurable signal, while narrowing teaching to syllabus knowledge rather than target questions.
Example
An evaluation team pins a benchmark subset, four student models, teacher and student prompts, turn budget, decoding settings and endpoint versions. Each candidate teacher guides the same initial student responses without seeing answer choices. The team reports direct accuracy beside student lift, per-ability metrics, variance across students and turns, token cost and failure cases. Re-running with alternative student pools tests whether the ordering is robust to the evaluator design.
How it differs
LLM-as-a-judge
LLM-as-a-judge asks a model to grade or compare outputs. Teach2Eval instead scores a candidate teacher through measured changes in student answers, although models still participate in data preparation and parts of the evaluation pipeline.
Benchmark contamination
Benchmark contamination is exposure to evaluation material. Blinding teachers to choices and labels weakens option matching, but the questions, derived MCQs, student models and generated guidance can introduce other dependencies; Teach2Eval is mitigation, not a contamination certificate.
Judge Calibration
Judge calibration concerns whether an evaluator's scores track a target standard. Teach2Eval's leaderboard correlations are calibration evidence for one configuration, while sensitivity to student selection and task construction remains a separate question.
Maturity and evidence
Maturity is 3. Teach2Eval has a dated origin, peer-reviewed ICLR publication, public implementation artifacts, multi-model experiments and independent follow-on scholarship that both recognizes the approach and sharpens its boundary. It is not rated higher because independent score reproduction was not located and the method remains configuration-sensitive.
Limits and open questions
Teach2Eval measures improvement in simulated LLM students, not learning, safety or instructional quality for people. Rankings can change with the student pool, baseline accuracy, task difficulty, generated distractors, number of turns, prompts and current model endpoints. The paper's high correlations and contamination experiment were produced by the authors; they should not be described as independent validation or universal robustness. Because the teacher sees the target question, later TeachBench work identifies possible problem-level information leakage and evaluates a different syllabus-grounded setting. Reports should preserve the exact protocol and show direct scores, student lift, variance and cost rather than collapsing everything into one timeless rank.
Related terms
References
- Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It TeachesarXiv · 2025-05-18 · class A
- Teach2Eval: An Interaction-Driven LLMs Evaluation Method via Teaching EffectivenessInternational Conference on Learning Representations · 2026 · class A
- Teach2Eval source repositoryTeach2Eval authors · 2025 · class A
- TeachBench: A Syllabus-Grounded Framework for Evaluating Teaching Ability in Large Language ModelsPeking University, ByteDance BandAI, and Institute of Automation, Chinese Academy of Sciences · 2026-01-29 · class B
Last updated: 2026-09-07
This term is also covered in the Skills Atlas as model evaluation skill.
This term is also covered in the Skills Atlas as llm benchmarking skill.
This term is also covered in the Skills Atlas as llm evaluation design skill.
This term is also covered in the Skills Atlas as evaluation data engineering skill.