Challenge
Evaluating the validity of exam questions required specialized judgment, and creating and verifying questions took a great deal of time and effort.
Learn more about Challenge
日本語版Case study / Aug 5, 2025
Evaluating the validity of exam questions required specialized judgment, and creating and verifying questions took a great deal of time and effort. In particular, differences in question-writing experience and inconsistent evaluation criteria made it hard to maintain consistent judgment, raising concerns about the risk of misjudgment.

Evaluating the validity of exam questions required specialized judgment, and creating and verifying questions took a great deal of time and effort.
Learn more about ChallengeBased on past questions and expert guidelines, an LLM automatically judges the validity of exam questions.
Learn more about ApproachGreatly reduced the effort required for creating and verifying questions, easing the burden on educators.
Learn more about OutcomesEvaluating the validity of exam questions required specialized judgment, and creating and verifying questions took a great deal of time and effort.
In particular, differences in question-writing experience and inconsistent evaluation criteria made it hard to maintain consistent judgment, raising concerns about the risk of misjudgment.
Based on past questions and expert guidelines, an LLM automatically judges the validity of exam questions.
It also automatically generates explanatory text to support each judgment, achieving both transparency and explainability.
Built a mechanism that achieves stable quality control without relying on subjective judgment.

Greatly reduced the effort required for creating and verifying questions, easing the burden on educators.
By reducing variation in judgment, established a structure that maintains objective, consistent quality.
