基于多智能体协同的基础数学解答题手写答案自动评阅研究

Research on Automatic Grading of Handwritten Answers to Basic Mathematics Solution Problems Based on Multi-Agent Collaboration

  • 摘要: 对数学试卷中主观题的手写答案进行自动评阅是人工智能赋能教育的重要应用场景之一。现有的依赖多模态大模型以提示词的方式对解答题手写答案评阅的方法,存在多行答案识别准确度有限和评分依据不透明等问题,进而影响评阅公平性难以保障的瓶颈。该研究提出了一个多智能体协作的智能处理框架,将自动评阅任务拆分为子任务并分配至四类智能体:主管智能体负责跨智能体协同指挥;识别智能体负责手写文本识别;分析智能体负责生成评阅基准图和二次分析;评阅智能体负责得分点匹配。智能体分工专精、协同互补,既突破了单一多模态大模型在自动评阅任务上的瓶颈,也通过模块化设计为后续突破难度更高的评阅任务提供了扩展便利性。以基础数学解答题及其手写答案为测试样本,与GPT-4o和Qwen-VL-30B两个模型进行了手写识别的对比实验,与DeepSeek-V3.2、Doubao-Seed-1.6和Qwen3-Plus三个大模型进行了人机评阅对比实验。结果显示:在手写识别中整体平均编辑距离和精确识别率(ARR)分别为0.40和85%;在评分中人机平均绝对误差、均方根误差和一致率分别为0.96、1.92和61%,均优于单一的大模型方法,为智慧教育中高效、透明的自动评阅系统落地提供了可行的技术参考。

     

    Abstract: Automatically grading handwritten answers to subjective questions in math exams is one of the important application scenarios of AI-empowered education. The existing methods, which rely on multimodal large models to grade handwritten answers to open-ended questions using prompts, face bottlenecks in accurately recognizing multi-line answers and ensuring fairness due to opaque grading criteria. This study proposes a multi-agent collaborative intelligent processing framework that divides the automatic grading task into subtasks and assigns them to four types of agents: the supervisor agent responsible for cross-agent collaborative coordination; the recognition agent responsible for handwriting text recognition; the analysis agent responsible for generating grading benchmark images and secondary analysis; and the grading agent responsible for matching scoring points. The specialized and complementary collaboration among agents not only breaks through the bottlenecks of single multimodal large models in automatic grading tasks but also provides scalability for future grading tasks with higher difficulties through modular design. The experiments in this study use basic math open-ended questions and their handwritten answers as test samples. Comparative experiments were conducted with GPT-4o and Qwen-VL-30B for handwriting recognition, and with DeepSeek-V3.2, Doubao-Seed-1.6, and Qwen3-Plus for human-machine grading comparison. The experimental results show that the overall average edit distance and accurate recognition rate in handwriting recognition are 0.40 and 85%, respectively; in scoring, the average human-machine absolute error, root mean square error, and consistency rate are 0.96, 1.92, and 61%, respectively, all outperforming single large model methods. This provides a feasible technical reference for the implementation of efficient and transparent automatic grading systems in intelligent education.

     

/

返回文章
返回