RabbitBench: Evaluating Large Language Models across Diverse Occupations through Occupational Qualification Examinations
Abstract
As large language models (LLMs) increasingly participate in occupational work, there is a growing need to comprehensively evaluate their knowledge and capabilities across diverse occupational domains. However, existing benchmarks still lack comprehensive and robust coverage across a broad range of occupations. Occupational qualification examinations support comprehensive and robust LLM evaluation through expert-designed questions that systematically cover occupational tasks and the knowledge and capabilities they require. Following official Chinese guidelines, we collect and organize accessible and evaluable examinations from 26 types of occupational qualifications to construct RabbitBench. The benchmark spans 8 mid-level and 23 minor-level occupational groups in the Chinese occupational classification system, providing broader coverage than existing benchmarks. It includes single-choice, multiple-choice, true-or-false, and subjective questions, together with high-quality reference answers, enabling the evaluation of multiple aspects of model performance. We evaluate a diverse set of LLMs on RabbitBench. GPT-5.6-Sol and Gemini-3.1-Pro achieve the strongest performance but at substantially higher API costs, while Qwen3.5 shows a favorable cost–performance trade-off. Across occupations, general capability dominates overall performance, but similarly capable models exhibit distinct occupational strengths. Within engineering qualifications, models consistently perform better on knowledge-oriented than application-oriented subjects. Together, these results show that RabbitBench enables broad and fine-grained analysis of occupational capabilities beyond aggregate model scores. An anonymous project page is available at https://rabbitbench.github.io/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.