acceptodds
Under review as a conference paper at ICLR 2027

RabbitBench: Evaluating Large Language Models across Diverse Occupations through Occupational Qualification Examinations

Abstract

As large language models (LLMs) increasingly participate in occupational work, there is a growing need to comprehensively evaluate their knowledge and capabilities across diverse occupational domains. However, existing benchmarks still lack comprehensive and robust coverage across a broad range of occupations. Occupational qualification examinations support comprehensive and robust LLM evaluation through expert-designed questions that systematically cover occupational tasks and the knowledge and capabilities they require. Following official Chinese guidelines, we collect and organize accessible and evaluable examinations from 26 types of occupational qualifications to construct RabbitBench. The benchmark spans 8 mid-level and 23 minor-level occupational groups in the Chinese occupational classification system, providing broader coverage than existing benchmarks. It includes single-choice, multiple-choice, true-or-false, and subjective questions, together with high-quality reference answers, enabling the evaluation of multiple aspects of model performance. We evaluate a diverse set of LLMs on RabbitBench. GPT-5.6-Sol and Gemini-3.1-Pro achieve the strongest performance but at substantially higher API costs, while Qwen3.5 shows a favorable cost–performance trade-off. Across occupations, general capability dominates overall performance, but similarly capable models exhibit distinct occupational strengths. Within engineering qualifications, models consistently perform better on knowledge-oriented than application-oriented subjects. Together, these results show that RabbitBench enables broad and fine-grained analysis of occupational capabilities beyond aggregate model scores. An anonymous project page is available at https://rabbitbench.github.io/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.