PRISM: Modular Allocation and Complementary Fusion for Argument Quality Assessment
Abstract
Continuous argument quality assessment (AQA) must rank arguments across topics and annotation scales, while aggregate scores obscure the distinct roles of model capacity, complementary signals, and score calibration. We introduce , a modular framework that allocates these roles to separate interfaces. An independently trained eight-replica Qwen3-8B ensemble provides one score, while a taxonomy-profile encoder provides a complementary signal; corpus-specific experts retain separate scales for closed-set IBM/Webis routing. On IBM-Rank-30k, the Qwen ensemble reaches Spearman , profile calibration raises encoder Spearman from to , and the frozen encoder complement raises the final fusion to ( over Qwen alone). After fusion, Pearson and Spearman show relative improvements of and , respectively, over the strongest published reference in our comparison. On Webis-ArgQuality-20, target-domain supervision substantially outperforms IBM zero-shot transfer; in a fixed IBM/Webis mixture, confidence-based abstention over development-fitted scale maps significantly improves Webis ranking over hard routing while evaluating both experts only on low-confidence samples. therefore makes the scorer, encoder calibration, and protocol-aware scale adaptation separately measurable.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.