acceptodds
Under review as a conference paper at ICLR 2027

Learning Cost and Token Use with Controlled Literary Chinese: A Paired State-Tracking Study

Abstract

We measure learning cost and generation token use separately: shorter examples need not reduce the cost of learning. We compare controlled Literary Chinese (WY) and controlled English (EN) in 80 pre-specified pairs of from-scratch Transformer classifiers across three scales on a fixed 720-item training inventory. Certification requires at least 75% inventory accuracy throughout evaluations spanning 7,200 presentations, charging training through certification. All 80 pairs attain it with lower WY costs; scales are analyzed separately. The 40 primary 10M pairs yield a mean cost difference of −0.7447 S (approximate paired-t 95% interval [−0.7811, −0.7084] S), where S is the scale-specific training-FLOPs unit: descriptive mean paired savings are 50.38%, with similar repetitions. A development-informed supplement trains 20 paired 11.24M-parameter generators. At equal exposure, WY uses 62.12% fewer training tokens; complete-output test accuracy is WY 97.86%, EN 95.03%. At nominal 16-million-token budgets, accuracy is WY 98.09%, EN 95.15%, with 2.64 times as many WY presentations. Input-plus-generated-output tokens per request, including visible steps and errors, are then 61.98% lower. One training-unseen layout accounts arithmetically for 39–43% of generator quality gaps in a post-hoc audit. Classifier transfer is poor, and generator evaluation data were development-reused. The evidence quantifies acquisition cost, checkpoint quality, and visible-output length for these task-specific packages, without isolating language effects or establishing matched-test-quality deployment savings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.