acceptodds
Under review as a conference paper at ICLR 2027

Beyond Frontier Evaluation: Auditable Energy, Affordability, and Data-Governance Metrics for AI Deployed in the Global South

Abstract

AI evaluation and governance are indexed to frontier models, yet most people meet AI through distilled, fine-tuned, or platform-embedded systems, often in low- and middle-income countries where rules written for frontier labs propagate weakly. We argue that evaluation should target diffusion capacity in the technology-policy sense: whether a deployed system is affordable, runs on local infrastructure, and can be governed where it lands. We operationalize it with three instruments an auditor can compute or verify within one jurisdiction from public evidence. Joules per Task (JpT) labels inference energy with its measurement boundary and task harness, since rankings can reverse when either changes; in a provenance-preserving collation, published per-prompt figures for text generation (measured and estimated) span nearly three orders of magnitude, and boundary choice alone moves one provider's median-prompt estimate by 2.4×. A Price Accessibility Index (PAI) relates price to income: on a 20-economy panel, a USD 20 monthly subscription exceeds the UN 2%-of-monthly-income affordability line for broadband in 19 economies, burdens range from 6.5× to about 82× the US level, and a purchasing-power discount alone leaves half the panel above the line. A data-value-retention (DVR) vector grades disclosure of onshore storage, consent, local revenue sharing, and portability. Verdict rules grade each axis separately, withhold positive verdicts below a host-country governance threshold, and let oversight and rights failures cap a grade rather than be averaged away; by construction they are monotone in adverse evidence and admit no aggregate that can offset a veto, and each identified gaming route has a closing rule. On three focal cases (an open-weight model release, a data-centre build-out, and a mobile-money platform), the rules return capped and ungraded verdicts alongside positive ones; a re-verified dataset of 92 documented engagements by 39 Chinese AI firm groups in 24 countries grounds the cases descriptively. We keep in-sample consistency checks separate from a forward test specified in advance, and release the dataset, code, and JpT task suite.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.