From Tokens to Files: Multi-Granularity Detection and Localization of AI-Involved Code
Abstract
The increasing use of large language models for code generation has created an urgent need for reliable methods to distinguish purely human-written code from *AI-involved* code — files in which any part was generated or substantially modified by a model, however small that part is. Our goal is not only to determine whether a program is purely human-written or has been generated or modified by AI, but also to localize the specific lines or blocks attributable to AI. Existing approaches typically compress a file into a single score, either from surface-level stylistic features or from a global code representation, which averages away the evidence when only a few blocks of a file are AI-involved. We propose a detection framework that models AI authorship at the token level and supports predictions at the line, block, and file levels. A frozen code language model is equipped with two low-rank adapters, trained separately on machine-written and human-written files. The per-token log-likelihood ratio between the two adapters provides local evidence of AI authorship. Lightweight prediction heads then aggregate this evidence within lines and blocks, using the surrounding code context to estimate the probability that each line or block was generated or substantially modified by AI. The resulting local predictions are further aggregated to produce a file-level decision. We train these predictions using mixed-authorship files constructed from human originals and their AI-generated or AI-modified counterparts, including continuation and fill-in-the-middle scenarios. On an internal test set of 11,773 AI-involved files across six generation regimes against a shared pool of 2,541 human files, our method reaches 72.4 macro F1 at a 1% FPR operating point, compared with 60.4 for a classifier that shares its backbone, training files, truncation, and optimisation schedule, 58.4 and 56.5 for fine-tuned UniXcoder and CodeBERT, and 48.5 for a style-preference baseline. The gain concentrates on the mixed-authorship regimes, where our method achieves 60.8 and 50.4, compared with 42.2 and 41.3 for the corresponding baseline. With no adaptation at all, the method reaches 62.5 line-level F1 on HybridCodeAuthorship, an external human–AI co-writing benchmark that ships no training split, above the 56 that the benchmark's own paper reports as its best line-level result. We further run an extra transfer evaluation on a human-written-only test set: across 19,445 files from repositories created before 2022, our false-positive rate stays between 0.55 and 0.81% when the false-positive rate threshold is transported at 1%. So the operating point remains valid after transferring from paired synthetic negatives to real human code.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.