acceptodds
Under review as a conference paper at ICLR 2027

AutoFHE-Bench: Beyond Pass Rates in Agent-generated Fully Homomorphic Encryption Programs

Abstract

Fully Homomorphic Encryption (FHE) enables computation over encrypted data without decryption, serving as a foundation for privacy-preserving machine learning. Writing an FHE program remains expert-only work: developers must manually pick encryption parameters, layout ciphertexts, and approximate unsupported operations, where any mistake fails silently under encryption. While LLM agents excel at general software synthesis, applying them to FHE exposes a fundamental vulnerability: a program can produce the correct answer while providing zero confidentiality. Prior evaluations infer security merely from static source code, a check easily satisfied by programs that invoke homomorphic APIs while computing in the clear. We show that their remaining axis, functional correctness, is nearly saturated: on ten end-to-end tasks from established FHE benchmarking suites, each with a verified reference, three frontier models succeed in 147 of 150 attempts. The open question is therefore no longer whether an agent can write a working FHE program, but whether that program is secure and deployable. We introduce AutoFHE-Bench, which treats security as an observable runtime property of an execution. Across multi-stage pipelines from key generation to decryption, an independent judge holds the secret keys, draws fresh inputs, enforces OS-level client–server isolation, and decrypts results itself, and it bounds multiplicative depth and security level on the tasks that admit such a bound. Nine adversarial control types score zero across 30 executions. The benchmark shows that submissions receiving the same verdict differ, at the extremes, by in encrypted-computation time on matrix inversion and in evaluation-key size on encrypted retrieval, and that stating and enforcing a depth budget reduces one model's median key material by at an unchanged pass rate. Two exploratory challenge tasks, reported separately, bound what the judge establishes: on one a model synthesizes a bootstrapped ResNet-20 for CIFAR-10 at one quarter of the reference key material, and on the other a passing submission shows that verifying an encrypted execution does not verify which computation it performed. We open-source the tasks, the judge, and the controls.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.