MPC-Patch-Bench: Beyond Functional Tests for LLM Repair of MPC Software
Abstract
Repairing secure multi-party computation (MPC) software requires more than producing outputs that pass existing tests. Patches must also respect intended information disclosure and numerical behavior. We introduce MPC-Patch-Bench, a repository-level benchmark containing 205 executable repair tasks from five MPC frameworks. Its construction combines MPC-specific pull-request filtering with expert-reviewed synthesis of missing problem statements and regression tests. Evaluation supplements functional tests with differential checks against plaintext references and static rules targeting potential MPC-specific violations. Across eight evaluated LLMs, Sonnet 4.6 achieves the highest functional resolution rate, 22.9%, but falls to 16.1% after the additional checks; Opus 4.6 then ranks highest at 17.1%. Across models, these checks reject from 11% to 40% of functionally passing patches. The results show that functional-test performance and performance under additional MPC-specific checks are distinct evaluation signals. The benchmark supports studying this gap without treating check-passing patches as formally secure.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.