BedrockBench: Can LLMs Speak Machine Code?
Abstract
Large language models achieve strong results in high-level code generation, but this success does not establish whether they can directly construct executable machine-code bytes. We introduce BedrockBench, the first benchmark for direct machine-code function synthesis from natural-language specifications, with 100 tasks targeting x86-64. Models submit hexadecimal function bytes without external tools or execution feedback. A shared evaluator executes these bytes unchanged and checks functional behavior, memory accesses, return state, and instruction and resource constraints across legal execution environments. Diagnostics distinguish submission failures, execution violations, and wrong answers. Across seven model configurations, pass@1 ranges from 13% to 100%, with GPT-6-Astra passing all 100 tasks. Among 240 format-valid unsuccessful candidates, 160 (66.7%) exhibit unauthorized memory accesses, making memory-contract violations a prominent observed failure mode. BedrockBench provides an explicit and auditable basis for comparing models and studying how to improve the reliability of direct machine-code synthesis. The benchmark data and code are publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.