OpenJarvis: Personal AI, On Personal Devices
Abstract
Personal AI stacks for writing, research, coding, and scheduling are becoming central to daily work, yet most still route every query to cloud-hosted frontier models. This paradigm exposes private data and creates recurring API spend. Excitingly, open-source models and consumer accelerators are significantly improving, enabling the possibility of a personal AI stack running on-device. However, we find that replacing cloud models with local models in existing personal AI stacks, such as OpenClaw and Hermes Agent, leads to significant deterioration in performance. For example, replacing Claude Opus 4.6 with Qwen3.5-9B drops accuracy by 25–39 percentage points across PinchBench and GAIA. This collapse reflects both a model-capability gap and harness incompatibility: agentic prompts, tool descriptions, memory configuration, and inference runtime settings carry model-specific defaults that materially contribute to the substitution gap. Towards building an on-device personal AI stack, we present OpenJarvis, an architecture that represents a personal AI system as a typed spec over five primitives: Intelligence, Engine, Agents, Tools & Memory, and Learning. By exposing each primitive as an independently editable field, the spec makes the stack portable, measurable, and end-to-end optimizable around any choice of model. On-device specs match or exceed cloud accuracy on 4 out of 8 benchmarks, and the best one (Qwen3.5-122B) lands within 3.2 pp of the best cloud baseline on average with up to a 6,600× reduction in marginal API cost; Qwen3.5-35B achieves 4× lower end-to-end latency. To close the remaining accuracy gap between the best cloud model and the best local model, OpenJarvis uses LLM-guided spec search, in which a frontier teacher proposes edits across the spec and a held-out gate accepts only non-regressing improvements. Search improves local students by 13–32 pp on average at 7–11× lower optimization cost than LoRA fine-tuning, with student inference and training on-device.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.