acceptodds
Under review as a conference paper at ICLR 2027

OpenJarvis: Personal AI, On Personal Devices

Abstract

Personal AI stacks for writing, research, coding, and scheduling are becoming central to daily work, yet most still route every query to cloud-hosted frontier models. This paradigm exposes private data and creates recurring API spend. Excitingly, open-source models and consumer accelerators are significantly improving, enabling the possibility of a personal AI stack running on-device. However, we find that replacing cloud models with local models in existing personal AI stacks, such as OpenClaw and Hermes Agent, leads to significant deterioration in performance. For example, replacing Claude Opus 4.6 with Qwen3.5-9B drops accuracy by 25–39 percentage points across PinchBench and GAIA. This collapse reflects both a model-capability gap and harness incompatibility: agentic prompts, tool descriptions, memory configuration, and inference runtime settings carry model-specific defaults that materially contribute to the substitution gap. Towards building an on-device personal AI stack, we present OpenJarvis, an architecture that represents a personal AI system as a typed spec over five primitives: Intelligence, Engine, Agents, Tools & Memory, and Learning. By exposing each primitive as an independently editable field, the spec makes the stack portable, measurable, and end-to-end optimizable around any choice of model. On-device specs match or exceed cloud accuracy on 4 out of 8 benchmarks, and the best one (Qwen3.5-122B) lands within 3.2 pp of the best cloud baseline on average with up to a 6,600× reduction in marginal API cost; Qwen3.5-35B achieves 4× lower end-to-end latency. To close the remaining accuracy gap between the best cloud model and the best local model, OpenJarvis uses LLM-guided spec search, in which a frontier teacher proposes edits across the spec and a held-out gate accepts only non-regressing improvements. Search improves local students by 13–32 pp on average at 7–11× lower optimization cost than LoRA fine-tuning, with student inference and training on-device.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.