acceptodds
Under review as a conference paper at ICLR 2027

Image Editing without Typing

Abstract

Text-guided Image Editing (TIE) uses a typing box to replace the button-driven interfaces in traditional software such as Adobe Photoshop, offering remarkable convenience. However, it suffers from a critical limitation: on mobile devices, the virtual keyboard occupies nearly half of the screen, which occludes or resizes the image to be edited and deteriorates user experience. We propose a potential new framework: the text box is directly eliminated, and users express all editing requirements by drawing or writing on the image with a virtual pen. We present the Jasmine Project, involving: (1) a new task NoTIE: No-Typing Image Editing. We define 14 subtasks, and conduct a user survey to identify how users express editing demands using the virtual pen; (2) a new dataset Jasmine-Dataset-15K: a high-quality 2K-resolution dataset, strictly filtered by real humans with a 50.4% high discard rate, with Style Change subtask samples exclusively reviewed by art-major students; (3) a new benchmark Jasmine-Benchmark-1K: 1,381 high-quality human-selected samples, equipped with an MLLM-as-a-Judge evaluation pipeline. We benchmark 19 SOTA models and our 3 models. All evaluation prompts are twice the length of the previous TIE work, and all scoring criteria are derived from the first-hand feedback during the human filtering process; (4) a new model series Jasmine-Model-Family: a Thinker-Editor architecture that outperforms all open-source models and ByteDance's Seedream series. We hope this work can make image editing more convenient on mobile devices, and we call for community help to extend this work.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.