acceptodds
Under review as a conference paper at ICLR 2027

ViralBench: A Bilingual Benchmark and Taxonomy for Quantifying the Memorization Gap in LLMs

Abstract

Every few months, a single "trick question" goes viral on social media by making state-of-the-art large language models (LLMs) fail in embarrassingly simple ways: counting the letter *r* in *strawberry*, deciding whether *9.9* or *9.11* is larger, or asking whether one should walk or drive to a car wash 50 meters away. These questions spread rapidly, and within weeks the newest model checkpoints "fix" them. But is such fixing genuine capability improvement, or targeted memorization of the specific viral instance? We argue this question has been impossible to answer because the community's trick questions are scattered, lack canonical answers, and are never evaluated under a controlled protocol. We introduce **ViralBench**, a bilingual (Chinese–English) benchmark of **64 seed questions** that genuinely went viral—each annotated with its source URL and the period it circulated—plus **292 programmatically generated perturbed variants** whose answers are computed by construction, for **356 items** in total. Our central methodological contribution is a **two-axis taxonomy**: a *surface-form* axis (`category`, 7 classes describing how a question circulates online) crossed with a *failure-mechanism* axis (`failure_mode`, 6 hypotheses about why models err). The seed-versus-variant design lets us measure a **memorization gap**: the accuracy drop from a memorized viral instance to its structurally identical but novel perturbation. On [FABRICATED] N models we find [FABRICATED] that non-reasoning models exhibit a mean memorization gap of 0.13, that this gap correlates with whether a question went viral before the model's training cutoff, and that cross-lingual answer consistency is lowest for sub-token perception failures. We release the dataset, the deterministic variant generator, and the evaluation harness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.