Retrieving by Frequency, Not Position: Frequency-Division Multiplexed Memory for Structured Retrieval in Language Models
Abstract
We introduce a frequency-division multiplexed (FDM) context-encoded token stream with functional natural language coexistence in transformer models, wherein multiple independent information channels for structured data are encoded as amplitude-shift keyed (ASK) sinusoidal carriers, superposed into a single token sequence, enabling accurate, uniform forward-pass retrieval of in-context memory. Each fact targeted for storage and accurate retrieval access is a KEY=VALUE pair encoded into K discrete-valued channels as ASK sinusoidal carriers superposed as a composite signal into a fixed token sequence. Every token therefore carries information about every channel, making retrieval frequency-addressed rather than position addressed, without explicit Fourier, demodulation, or retrieval-specific modules. We finetune four transformer architectures (GPT-2 Medium, Qwen3-0.6B, LFM2.5-1.2B, Hermes3-3B) on up to 215K training samples of structured, non-linguistic FDM encoded data, achieving 93–100% per-channel retrieval accuracy across a 32-channel Context block (ch8–39). We then test whether the learned decoding is tied to a single position range. A brief additional finetuning stage teaches the FDM-finetuned model to decode the block at a second position range, and a learned key-value (K/V) prefix placed in front of the block does not degrade retrieval. We then finetune under a five-stage curriculum of decreasing carrier amplitude contrast, interleaving general-text replay where ordinary English text is mixed into every batch (annealed from 30% to 10%) specifically to guard the model from losing its natural-language abilities while it learns to read the FDM-encoded channels. This is performed on a 10-channel mixed-domain configuration, wherein all four architectures sustain 96.9–97.5% FDM action accuracy and 99.9–100% per-channel retrieval while answering 8–10 of 10 held-out natural-language questions correctly (10/10 for Qwen3-0.6B), with the same forward pass and no routing or gating mechanism handling both input types. These results show that transformer attention can learn to retrieve from a frequency-encoded in-context memory substrate while basic natural-language question answering remains functional. These results suggest a new design space for structured, frequency-addressed memory substrates that support short, accurate retrieval of structured facts within a single forward pass.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.