acceptodds
Under review as a conference paper at ICLR 2027

RiboUnmix: Learning Shared Translational Dynamics from Biased and Noisy Ribo-seq Measurements

Abstract

Ribosome profiling (Ribo-seq) is used to study translation dynamics by measuring how ribosomes are distributed along mRNA sequences, but the resulting occupancy profiles also reflect experiment-specific distortions and stochastic measurement variability. Models trained to predict these profiles from mRNA sequences can also learn these distortions, so accurate profile prediction alone does not establish recovery of the underlying biological behavior. We ask whether combining Ribo-seq datasets affected by different experimental conditions can enable us to reveal shared sequence-dependent patterns in the underlying ribosome distribution. We introduce RiboUnmix, a probabilistic multi-dataset framework that jointly learns from multiple Ribo-seq datasets collected in different experiments. RiboUnmix represents each expected measured profile as a shared sequence dependent signal modulated by a dataset-specific multiplicative factor, while a negative-binomial observation model captures variability across individual Ribo-seq replicates. To evaluate recovery of the shared profile and dataset-specific effects, we construct a controlled synthetic Ribo-seq benchmark combining programmed translation kinetics, ribosome traffic, stochastic count sampling, and multiple sequence-dependent experimental distortions. The programmed kinetics and injected distortions provide known targets for evaluating the inferred shared profile and dataset-specific effects separately. Both inferred components show high correlation with their synthetic targets, showing RiboUnmix's ability to separate shared kinetic patterns from dataset-specific effects under controlled conditions. Across four organism-specific real-data benchmarks, RiboUnmix outperforms the evaluated sequence-to-profile baselines in predicting measured Ribo-seq profiles. The model independently trained on separate subsets of 114 HEK-derived datasets recovers concordant shared profiles for holdout transcripts, remarking reproducibility of the shared signal. Complementary experiments varying the number and composition of training datasets further show that the inferred shared representation remains substantially stable as the experimental evidence base changes. RiboUnmix turns variation across experiments into evidence for reproducible sequence-dependent patterns of ribosome occupancy, providing a foundation for biological hypothesis generation from diverse Ribo-seq datasets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.