Iditarod: What Makes a Good Distributed Trace Representation?
Abstract
Distributed traces—graph-based records of individual request executions—are critical for understanding and operating production software systems across industries, supporting security, performance optimization, and debugging. However, they have received comparatively little attention in machine learning (ML), despite rapid progress in both AI-driven software engineering and graph representation learning. We address this gap with Iditarod, the first multitask benchmark for distributed trace representation learning on real-world data. Iditarod is based around a new corpus of 3.18 million traces drawn from a commercial observability platform's own telemetry, combined with a subset of the Uber trace corpus. Its four tasks evaluate two complementary aspects of representation quality: (a) preserving key properties of individual traces through two intrinsic tasks and (b) providing efficient and performant embeddings for downstream applications through the Trace Retrieval task and the multimodal Faulty Deployment detection task over traces and time series. Our evaluation of sixteen different trace embedding methods on Iditarod includes models that process traces as either serialized text, call-path collections, or graphs. We find that to match or outperform language encoders on the intrinsic tasks and Trace Retrieval, graph neural networks often need to exploit the directed acyclic structure of traces. On Faulty Deployment, meanwhile, we find that no encoder produces trace embeddings that improve upon using time series alone, while a handcrafted feature extractor does, revealing a gap in current graph representation capabilities. Iditarod's data and evaluation code are available under the Apache 2.0 License.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.