Trace Sampling at the Collector Boundary: Costs and Diagnostic Evidence

Trace retention is an incomplete predictor of observability cost: a sampler changes where work is avoided, how spans are grouped for export, and which evidence remains available. We study these effects in OpenTelemetry Collector Contrib v0.136.0 on one shared node with loopback transport. Five randomized blocks cross sampler placement with export path at 40,000 offered spans/s. Native uniform sampling at nominal 10% retention increases Collector CPU by 3.2% with JSON-only export and reduces it by 24.2% with JSON plus Jaeger. A pre-ingress gate retains identical trace IDs but avoids ingestion and excludes selection work from the Collector endpoint. Retain-all controls show that stateful release changes batching and CPU without discarding traces. A short-timeout CPU reduction at 250 spans/s disappears at 5,000 spans/s, where tail increases CPU at both tested timeouts. A separate SDK comparison, load sweep, and delivery probes distinguish application work, Collector resources, and complete evidence delivery. An illustrative localization task uses measured HTTP timings and ideal offline sampling without an SDK or Collector. Its results show how window size and selection-dependent reference evidence affect a fixed median-change scorer, conditional on the observed corpus. The study supports evaluating sampling at explicit component boundaries, measuring batching and serialization alongside trace counts, and defining the diagnostic evidence objective before selecting a rate. The research archive preserves frozen protocols, raw observations, failed attempts, and executable checks. Research artifact. Version 1.8.0 contains the complete 22-page article, including the mathematical supplement as Appendix A. The complete research artifact is available from the immutable GitHub v1.8.0 release linked below. It includes all 9,242 raw evidence files, including experiment logs, test-environment metadata, pilots, failed attempts, frozen execution sources, derived results, analysis, limited Lean proofs, licenses, and validation records. The 2.48 GB archive is distributed there in three numbered parts with RELEASE.json, VERIFICATION.log, and SHA256SUMS. This Zenodo record deposits the complete article PDF; the research archive is hosted in the linked GitHub release. Historical availability statements in the unchanged paper and sealed archive describe their preparation checkpoint before public deposition. This is a preprint and has not been peer reviewed. AI-assisted research workflow. OpenAI Codex assisted protocol design, literature discovery, workload and orchestration programs, analysis and validation code, mathematical derivations and proof code, figure-generation code, manuscript drafting, and revision. Automated workloads collected the measurements in the Kubernetes testbed; timings and resource counters are archived program observations. The validation implementations are part of the same AI-assisted workflow, not external replication. The author retains responsibility for the methods, claims, source attribution, and final text. License scopes. The deposited manuscript is licensed under Creative Commons Attribution 4.0 International (CC BY 4.0). Original research data, figures, derived results, and narrative documentation in the linked artifact use CC BY 4.0; original software, formalization, build, and deployment code use MIT. These licenses apply to their respective components. Third-party material retains its applicable terms; LICENSE.md in the archive details the scope. Research archive: https://github.com/vicotrbb/trace-sampling-strategies/releases/tag/v1.8.0

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-29
DOI
https://doi.org/10.5281/zenodo.23032244
Primary Topic
Scientific Computing and Data Management
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Trace Sampling at the Collector Boundary: Costs and Diagnostic Evidence

Victor Bona
Zenodo (CERN European Organization for Nuclear Research)
Scientific Computing and Data Management
preprint

Trace Sampling at the Collector Boundary: Costs and Diagnostic Evidence

Victor Bona
preprint en

Abstract

Trace retention is an incomplete predictor of observability cost: a sampler changes where work is avoided, how spans are grouped for export, and which evidence remains available. We study these effects in OpenTelemetry Collector Contrib v0.136.0 on one shared node with loopback transport. Five randomized blocks cross sampler placement with export path at 40,000 offered spans/s. Native uniform sampling at nominal 10% retention increases Collector CPU by 3.2% with JSON-only export and reduces it by 24.2% with JSON plus Jaeger. A pre-ingress gate retains identical trace IDs but avoids ingestion and excludes selection work from the Collector endpoint. Retain-all controls show that stateful release changes batching and CPU without discarding traces. A short-timeout CPU reduction at 250 spans/s disappears at 5,000 spans/s, where tail increases CPU at both tested timeouts. A separate SDK comparison, load sweep, and delivery probes distinguish application work, Collector resources, and complete evidence delivery. An illustrative localization task uses measured HTTP timings and ideal offline sampling without an SDK or Collector. Its results show how window size and selection-dependent reference evidence affect a fixed median-change scorer, conditional on the observed corpus. The study supports evaluating sampling at explicit component boundaries, measuring batching and serialization alongside trace counts, and defining the diagnostic evidence objective before selecting a rate. The research archive preserves frozen protocols, raw observations, failed attempts, and executable checks. Research artifact. Version 1.8.0 contains the complete 22-page article, including the mathematical supplement as Appendix A. The complete research artifact is available from the immutable GitHub v1.8.0 release linked below. It includes all 9,242 raw evidence files, including experiment logs, test-environment metadata, pilots, failed attempts, frozen execution sources, derived results, analysis, limited Lean proofs, licenses, and validation records. The 2.48 GB archive is distributed there in three numbered parts with RELEASE.json, VERIFICATION.log, and SHA256SUMS. This Zenodo record deposits the complete article PDF; the research archive is hosted in the linked GitHub release. Historical availability statements in the unchanged paper and sealed archive describe their preparation checkpoint before public deposition. This is a preprint and has not been peer reviewed. AI-assisted research workflow. OpenAI Codex assisted protocol design, literature discovery, workload and orchestration programs, analysis and validation code, mathematical derivations and proof code, figure-generation code, manuscript drafting, and revision. Automated workloads collected the measurements in the Kubernetes testbed; timings and resource counters are archived program observations. The validation implementations are part of the same AI-assisted workflow, not external replication. The author retains responsibility for the methods, claims, source attribution, and final text. License scopes. The deposited manuscript is licensed under Creative Commons Attribution 4.0 International (CC BY 4.0). Original research data, figures, derived results, and narrative documentation in the linked artifact use CC BY 4.0; original software, formalization, build, and deployment code use MIT. These licenses apply to their respective components. Third-party material retains its applicable terms; LICENSE.md in the archive details the scope. Research archive: https://github.com/vicotrbb/trace-sampling-strategies/releases/tag/v1.8.0

Zenodo (CERN European Organization for Nuclear Research)
Scientific Computing and Data Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.