Sampling without losing the spans that moved money

2 min read

Short answer

Sampling decides which traces you keep. Head sampling decides when the trace starts, usually from the trace ID and a fixed rate. Tail sampling decides after the spans have arrived, so it can keep every trace with an error. A reverted payment is only known at the end, so a head sampler drops it as often as any other trace.

On this page

Sample 10% of agent runs and you keep about 10% of your reverted payments. The sampler decided before the transaction was even sent.

Head sampling at 10% keeps only some reverted payments because it decides before they happen, while Collector tail sampling keeps every error and every transaction.

Head or tail?

Head sampling Tail sampling
Decides When the trace starts After its spans arrive
Knows about errors? No Yes
Runs in The SDK A stateful component, e.g. the Collector
Cost Cheap Memory for every pending trace

The sampling docs list the catch with head sampling: it can’t decide based on data in the entire trace, so you can’t make sure traces with errors are kept.

What does a 10% head sampler drop?

1,000 agent runs, every tenth payment reverts, sampled with the usual ParentBased + TraceIdRatioBased pair:

import { SpanStatusCode } from '@opentelemetry/api';
import { NodeTracerProvider } from '@opentelemetry/sdk-trace-node';
import { ParentBasedSampler, SimpleSpanProcessor, TraceIdRatioBasedSampler } from '@opentelemetry/sdk-trace-base';
new NodeTracerProvider({
sampler: new ParentBasedSampler({ root: new TraceIdRatioBasedSampler(0.1) }), // keep 10% of runs
spanProcessors: [new SimpleSpanProcessor(exporter)],
}).register();
for (let run = 0; run < 1000; run++) {
tracer.startActiveSpan('invoke_agent treasury-agent', (agent) => {
const confirm = tracer.startSpan('confirm 8453');
if (run % 10 === 0) confirm.setStatus({ code: SpanStatusCode.ERROR }); // every 10th payment reverts
confirm.end();
agent.end();
});
}

Three runs:

Run Confirm spans kept Reverted payments kept
1 108 of 1,000 11 of 100
2 100 of 1,000 15 of 100
3 88 of 1,000 9 of 100

The confirm span follows the root’s decision. Raising its own priority would only produce orphan spans whose parents were dropped.

Keep the traces that moved money

Make the decision at the tail. With the Collector’s tail sampling processor, a trace is kept if any policy says so:

processors:
tail_sampling:
decision_wait: 30s
policies:
- name: errors
type: status_code
status_code: { status_codes: [ERROR] }
- name: moved-money
type: string_attribute
string_attribute: { key: blockchain.operation.name, values: [send, payment] }
- name: ten-percent-of-the-rest
type: probabilistic
probabilistic: { sampling_percentage: 10 }

Every failed span and every trace with a transaction stays. Plain chat runs are sampled at 10%. The SDK must then export everything to the Collector, so keep head sampling at 100%.

FAQ

What is the difference between head and tail sampling?

Head sampling decides as early as possible, before the trace is complete, which makes it cheap but blind to errors and latency. Tail sampling looks at all or most of a trace's spans first, which needs a stateful component such as the Collector's tail sampling processor.

What sampling rate should I use?

There's no single number. Keep every trace you'd need for an incident, such as errors and anything that moved money, and sample the rest at whatever rate your backend budget allows.

Does ParentBased sampling keep child spans of a dropped trace?

No. A ParentBased sampler follows the parent's decision, so when the root span isn't sampled, its children aren't either.

Go further

Type to search the docs, Learn topics and the blog.