Cardinality that kills the metrics bill

Back to list2025-10-08
You

Our metrics cost exploded after adding a user id label.

ChatGPT

That is the mechanism: every distinct label value creates a series. User ids, request ids and full URLs never belong in labels. Put them in logs and traces, where high cardinality is the point rather than the problem.

You

Head sampling drops exactly the requests I want to see.

ChatGPT

Tail sampling: buffer the spans and decide after the outcome is known, keeping everything that errored or exceeded a latency threshold plus a small random slice of the rest. Costs more memory at the collector and is almost always worth it.

You

Small service, limited time. Which one first?

ChatGPT

Logs with structure, because they answer "what happened" for incidents you did not anticipate. Metrics come next for the handful of numbers you would page on. Traces are worth it once a request crosses more than two services.