You
Our metrics cost exploded after adding a user id label.
ChatGPT
That is the mechanism: every distinct label value creates a series. User ids, request ids and full URLs never belong in labels. Put them in logs and traces, where high cardinality is the point rather than the problem.
You
Head sampling drops exactly the requests I want to see.
ChatGPT
Tail sampling: buffer the spans and decide after the outcome is known, keeping everything that errored or exceeded a latency threshold plus a small random slice of the rest. Costs more memory at the collector and is almost always worth it.