Our metrics cost exploded after adding a user id label

Back to list2024-06-07
You

Our metrics cost exploded after adding a user id label.

ChatGPT

That is the mechanism: every distinct label value creates a series. User ids, request ids and full URLs never belong in labels. Put them in logs and traces, where high cardinality is the point rather than the problem.

You

Head sampling drops exactly the requests I want to see.

ChatGPT

Tail sampling: buffer the spans and decide after the outcome is known, keeping everything that errored or exceeded a latency threshold plus a small random slice of the rest. Costs more memory at the collector and is almost always worth it.

You

Small service, limited time. Which one first?

ChatGPT

Logs with structure, because they answer "what happened" for incidents you did not anticipate. Metrics come next for the handful of numbers you would page on. Traces are worth it once a request crosses more than two services.

You

Does this still hold if I am working alone?

ChatGPT

It gets easier, not harder. A fair share of the difficulty is coordination, and that disappears. What remains is remembering your own decisions, which is what writing them down is for.