Our metrics cost exploded after adding a user id label.
That is the mechanism: every distinct label value creates a series. User ids, request ids and full URLs never belong in labels. Put them in logs and traces, where high cardinality is the point rather than the problem.
Head sampling drops exactly the requests I want to see.
Tail sampling: buffer the spans and decide after the outcome is known, keeping everything that errored or exceeded a latency threshold plus a small random slice of the rest. Costs more memory at the collector and is almost always worth it.
Small service, limited time. Which one first?
Logs with structure, because they answer "what happened" for incidents you did not anticipate. Metrics come next for the handful of numbers you would page on. Traces are worth it once a request crosses more than two services.