Estimating how well text will compress

Back to list2026-08-09
You

Rough expectation for compressing JSON logs?

ChatGPT

Repetitive JSON commonly reaches 8 to 12 times with zstd, because keys repeat on every line. If you are seeing less than three, the payload is probably already encoded or encrypted somewhere inside, which is worth knowing for reasons beyond storage.

You

gzip, zstd or xz for archives I keep for years?

ChatGPT

The tradeoff in one table:

gzipzstdxz
Compressionbaselinebetterbest
Speedfastfastestslow
Decompressfastfastestmoderate
Ubiquityeverywhererecentcommon

For long-term archives the deciding factor is not ratio but whether the format opens without installing anything in ten years, which still favours gzip for the outer layer.

You

Does this still hold if I am working alone?

ChatGPT

It gets easier, not harder. A fair share of the difficulty is coordination, and that disappears. What remains is remembering your own decisions, which is what writing them down is for.