Making sense of an old research folder

Back to list2026-04-27
You

I have 4000 files from a project I stopped in 2021. No naming convention. Where do I even start?

ChatGPT

Do not open them in order. Build an inventory first, then decide:

find . -type f -printf "%s\t%TY-%Tm\t%p\n" \
  | sort -rn | head -50 > biggest.tsv
find . -type f | sed 's/.*\.//' \
  | sort | uniq -c | sort -rn

Size and extension together usually reveal that 90% of the bytes are three or four kinds of artefact you can decide about in one pass, and the actual thinking lives in a few dozen small text files.

You

Half of them are duplicates from different backups.

ChatGPT

Hash first, compare second:

find . -type f -size +1M -exec sha256sum {} + \
  | sort | uniq -w64 -D > duplicates.txt

Only after the list exists should you delete anything, and delete by moving to a quarantine folder you erase in a month. Every deduplication story that ends badly skipped that step.