Why my search index got 19 MB and how to keep it fast

Back to list2026-03-02
You

Full text search over my archive means shipping a huge index to the browser. How do people keep this usable?

ChatGPT

Split it in two. The list view needs titles, dates and counts, which is small and loads immediately. The full text is only needed the moment someone types, so it can be a separate file fetched on first keystroke.

The difference is the whole user experience: 0.2 MB before interaction, and the remaining megabytes only for people who actually search.

You
index-size-growth.png
ChatGPT

The flat line is the lazy version. It grows too, but nobody pays for it until they use it, and by then they have already decided the archive is worth searching.

You

Any trick for the search itself, or is substring matching enough?

ChatGPT

For an archive this size, plain substring matching over a normalised copy is enough and has one large advantage: no index build step, no stemming surprises, no ranking to explain. The part worth spending effort on is showing the matching sentence under each hit, because that is what turns a list of titles into an answer.