Are embeddings worth it for searching 5000 documents

Back to list2026-03-12
You

Are embeddings worth it for searching 5000 documents?

ChatGPT

Often not. Lexical search handles exact terms, names and identifiers, which is most of what people type. Embeddings pay off for paraphrase-heavy queries, and the honest setup is hybrid, with lexical results ranked first when the query looks like an identifier.

You

Classic overfitting or something else?

ChatGPT

Classic overfitting if the curves separate smoothly. If validation loss jumps around, suspect the split instead: leakage between sets, or a validation set too small to be stable. Plot both losses per epoch before touching regularisation.