You
One class is 3% of the data and oversampling barely helps.
ChatGPT
Change the metric before changing the data. Accuracy is meaningless here; precision-recall AUC and a threshold chosen from the business cost of each error type usually reveal that the model was fine and the decision rule was wrong.
You
Are embeddings worth it for searching 5000 documents?
ChatGPT
Often not. Lexical search handles exact terms, names and identifiers, which is most of what people type. Embeddings pay off for paraphrase-heavy queries, and the honest setup is hybrid, with lexical results ranked first when the query looks like an identifier.
