You
With 800 labelled examples, is an 80/20 split reasonable?
ChatGPT
Use cross validation instead. At 800 examples a single 20% test set is 160 rows, and the variance between splits will exceed the differences you are trying to measure. Five fold gives you a mean and a spread, and the spread is the honest part.
You
One class is 3% of the data and oversampling barely helps.
ChatGPT
Change the metric before changing the data. Accuracy is meaningless here; precision-recall AUC and a threshold chosen from the business cost of each error type usually reveal that the model was fine and the decision rule was wrong.