You
With 800 labelled examples, is an 80/20 split reasonable?
ChatGPT
Use cross validation instead. At 800 examples a single 20% test set is 160 rows, and the variance between splits will exceed the differences you are trying to measure. Five fold gives you a mean and a spread, and the spread is the honest part.
You
One class is 3% of the data and oversampling barely helps.
ChatGPT
Change the metric before changing the data. Accuracy is meaningless here; precision-recall AUC and a threshold chosen from the business cost of each error type usually reveal that the model was fine and the decision rule was wrong.
You
What would you skip entirely?
ChatGPT
The reporting layer, until something forces it. It is the part that feels productive to build and the part nobody opens twice.