Same code, same data, different results on another machine.
Seed every source of randomness, pin the library versions, and disable nondeterministic kernels on the GPU. Even then expect small differences from floating point order; the goal is results that agree within noise, not bitwise equality.
With 800 labelled examples, is an 80/20 split reasonable?
Use cross validation instead. At 800 examples a single 20% test set is 160 rows, and the variance between splits will exceed the differences you are trying to measure. Five fold gives you a mean and a spread, and the spread is the honest part.
One class is 3% of the data and oversampling barely helps.
Change the metric before changing the data. Accuracy is meaningless here; precision-recall AUC and a threshold chosen from the business cost of each error type usually reveal that the model was fine and the decision rule was wrong.