You
Same code, same data, different results on another machine.
ChatGPT
Seed every source of randomness, pin the library versions, and disable nondeterministic kernels on the GPU. Even then expect small differences from floating point order; the goal is results that agree within noise, not bitwise equality.
You
With 800 labelled examples, is an 80/20 split reasonable?
ChatGPT
Use cross validation instead. At 800 examples a single 20% test set is 160 rows, and the variance between splits will exceed the differences you are trying to measure. Five fold gives you a mean and a spread, and the spread is the honest part.