Classic overfitting or something else?
Classic overfitting if the curves separate smoothly. If validation loss jumps around, suspect the split instead: leakage between sets, or a validation set too small to be stable. Plot both losses per epoch before touching regularisation.
Same code, same data, different results on another machine.
Seed every source of randomness, pin the library versions, and disable nondeterministic kernels on the GPU. Even then expect small differences from floating point order; the goal is results that agree within noise, not bitwise equality.
With 800 labelled examples, is an 80/20 split reasonable?
Use cross validation instead. At 800 examples a single 20% test set is 160 rows, and the variance between splits will exceed the differences you are trying to measure. Five fold gives you a mean and a spread, and the spread is the honest part.
Is there a simpler version that gets most of the benefit?
Yes: do the first step, skip the automation, and revisit in a month. Most of the value is in the first step, and most of the cost is in making it repeatable before you know it is right.