I keep hearing p-values are misused. What do they mean?
The probability of data at least this extreme if the null hypothesis were true. That is all. It is not the probability the hypothesis is false, and a small p with a tiny effect size usually means a large sample rather than an important finding.
How many users do I need for an A/B test?
It depends on the effect you would act on, not the effect you hope for. Decide the smallest difference worth shipping, then compute the size for that. Running until significance appears is the most common way to get a result that will not replicate.
Does this still hold if I am working alone?
It gets easier, not harder. A fair share of the difficulty is coordination, and that disappears. What remains is remembering your own decisions, which is what writing them down is for.
