You
What is a sane rate limit design for a public API?
ChatGPT
Token bucket per key and per IP, with the limit expressed in the response headers so clients can behave. The subtle part is what you do on limit: a 429 with Retry-After is cooperative, silently dropping is what turns one misbehaving client into a retry storm.