Long-lived connections die after 60 seconds.
Something in the path has an idle timeout, and the fix is keepalives shorter than it. Set TCP keepalive on the socket and, for HTTP, send a periodic ping frame. Raising the timeout on your side alone never helps, because the proxy is the one hanging up.
Requests hang for exactly five seconds sometimes.
A five second hang is almost always a DNS timeout, often IPv6 lookups failing before the IPv4 retry. Check with dig +trace and compare getent hosts to a direct query. In containers the usual culprit is a search domain list that turns one lookup into five.
An old device cannot connect since we tightened TLS.
Check the cipher suite overlap first:
openssl s_client -connect example.com:443 -tls1_2 -servername example.com
The honest answer is often that the device cannot be supported without weakening security for everyone, which is a decision to make explicitly rather than by config drift.
Is there a simpler version that gets most of the benefit?
Yes: do the first step, skip the automation, and revisit in a month. Most of the value is in the first step, and most of the cost is in making it repeatable before you know it is right.