Failure is normal
Phones move between Wi-Fi and mobile data, pass through areas of weak signal, join networks with captive sign-in pages and sit behind congested links. None of this is exceptional. A system that treats a dropped connection as a rare event will behave badly in daily use.
Recovery should be designed in from the start, not added after the first complaint.
A simple loop: detect, diagnose, act, verify
- Detect. Notice trouble through timeouts, failed handshakes or rising delay, while avoiding false alarms from a single slow response.
- Diagnose. Work out which layer is affected before reacting, because the right response depends on it.
- Act. Retry, switch to another resolver, rebuild a tunnel or move to another network path.
- Verify. Confirm that the connection actually works again, instead of assuming that the action succeeded.
Principles that keep recovery safe
- Back off with jitter. Retrying instantly and constantly drains battery and can overload servers. Waiting a little longer each time, with random variation, avoids both.
- Avoid flapping. If a system switches state at the first sign of change, it will oscillate. Requiring a condition to hold for a short time before acting keeps it stable.
- Bound the effort. Recovery should have limits on time and battery so that it never becomes the problem.
- Be honest in the interface. A message like connection unstable, retrying tells the person what is happening instead of implying that everything is fine.
The hardest decision: fail open or fail closed
When a protection layer cannot work, should traffic be blocked or allowed? Blocking keeps people private but can leave them offline. Allowing keeps them connected but unprotected. There is no universal answer. The right choice depends on context, and it should be made deliberately and communicated clearly, not left to chance.
Keeping sessions alive across changes
Many connections break when a device changes network because they are tied to the addresses at each end. Newer transport designs, such as QUIC, can carry a session across a change of network, which is a good example of recovery handled at the protocol level rather than patched on top.
KEY TAKEAWAYS
- Treat failure as normal and design recovery in advance.
- Use a loop of detect, diagnose, act and verify.
- Back off, avoid flapping and limit effort.
- Decide fail-open or fail-closed on purpose, and tell people what is happening.