Customer support · 2026
Putting a SIP layer in front of the dialers so maintenance stops costing calls
Every patch meant a maintenance window, and every maintenance window meant lost dialing. Moving signalling off the dialers made nodes disposable and upgrades boring.
- Maintenance windows needed
- 0
- Rolling patches, business hours
- Calls dropped per node drain
- 0
- Verified under live traffic
- Carrier failover
- Tested
- Each route failed over deliberately
- Time to patch the cluster
- 2 wks → 1 day
- No coordination overhead
The situation
Signalling and media both terminated on the dialer nodes, so taking any node out of service dropped its live calls. Every upgrade therefore needed an agreed window, and windows kept getting deferred.
Deferred patching is how platforms end up years behind. By the time I saw it, the cluster was far enough back that catching up in one jump was itself the risk.
Carrier failover existed on paper. It had never been tested by actually failing a carrier over, which usually means it does not work.
What I did
- 01
Introduced a SIP proxy layer in front of the dialer cluster so signalling scales and fails over independently of media.
- 02
Moved registration and edge handling to the proxy, letting individual dialer nodes drain and return without dropping live calls.
- 03
Configured dispatcher-based failover across backends, then tested it the only way that counts — by taking nodes out during business hours.
- 04
Rebuilt carrier failover at the SIP layer and verified it by failing each carrier over deliberately under real traffic.
- 05
Established a rolling patch process the client’s own team runs, without booking a window.