From about 12:00 to 13:05 UTC, some clients got connection timeouts to their Redis databases. Affected were databases hosted in fra or gig, and clients whose connections were routed through fra or gig, even if their database is hosted in another region.
During a planned rolling upgrade, replicas in fra and gig did not finish draining and were left out of service. They were returned to Fly routing before they were ready to accept connections. Connections routed to them then timed out.
We will be adding tooling to return a replica to service safely after maintenance.
We will be adding safety checks to our upgrade process.
We will make the upgrade procedure more resilient to errors like this one.