Microsoft’s March 2024 Azure outage in South Africa showed why a network can survive a cable failure and still be in a deteriorating state: the February Red Sea cuts had already consumed enough spare capacity that a later West African cable incident and an ordinary router-optic failure pushed services into congestion. For Windows administrators running Azure workloads, Microsoft 365 integrations, or globally distributed applications, the practical lesson is blunt: successful rerouting is the beginning of an incident, not its resolution.

KoreaTechDesk frames the Red Sea cable crisis as a call for operators to plan for the next disruption before repairs from the first are complete. The reporting is directionally right, but Microsoft’s unusually detailed post-incident account supplies the harder evidence. It documents a sequence in which redundancy worked exactly as engineered on February 24, then became materially less useful while damaged capacity remained unavailable.

The error in many outage assessments is treating service availability as the whole resilience story. It is not. A service can remain online while its operational headroom—the capacity and route diversity left for the next failure—has fallen to a level where a routine fault becomes customer-visible.

A cybersecurity operations center monitors global network traffic, connections, and threat alerts on multiple screens.Microsoft’s 77% network-availability low point​

Microsoft’s Azure status history says its South Africa regions were designed around four physically diverse submarine cable systems, with an intended failure mode in which three of four could be lost without customer impact. That is a substantial design target, and the February 24 Red Sea damage initially appeared to validate it.

The February failures affected Asia Africa Europe-1, Europe India Gateway, and the SEACOM/Tata TGN-Eurasia system. TeleGeography made an important correction early in the event: SEACOM and Tata TGN-Eurasia were two names for the same system in the relevant portion of the network, so reports of four separate cable failures overstated the number. Three damaged systems were serious enough; precision matters because an operator’s resilience model depends on independent failure domains, not headlines.

Microsoft says the Red Sea cuts removed east-coast capacity but did not initially affect customers. The company had already begun requesting African capacity additions on February 5 after evaluating geopolitical risk around Red Sea infrastructure. Yet those additions had not entered service when the second event arrived.

On March 14, cuts affecting West African capacity—Microsoft named WACS, MainOne and SAT-3 in its timeline, while regional reporting also identified ACE—further reduced available paths to its South Africa regions. The Internet Society’s post-event report concluded that the West African incident affected four systems, including ACE, likely around a shared failure area off Côte d’Ivoire. That distinction matters: separate cable names did not mean four independent risks when their routes converged near the same location.

Then a line-card optic failed in a Microsoft router within the region. Microsoft described that kind of component failure as routine across its fleet, saying it sees hundreds each day among more than 500,000 network devices. Under normal conditions, customers would never know it happened. With cable capacity already missing, the optic failure removed enough residual headroom on a failover path to cause congestion, latency, and packet loss.

Network availability fell as low as 77% on March 14. Azure Compute, Storage, Networking, Databases, and App Services were affected, as were Microsoft 365 services that depended on calls beyond South Africa. Microsoft shifted capacity from its Lagos edge, changed traffic engineering, failed Azure Front Door out of the South Africa regions, and brought emergency capacity online over March 17 and 18.

The key finding is not that Microsoft lacked redundancy. Its architecture absorbed the first crisis. The finding is that redundancy had become partially spent inventory by the time a much smaller, everyday equipment fault occurred.


Repair time is part of the failure model​

Cable-cut stories often focus on the moment connectivity drops, then move on once traffic takes another route. The repair window receives less attention, even though it determines how long the network must operate in its weakened state.

SEACOM said in March 2024 that permissions for its repair partner could take up to eight weeks because access required regulatory approval. That estimate proved optimistic. Yemen’s telecommunications and transport authorities announced completion of repair work on AAE-1, EIG, and SEACOM in late July 2024, roughly five months after the February faults.

The cause of the Red Sea damage was widely discussed but not conclusively established by all involved parties. Cloudflare reported at the time that the cables were believed to have been cut by the anchor of the Rubymar, a ship damaged in a Houthi missile attack days earlier. What is indisputable is the operational consequence: repair crews had to work amid regional conflict, contested authority, permit requirements, and limited specialist-vessel availability.

Those dependencies mean the phrase “we have backup routes” is incomplete. Operators also need answers to less glamorous questions:

  • How long can the surviving paths carry peak business traffic without violating latency or packet-loss objectives?
  • Which backup paths share a landing station, coastal approach, repair vessel pool, or government approval process?
  • How quickly can contracted emergency capacity be activated, tested, and put into production?
  • Which applications still depend on cross-region API calls even when the primary workload is hosted locally?

Microsoft’s own incident response points toward the right measures. It later added capacity through a new physically diverse cable system, reviewed how quickly it could procure urgent capacity, evaluated a fifth WAN path between South Africa and the United Arab Emirates, and worked to reduce dependencies on international regions by running more services locally.

This is not a purely telecom concern. For an IT team, a cloud region’s healthy status page does not guarantee that international dependencies are healthy. Identity flows, security inspection, SaaS APIs, content delivery, control planes, telemetry, backup replication, and multi-region databases can all turn a constrained WAN into an application incident.

May 2024 proved that overlap was no longer theoretical​

The February damage was still unrepaired when faults involving EASSy and SEACOM disrupted East African connectivity on May 12. Cloudflare Radar observed immediate traffic declines of 10% to 25% in Kenya, Uganda, Madagascar, and Mozambique. Rwanda, Malawi, and Tanzania fell by one-third or more against the prior week’s traffic levels.

That second event did not create the same cloud-service sequence Microsoft documented in March, but it confirmed the broader pattern. A network running with damaged routes is not experiencing one long incident; it is operating under a changed risk profile until capacity is restored.

Cloudflare’s analysis also makes clear that redundancy is uneven. Safaricom and Airtel Kenya said they activated alternative capacity and redundancy measures, while the Communications Authority of Kenya identified TEAMS as an unaffected route used to preserve local traffic flow. Some networks maintained service with degraded performance; others took a far steeper traffic hit.

The difference was not merely how many cable systems appeared on a map. It was whether remaining routes had usable capacity, whether they were genuinely independent from the affected infrastructure, and whether operators could redirect traffic quickly enough.

This is where capacity planning and disaster recovery discipline meet. A failover route that exists but cannot take production traffic at business-hour load is not a recovery path. It is a partial mitigation whose limits need to be measured before an incident.


The ITU has moved beyond cable-counting​

The International Telecommunication Union’s International Advisory Body on Submarine Cable Resilience formally recognized the same problem in its final report approved on July 10, 2026. The ITU identified longer repair times, geographic concentration, physical risks, and countries’ dependence on a small number of systems as central resilience problems.

Its recommendations include faster permitting and repair, stronger government-industry coordination, risk monitoring, preparedness, and greater geographic route diversity. The significance is that the ITU is treating permitting and repair readiness as part of communications resilience rather than administrative afterthoughts.

The body also reported more than 170 cable repairs globally in 2025—close to four per week. That does not mean every repair becomes a public outage. It does mean operators should stop modeling concurrent or overlapping cable issues as freak events.

KoreaTechDesk quotes RETN chief executive Tony O’Sullivan recommending at least four independent routes, preferably five or more, for any geography. That is a company executive’s planning recommendation rather than an industry standard, and RETN naturally highlights its TRANSKZ terrestrial corridor as a Red Sea bypass. Still, the underlying point holds: independence must be judged by route family, not just cable count.

Four submarine systems crossing the same narrow maritime corridor, relying on the same repair jurisdiction, or terminating at the same landing geography do not deliver the resilience implied by “four routes.” A terrestrial corridor may provide valuable separation from maritime failure risks, but it can introduce its own border, power, political, and capacity dependencies. No route deserves to be called diverse until those shared constraints are mapped.

What Windows and cloud teams should change​

Most enterprise IT teams cannot buy a new submarine cable route. They can, however, stop relying on regional availability as their only resilience indicator.

For Azure and Microsoft 365-dependent environments, that means testing whether applications can tolerate a constrained connection to a region rather than only a complete regional outage. It means identifying services that quietly leave a local region for APIs, authentication, security services, management planes, or data replication. It also means demanding clearer answers from carriers and cloud providers about emergency-capacity activation, physical route overlap, and restoration assumptions.

South Korea’s policy research is relevant beyond its borders. A 2024 Science and Technology Policy Institute study characterized submarine cables as part of the global data supply chain and called for changes in governance, maintenance cooperation, and route diversification. The Korea Maritime Institute returned to cable-security policy in March 2026. The broader lesson applies to every cloud-reliant economy: counting international links says little about the risk if those links share the same geography, repair constraints, or political chokepoints.

Microsoft’s March 2024 incident supplied the operational proof. The first cable cuts caused no customer impact; the next set, combined with a normally invisible optic fault, did. Any resilience plan that ends with “traffic was rerouted” leaves out the number that matters most: how much capacity remains before the next fault arrives.


References​

  1. Primary source: koreatechdesk.com
    Published: August 8, 2026 at 10:17 PM UTC