When Direct Connect goes dark: how our edge network rode out the July 24 AWS us-west-2 outage
By MailChannels | 9 minute read
On July 24, 2026, between 10:55 and 12:12 UTC (3:55 to 5:12 AM Pacific), AWS experienced a networking failure in its us-west-2 (Oregon) region that broke routing between the region and the Seattle Metro. One of the three AWS Direct Connect circuits linking our AWS infrastructure to our global edge nodes, the circuit serving our Seattle node, carried zero traffic for 51 minutes. Our other two circuits absorbed the Seattle node’s share within about a minute, email kept flowing, and nobody at MailChannels had to wake up.
The circuit loss itself caused no customer-visible impact. The broader AWS event also degraded general connectivity into us-west-2 for 20 minutes, which temporarily reduced the volume of mail reaching our infrastructure until AWS restored regional connectivity at 11:15 UTC. Standard SMTP retry behavior absorbed that dip, and our telemetry shows the accumulated backlog draining within about ten minutes of recovery. AWS attributed the failure to networking devices responsible for routing from the region to the Seattle Metro. Nothing in AWS’s post-event summary or in our own telemetry suggests malicious activity, and no email or data was lost.
Automatic failover is easy to claim and harder to demonstrate. So this post shows our work: the architecture that makes rerouting possible, minute-by-minute traffic data from all three circuits during the event, how flows rebalanced once the Seattle path returned, and the one place where our margins were thinner than we would like.
How email leaves MailChannels
Our mail processing infrastructure runs in AWS us-west-2. Email does not leave for the internet from there, though. It egresses from edge nodes we operate around the world, each announcing MailChannels-owned IP space. Sending from IP addresses we control end to end is what lets us protect the sending reputation our customers depend on.
Between AWS and each edge node sits AWS Direct Connect: a dedicated private circuit into the AWS network at a colocation facility, rather than a path across the public internet. On each circuit we run a virtual interface (VIF), a logical link with its own BGP session. BGP, the Border Gateway Protocol, is how routers advertise which destinations they can reach. When a session dies, its routes are withdrawn and traffic converges onto the remaining advertised paths. No human is in that loop, which is the entire point.
Three nodes matter in this story: Seattle, Astute, and TechFutures, each connected over a 1 Gbps Direct Connect virtual interface. The Seattle circuit enters AWS through the Direct Connect location at the Westin Building Exchange (EqSe2) in Seattle. The other two circuits land at different Direct Connect locations. That one topological detail is the whole story of July 24.
We deliberately spread production traffic across all three circuits at all times. In the hour before the event, Seattle carried an average of 415 Mbps, Astute 295 Mbps, and TechFutures 426 Mbps, about 1.14 Gbps in total, with Seattle handling 37% of egress. A standby circuit that never carries traffic is a circuit you cannot trust, so we do not have any.
What AWS reported
AWS’s post-event notes describe the sequence concisely. Connectivity issues to us-west-2 began at 10:55 UTC, caused by networking devices that handle routing between the region and the Seattle Metro. AWS’s engineers were automatically engaged within six minutes, and their mitigations produced initial recovery of general regional connectivity at 11:15 UTC. A brief route reconvergence event between 11:47 and 11:59 UTC caused intermittent connectivity as paths were re-established. Customers whose Direct Connect circuits ran through EqSe2 at the Westin Building Exchange had an extended impact window until 12:12 UTC, when routes for that specific path were fully restored. Connectivity within the region was never affected, and AWS noted that customers connected redundantly through other Direct Connect locations were not impacted.
That last sentence describes our other two circuits exactly. Credit to AWS for fast automated engagement and a clear causal writeup.
What our telemetry shows
The chart below plots the CloudWatch metric aws.dx.virtual_interface_bps_egress for each of the three virtual interfaces: one-minute average egress, measured in the AWS-to-edge direction. The y axis is normalized to an internal baseline volume unit, so it shows relative traffic levels rather than absolute throughput; the dashed line marks nominal circuit capacity on the same scale. During the outage the Seattle interface reported no datapoints because no traffic flowed, so we render that gap as zero. Four isolated single-minute collector gaps are bridged; one of them, at 11:55 UTC on the Seattle interface, lands inside AWS’s reconvergence window and we call it out below.

The break
Seattle’s last full-rate sample was 380 Mbps at 10:55 UTC. The 10:56 sample shows 151 Mbps, a partial minute as the path died underneath established flows. From 10:57 the interface carried nothing.
The same 10:57 minute tells the failover story: Astute jumped to 425 Mbps and TechFutures to 603 Mbps, a combined 1.03 Gbps, about 90% of the pre-event total, flowing on two circuits within roughly sixty seconds of the path loss. BGP withdrew the Seattle routes, the remaining paths absorbed the flows, and the convergence was effectively instantaneous at the resolution of our metrics.
The dip that wasn’t ours
From 10:57 to 11:14 UTC, combined egress averaged 599 Mbps, down 47% from baseline. This was not a failover problem. Our egress mirrors our ingress, and during those minutes AWS’s general connectivity impairment meant less mail was reaching us-west-2 in the first place. Mail transfer is store-and-forward by design: sending systems that could not reach the region queued their messages and retried, exactly as SMTP intends.
The drain
At 11:15 UTC, AWS’s mitigation restored general regional connectivity, and the queued volume arrived all at once on two circuits instead of three. TechFutures reported one-minute averages at or above 1 Gbps for nine consecutive minutes, peaking at a reported 1.16 Gbps, and Astute peaked at 988 Mbps. Combined reported egress touched 2.13 Gbps at 11:17 UTC. Reported values above the 1 Gbps line rate are an artifact of computing rates from per-minute byte counters whose timestamps shift slightly between collection intervals; the honest reading is simply that the circuit was saturated. The shape of the surge, roughly ten minutes of line-rate transfer followed by a settling curve, is consistent with retried and queued mail flushing through the system, though we infer that from the traffic pattern rather than from queue-level measurements.
Once the drain completed, the two circuits settled into a steady two-circuit state. Between 11:15 and 11:46 UTC, Astute averaged 567 Mbps (up 92% from its baseline) and TechFutures averaged 846 Mbps (up 99%). Both circuits roughly doubled their load without intervention.
| Window (UTC) | Seattle | Astute | TechFutures | Combined |
|---|---|---|---|---|
10:00-10:54 (baseline) | 415 Mbps | 295 Mbps | 426 Mbps | 1,136 Mbps |
10:57-11:14 (regional impairment) | 0 | 248 Mbps | 351 Mbps | 599 Mbps |
11:15-11:46 (two-circuit failover) | 0 | 567 Mbps | 846 Mbps | 1,413 Mbps |
12:00-12:15 (restored) | 438 Mbps | 311 Mbps | 469 Mbps | 1,217 Mbps |
One-minute average egress per circuit. The failover-window total exceeds baseline because backlog catch-up overlapped the normal early-morning ramp in volume.
The rebalance
Our Seattle interface began passing traffic again at 11:48 UTC, squarely inside AWS’s 11:47 to 11:59 reconvergence window, which is exactly when you would expect withdrawn routes to reappear. The first minute moved barely a kilobit per second, then 67 Mbps at 11:49 and 390 Mbps by 11:52. The single missing sample at 11:55 is consistent with a brief route flap during reconvergence; the interface was stable from 11:56 onward.
Rebalancing was as automatic as the failover. As the Seattle routes were re-advertised, Astute and TechFutures shed the load they had been carrying, and by 12:00 UTC the three-way split was back to normal: Seattle at 438 Mbps, Astute at 311 Mbps, TechFutures at 469 Mbps, a combined 1,217 Mbps that matches baseline plus ordinary morning growth. Our traffic had fully rebalanced twelve minutes before AWS declared the EqSe2 path restored at 12:12 UTC.
Here is the full sequence in one place:
| UTC | Pacific | Event |
|---|---|---|
10:55 | 3:55 AM | AWS networking failure begins; Seattle circuit loses connectivity; flows shift to Astute and TechFutures within about a minute |
11:01 | 4:01 AM | AWS engineers automatically engaged |
11:15 | 4:15 AM | General regional connectivity restored; queued mail drains over two circuits at line rate for about ten minutes |
11:47-11:59 | 4:47-4:59 AM | AWS route reconvergence; our Seattle interface resumes carrying traffic at 11:48 |
12:00 | 5:00 AM | Traffic split back to normal across all three circuits |
12:12 | 5:12 AM | AWS declares the EqSe2 (Westin Building Exchange) Direct Connect path fully restored |
13:01 | 6:01 AM | AWS marks the event resolved |
Why nobody woke up
The event started at 3:55 AM Pacific and was over before the first coffee of the day. The first human involvement was reading dashboards after the fact. Three properties made that possible.
First, path redundancy across Direct Connect locations, not just redundant circuits. Two circuits terminating at the same facility would both have gone dark on July 24. AWS’s own summary drew the line between impacted and unimpacted customers at precisely this boundary.
Second, live headroom. Because every circuit carries production traffic continuously, we know each path works, and the surviving pair had room to roughly double their load. Failover onto a path you have never exercised is a hypothesis, not a capability.
Third, protocol semantics. SMTP’s store-and-forward and retry behavior converted a 20-minute regional impairment into a ten-minute catch-up surge rather than lost mail. The resilience budget of an email system includes the protocol itself.
We also owe you the caveat. TechFutures spent 12 minutes at or above 85% of nominal capacity, and was saturated for nine of them. This event began during our overnight traffic trough. Had a 51-minute circuit loss coincided with daily peak volume, two circuits might not have covered demand plus backlog, and mail would have queued at the edge. Queued mail is delayed, not lost, but delay is still impact, and we do not get to grade ourselves only on outages that arrive at 3:55 AM.
What we’re changing
We are re-deriving our per-circuit headroom targets against peak volumes rather than average volumes, so that the loss of any single circuit leaves the survivors below saturation at the worst hour of the day, backlog drain included. That review covers both circuit capacity and the weights that govern how traffic is distributed across nodes. We will publish the outcome in a follow-up post.
Takeaways for operators
If you run Direct Connect, the lesson of July 24 is that redundancy is a property of facilities, not of circuit counts. Check whether your “redundant” circuits share a Direct Connect location, because AWS’s impact boundary ran exactly along that line. Keep every failover path carrying production traffic so that convergence is a routine event rather than a first-time experiment. And when you review an incident like this one, measure the drain, not just the outage: the backlog surge is where your remaining capacity is actually tested.
AWS’s full post-event summary is available in the us-west-2 history on their status page and is worth reading as a model of a clear causal chain. If you want to see how our circuits behave day to day, our status page publishes ongoing service history, and this blog will carry the headroom review when it is complete.