The Internet Report

The Internet Report

By ThousandEyesTechnology
Download on the App Store

The Internet Report episodes

  • Understanding the Recent Workday and Cloudflare Outages | Pulse Update

    Backend-related incidents have been a recurring theme in outages across 2023, caused by everything from data center issues and hardware mishaps to failures at common (shared) services.

    Recently, we saw two examples of these backend issues when data center power problems led to outages at both Cloudflare and Workday.

    Tune in to hear more about what happened at Cloudflare and Workday, as well as our analysis of disruptions at OneLogin and GitLab.

    ———

    CHAPTERS
    00:00 Intro
    01:00 OneLogin Disruption
    05:22 GitLab.com Availability Issues
    09:14 Workday and Cloudflare Outages
    31:16 Get in Touch

    ———

    For more insights, check out these links:

    - The Internet Report: Pulse Update Blog: https://www.thousandeyes.com/blog/internet-report-pulse-update-workday-cloudflare-outages?utm_source=transistor&utm_medium=referral&utm_campaign=na_fy24q2_internetreportpulse23_podcast

    - Interested in more outage analysis? Check out our Internet Outages Timeline, which covers several notable Internet outages and application issues from the past year, along with the lessons they leave: https://www.thousandeyes.com/resources/internet-outages-timeline?utm_source=transistor&utm_medium=referral&utm_campaign=na_fy24q2_internetreportpulse23_podcast

    ———

    Want to get in touch?

    If you have questions, feedback, or guests you would like to see featured on the show, send us a note at [email protected]. Or follow us on X: @thousandeyes

    33 min
  • Halloween Special: Ghosts of Outages Past

    This Halloween, The Internet Report team is sharing some of their most thrilling (and chilling) networking tales.

    Pull up a chair (and a big bowl of your favorite Halloween candy) to hear what happened—and important lessons learned.

    ———

    CHAPTERS
    00:00 Intro

    01:40 Haunting obstacles with a dynamic routing protocol that thwarted crew changes on an oil platform

    10:00 A spooky code base rollout that unleashed memory leak mischief

    18:58 A chilling application rollout that failed to deliver on user expectations around the globe

    29:45 Mysterious application issues that sent shivers down spines, before they were discovered to be caused by a wicked broadcast storm

    42:43 Get in Touch

    ———

    Want to get in touch?

    If you have questions, feedback, or guests you would like to see featured on the show, send us a note at [email protected]. 

    And follow us on X: @thousandeyes

    44 min
  • Insights From Outages at Citibank, DBS, and Other News | Pulse Update

    In recent weeks, back-end infrastructure work and other backend-related issues impacted various online and consumer banking services, including DBS and Citibank in Singapore.

    Simple front-facing customer experiences that we’ve become accustomed to today can often mask considerable complexity on the backend. The service delivery chain of technologies powering the front end often comprises a mix of on-premises assets, cloud services, containers, and APIs.

    A degradation or outage to just one of those components can have massive impact. Depending on the architecture of the app and resilience of the backend, an incident in one part can be routed around in the best case scenario, or take down critical systems for hours in the worst case.

    Tune in to this episode to learn more about how backend changes led to outages at DBS, Citibank, and a number of Japanese banks—and how other backend issues appeared to contribute to a Google Cloud VMware Engine disruption and potentially also a Microsoft Exchange incident.

    For more insights, check out these links:

    - The Internet Report: Pulse Update Blog: https://www.thousandeyes.com/blog/internet-report-pulse-update-dbs-citibank-outages?utm_source=transistor&utm_medium=referral&utm_campaign=na_fy24q1_internetreportpulse22_podcast

    - Explore the Equinix issues that impacted DBS and Citibank: https://ajhrlohbopohbnmekzbcvrbeslqaijfr.share.thousandeyes.com/

    - Interested in more outage analysis? Check out our Internet Outages Timeline, which covers several notable Internet outages and application issues from the past year, along with the lessons they leave: https://www.thousandeyes.com/resources/internet-outages-timeline?utm_source=transistor&utm_medium=referral&utm_campaign=na_fy24q1_internetreportpulse22_podcast

    ———

    CHAPTERS
    00:00 Intro
    00:47 The Download
    04:10 By the Numbers
    06:40 Equinix Chiller Upgrade Leads to DBS, Citibank Outages in Singapore
    23:19 Get in Touch

    ———

    Want to get in touch?

    If you have questions, feedback, or guests you would like to see featured on the show, send us a note at [email protected]. Or follow us on X: @thousandeyes

    25 min
  • Talking Data Freshness + Slack, Cloudflare, and Google Outages | Pulse Update

    Outages and degradations can happen when underlying data isn’t fresh enough. In recent weeks, stale data may have contributed to incidents at both Slack and Cloudflare. 

    Slack began experiencing issues when, by our best guess, its app stopped trusting the freshness of the data in the cache; and, separately, Cloudflare’s 1.1.1.1 DNS resolver ran into some issues related to stale root zone data.

    Watch this Pulse Update episode to hear more about the Cloudflare and Slack outages, and also explore recent disruptions at Google.

    For more insights, check out these links:

    - Explore the Slack outage in the ThousandEyes platform (NO LOGIN REQUIRED): 
    https://apiyhhcphzaowmpqpyxrtdgggadiiujg.share.thousandeyes.com

    - Interested in more outage analysis? Check out our Internet Outages Timeline, which covers several notable Internet outages and application issues from the past year, along with the lessons they leave: https://www.thousandeyes.com/resources/internet-outages-timeline?utm_source=transistor&utm_medium=referral&utm_campaign=na_fy24q1_internetreportpulse21_podcast

    ———

    CHAPTERS
    00:00 Intro
    01:11 The Download
    04:46 By the Numbers
    09:30 Slack Outage: Cached Data Freshness Issues
    23:08 Cloudflare Outage: Resolvers Use Stale Root Zone
    29:55 Get in Touch

    ———

    Want to get in touch?

    If you have questions, feedback, or guests you would like to see featured on the show, send us a note at [email protected]. Or follow us on X (formerly Twitter): @thousandeyes

    31 min
  • Internet Outages: Why One Small Link Can Break the Whole Chain | Pulse Update

    Providing great digital experiences relies on a complex service delivery chain. The past few weeks brought multiple reminders that the root cause of cloud and app disruptions often comes down to one single link in this chain. While the component at issue may appear small, if it’s not functioning normally, the consequences can be significant. 

    Additionally, the impact of a malfunctioning “link” is often intensified by a lack of understanding or visibility into the entire end-to-end service delivery chain, especially in situations where a change is made outside standard operating procedures or pipelines.

    In this episode, explore how this phenomenon appeared to play out recently when .au domains failed to resolve, as well as during disruptions at Salesforce and Microsoft Azure.

    For more insights, check out these links:

    - Explore the .au incident in the ThousandEyes platform (NO LOGIN REQUIRED): 
    https://ajdaombojgmvnbvsclhtvgmrskaicjvo.share.thousandeyes.com/

    - Internet Outages Timeline: https://www.thousandeyes.com/resources/internet-outages-timeline?utm_source=transistor&utm_medium=referral&utm_campaign=internetreportpulseep20

    ———

    CHAPTERS
    00:00 Intro
    00:54 The Download
    06:01 By the Numbers
    08:23 .au Domains Fail to Resolve
    15:42 PlayStation Network Disruption
    21:10 Get in Touch

    ———

    Want to get in touch?

    If you have questions, feedback, or guests you would like to see featured on the show, send us a note at [email protected]. Or follow us on X (formerly Twitter): @thousandeyes

    23 min
  • Data Center Disruptions, Square Down, and More News | Pulse Update

    In a world that operates at “hyperscale,” the potential for hyperscale-sized problems is also very real. The measure of a good provider—and a well-engineered system—is how well they handle these anomalous conditions and minimize disruption.

    During recent weeks, some of these hyperscale-sized outages hit, including data center-focused disruptions that impacted companies like Square, Oracle OCI, NetSuite, and Microsoft Azure. 

    Tune into this Pulse Update episode to go under the hood of these outages and discover how the companies responded—and important lessons learned.

    For more insights, check out these links:

    - The Internet Report: Pulse Update Blog: https://www.thousandeyes.com/blog/internet-report-pulse-update-square-down?utm_source=transistor&utm_medium=referral&utm_campaign=internetreportpulseep19

    - Explore the Square outage in the ThousandEyes platform (NO LOGIN REQUIRED): 
    https://akinmwcoyjwhlwmhnykikzqkxcltwasv.share.thousandeyes.com


    ———

    CHAPTERS
    00:00 Intro
    00:59 The Download
    04:33 By the Numbers
    09:02 Square Outage
    23:08 Oracle OCI, NetSuite, and Microsoft Azure Outages
    32:19 Get in Touch

    ———

    Want to get in touch?

    If you have questions, feedback, or guests you would like to see featured on the show, send us a note at [email protected]. Or follow us on X (formerly Twitter): @thousandeyes

    34 min
  • Disruptions at Slack and X + Thoughts on “Take Twos” | Pulse Update

    An outage occurs, a change is rolled back, and everything stabilizes. But what happens when the change is attempted a second time?

    These second tries often go much more smoothly. While another outage might still occur during this “take two,” the impact is usually far less severe. The engineering team has learned from what went wrong the first time and is ready to stop at the first hint of trouble. 

    Slack recently experienced a pair of disruptions that appear to illustrate this “take two” scenario: a longer disruption resulting from a routine database cluster migration, followed by a much shorter outage a few weeks later that also involved database work, potentially indicative of related work that went more smoothly.

    And for more insights, check out these links:

    - The Internet Report: Pulse Update Blog: https://www.thousandeyes.com/blog/internet-report-pulse-update-slack-x-outage?utm_source=transistor&utm_medium=referral&utm_campaign=internetreportpulseep18

    - Explore the Slack and X disruptions in the ThousandEyes platform (NO LOGIN REQUIRED): 
    Slack: https://afkmcwbeszwdtqqpvouwgjolywiugryx.share.thousandeyes.com/
    X: https://adcsnhfupsardmzyocrxqdcvriengkew.share.thousandeyes.com

    ———

    CHAPTERS
    00:00 Intro
    00:47 The Download
    04:06 By the Numbers
    05:41 Slack Disruptions
    09:25 X Disruptions
    20:20 Get in Touch

    ———

    Want to get in touch?

    If you have questions, feedback, or guests you would like to see featured on the show, send us a note at [email protected]. Or follow us on X (formerly Twitter): @thousandeyes

    22 min
  • An August Slack Outage and Why Context Matters | Pulse Update

    Context matters when working on a distributed web-based application or service where everything is linked and dependent on each part functioning correctly. 

    It’s all too easy for one team to make a change that unexpectedly affects something another team is working on. Or the combined impact of both changes may also accidentally break something.

    To avoid such mishaps, teams should cut back on silos as much as possible.

    However, it’s hard to completely eliminate siloed operations or decision-making. But the potential negative effects of silos can be reduced if each team has a view of the end-to-end service that’s tailored to their specific area or domain—that is, presented to them in a context that they understand.

    Tune in to this week’s episode to learn more about mitigating silos and also explore lessons from recent disruptions at Slack, Spotify, and Wells Fargo.

    And for more insights, check out these links:

    - The Internet Report: Pulse Update Blog: https://www.thousandeyes.com/blog/internet-report-pulse-update-slack-outage?utm_source=transistor&utm_medium=referral&utm_campaign=internetreportpulseep17

    - Explore the Slack disruption in the ThousandEyes platform (NO LOGIN REQUIRED): https://apiijiaoljxvlpnzjkxwwlcqmmcovppx.share.thousandeyes.com/

    ———

    CHAPTERS
    00:00 Intro
    00:49 The Download
    04:35 By the Numbers
    06:55 Slack Disruption
    27:18 Spotify Disruption
    33:15 Get in Touch

    ———

    Want to get in touch?

    If you have questions, feedback, or guests you would like to see featured on the show, send us a note at [email protected]. Or follow us on X (formerly Twitter): @thousandeyes

    35 min
  • SharePoint Outage and Security Certificate Considerations | Pulse Update

    In an end-to-end service delivery chain, isolated changes can have broad consequences. This played out recently when an erroneous SSL certificate change at Microsoft appeared to cause a SharePoint Online and OneDrive for Business outage.

    While this incident definitely underscores the importance of valid security certificates, it’s also a reminder of what can happen when even one component in an end-to-end service delivery chain experiences issues. Every component needs to work in sync to maintain the service’s availability. 

    As a result, all changes, especially manual ones, should be made with care and teams should have a deep understanding of every dependency and interconnection within their service delivery chain.

    Watch this week’s episode to learn more and explore other recent outages that impacted Slack, Starbucks, and NASA.

    And for more insights, check out these links:

    - The Internet Report: Pulse Update Blog: https://www.thousandeyes.com/blog/internet-report-pulse-update-sharepoint-outage?utm_source=transistor&utm_medium=referral&utm_campaign=internetreportpulseep16

    - Explore the OneDrive outage in the ThousandEyes platform (NO LOGIN REQUIRED): 
    https://arcbfeptuostafskynuemdpgwcyerodc.share.thousandeyes.com

    ———

    CHAPTERS
    00:00 Intro
    00:43 The Download
    04:03 By the Numbers
    06:09 SharePoint Outage
    21:47 Slack Outage
    25:31 Get in Touch

    ———

    Want to get in touch?

    If you have questions, feedback, or guests you would like to see featured on the show, send us a note at [email protected]. Or follow us on Twitter: @thousandeyes

    27 min
  • Azure Disruption, Meta App Issues, and Navigating Edge Cases | Pulse Update

    Let’s face it. Not every contingency can be planned for. Sometimes an outlier scenario pops up and causes an unexpected outage or disruption.

    Over the past few weeks, multiple companies appeared to be impacted by such edge cases: Azure; GitLab; and Meta’s WhatsApp, Facebook, Instagram, and Threads—its newest addition.

    Tune into the latest Pulse Update episode to learn more about what happened during these disruptions and why robust visibility is so important for navigating unexpected outlier scenarios.

    And for more insights, check out these links:

    - Internet Report: Pulse Update Blog: https://www.thousandeyes.com/blog/internet-report-pulse-update-azure-disruption?utm_source=transistor&utm_medium=referral&utm_campaign=InternetReportPulseEp15

    - Explore the Azure disruption in the ThousandEyes platform (NO LOGIN REQUIRED): https://augfulkplwamllucivbbxxahisxddgay.share.thousandeyes.com

    - Cloud Performance Report: https://www.thousandeyes.com/resources/cloud-performance-report-2022?utm_source=transistor&utm_medium=referral&utm_campaign=InternetReportPulseEp15

    ———

    CHAPTERS
    00:00 Intro
    00:41 The Download
    03:34 By the Numbers
    05:26 Azure Disruption
    12:01 GitLab Outage
    18:20 Get in Touch

    ———

    Want to get in touch?

    If you have questions, feedback, or guests you would like to see featured on the show, send us a note at [email protected]. Or follow us on Twitter: @thousandeyes

    20 min

About The Internet Report

From the publisher's feed

This is The Internet Report, a podcast uncovering what’s working and what’s breaking on the Internet—and why.

More shows like The Internet Report

The Daily by The New York Times

The Daily

111,868 Listeners

The Hedge by Russ White

The Hedge

18 Listeners

The Art of Network Engineering by Andy and Friends

The Art of Network Engineering

84 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

15,915 Listeners