GitHub Outage Tracker: Is GitHub Cooked?

(isgithubcooked.com)

134 points | by toomanyrichies 4 hours ago

18 comments

  • kashnote 2 hours ago
    I think we need to have a little more sympathy for GitHub. You could justify the jabs when we could all blame any outage on the migration to Azure, but then they shared numbers around the scale they're dealing with now that everyone is constantly building and pushing with AI.

    I think it's commendable that they're not limiting access to the site or (intentionally) throttling newcomers. Yes, they need to get this figured out, but a little sympathy goes a long way. I personally wish them the best and hope their on-call people can go back to getting normal amounts of sleep soon.

    • gogobio 2 hours ago
      Sympathy? It's a Microsoft company that is being ran with a consistency of a startup in early seed rounds. Their downtime is abhorrent and unacceptable as far as enterprise goes. Their engineers look like absolute amateurs allowing for such low class work it results in their customers experiencing industry leading downtime.
      • __MatrixMan__ 2 hours ago
        Sympathy negated by their 2026 Pwnie award: "Lamest Vendor"

        https://this.weekinsecurity.com/microsoft-wins-lamest-vendor...

      • tyre 17 minutes ago
        This is uncharitable, rude, and pretty baseless. Unless you have a lot of direct, personal information about GitHub engineers, they’re dealing with a huge spike in traffic.

        Is their uptime acceptable? No. But personal attacks aren’t necessary or constructive.

      • zvmaz 20 minutes ago
        > Their engineers look like absolute amateurs allowing for such low class work it results in their customers experiencing industry leading downtime.

        I gather that you have intimate and deep knowledge on the teams and the problems they try to solve there.

        • californical 13 minutes ago
          If I own a restaurant and buy bread from a supplier — BreadHub.

          And 99% of the bread that I get is good but 1% of the loaves, they forgot to add flour. Consistently, for years, they always have loaves missing a key ingredient that I still end up paying for.

          I can be pretty sure that BreadHub have a pretty major internal issue, and should probably be questioning their competence, regardless of their “scale”, and without any knowledge of the “problems they’re solving”

          • tyre 10 minutes ago
            I’m not saying users don’t have a right to be pissed or aren’t justified in looking at other options.

            I’m saying that GH is operating at a huge scale with (probably) lots of technical debt and a forced migration to new infrastructure.

            I would be (and am) highly critical of leadership. I’m not going to make strong assertions about ICs without knowing their context. I’ve worked at a company with a sterling reputation for engineering excellence where brilliant ICs were kneecapped by poor leadership.

            I think a lot of us, at one point or another in our careers, have worked with potato leadership that can be short-sighted or political. It isn’t a comment on the engineers.

      • javier2 1 hour ago
        and it has been so damn slow for a number of years now, it drives me insane.
    • sandeepkd 48 minutes ago
      I think somewhere down the line you have an answer to problem itself. Its for a while that companies are not into the business of sympathies, we should probably not get into that side for the topic.

      1. Github has enterprise users who paid for the service, their day job requires Github to be available and working

      2. Github has generous free tier which is the one which is exploring a lot more with the AI generated code.

      It is a complexity in itself but the traffic should have been separated, the free users should not be allowed to bring down Github for enterprise customers (Just to clarify, I am free user myself). And if they do not have capacity it would have been perfectly fine to push back or throttle new users/repositories.

      • tyre 13 minutes ago
        I agree that they should reserve capacity for enterprises and paying users.

        But I think they’re in a tough spot. GitHub has historically been a huge supporter, proponent, and provider for open source projects. Engineers are difficult customers, to say the least, and the community would likely freak tf out of the segmented traffic.

        The logical, pragmatic, and justifiable answer doesn’t always align with your market.

      • IcyWindows 16 minutes ago
        Do we have a breakdown on how much traffic is free users vs paid users?
        • reassess_blind 0 minutes ago
          I’d be surprised if paid traffic was more than 1%.
    • herpdyderp 2 hours ago
      > I think it's commendable that they're not limiting access to the site or (intentionally) throttling newcomers.

      I don't think this is commendable at all. I give GitHub a lot of money and I'm tired of it being wasted with downtime.

      • jjice 2 hours ago
        It's a shame that the GitHub org that we use at my job that we pay a lot of money for gets affected the same way my personal nonsense does.

        I don't know the architecture or any of that, but I feel like there could be (and it's not like they would've really known this until the last year or two with the massive spike) separate infrastructure for paid users/orgs vs free the same way they make the distinction with enterprise.

        I get the massive load changes that they are under over the last two years, but why does a bunch of vibe coded slop take down the same resources that my company pays for every single month and has for years? I imagine properly splitting that out would be an absolute headache and not worthwhile for them vs stabilizing the rest of the service, but damn it sucks when I get blocked at work because GH is down.

        • tempest_ 1 hour ago
          Our GitlabCE instance has been sitting in the racks for nearly two years with almost 100% uptime running on 10 year old xeons that have long since paid for themselves.

          Of course it is not free of all management but for our use case it is working.

          There are hiccups with the CI runners from time time but nothing major and we have another machine in another rack that serves as a backup which can be brought up ~< 20 minutes.

          I know companies have long since tossed their expertise for hosting their own stuff in favour of SaaS but at some point its hard to beat the up time of a single machine.

          • Uvix 42 minutes ago
            We have the expertise. We don’t have the interest in fighting our internal “security” teams to keep running VMs.
    • mronetwo 2 hours ago
      We have a business relationship so no we shouldn’t have any sympathy. They sell a service and they’re failing to provide it.
      • amelius 2 hours ago
        Why? If we have sympathy for Apple then we can have sympathy for Microsoft ...
        • mirashii 1 hour ago
          You're the only mention of Apple in this thread, and I don't see what they have to do with it.
        • pimeys 1 hour ago
          Wait, I don't sympathize Apple at all... Or any other American corporation.
    • bluerooibos 11 minutes ago
      > have a little more sympathy for the billion dollar company

      No.

      Their leadership went all in on AI and in the last blog post essentially admitted some missing test coverage for a critical path. Time to learn lessons and fix your vibe coded shitslop and stop using "user graph go up" as some kind of excuse.

    • padjo 2 hours ago
      I don't typically have sympathy for businesses that fail to deliver a service as advertised.

      I can have sympathy for the humans caught in the crossfire but only managing one nine of availability on a commercial service is not acceptable.

    • alwa 10 minutes ago
      For that matter, I admire their relative transparency about their incidents. I can think of other big players who will cheerfully show green statuses across the board while everyone can see that their pants are down...
    • VCFundedGenYer 1 hour ago
      "Leave the big trillion dollar corporation alone" is not the right idea.
    • jraph 1 hour ago
      They did that to themselves, and doubly so.

      They decided they needed to capture the whole open source ecosystem by turning open source work into social networking... on a proprietary platform (because open source is great, especially when it's others' software). That was before they joined Microsoft.

      And then Microsoft pushed AI everywhere, including on GitHub itself with copilot.

      I would have liked if they had left the open source projects alone and didn't create that FOMO for not using them.

      I have no sympathy.

      • kstrauser 57 minutes ago
        > And then Microsoft pushed AI everywhere, including on GitHub itself with copilot.

        Ding ding ding, we have a winner. I like AI. I work for an AI company. Still, Microsoft aggressively pushed GitHub users toward Copilot. They don't get to do that and complain about increased volume from AI-generated changes.

        No Copilot + reasonable operation: the way things were

        Copilot + reasonable operation: Well done!

        No Copilot + being overwhelmed by AI commits: Sympathy.

        Copilot + being overwhelmed by AI commits: "Where did that petard come from that's hoisting us?"

    • dpz 2 hours ago
      Maybe for a free account.

      But we pay enterprise license and GitHub is a big dependency in our software flow.

      If this continues to be a problem as an enterprise product they need to do something. Otherwise theyre are going to to start losing business

    • niltecedu 1 hour ago
      idk man, we pay stupid amounts of money to microsoft, we are an enterprise customer, expecting better availablity compared to my laptop isnt really a high bar.
    • JDups 2 hours ago
      Constant building with AI is something that they (Microsoft) promote and are heavily invested in.
    • mcmcmc 53 minutes ago
      It’s commendable that they let people give them training data for free?
    • bamboozled 1 hour ago
      No, we don’t.
    • 0xbadcafebee 47 minutes ago
      Our company pays for GitHub. We're paying for a broken product that stops our work. I don't have sympathy for the companies whom I pay for a product and give me broken shit in return. This is entirely preventable and their own fault.

      A restaurant makes pizzas. They suddenly get 100x more popular. They can't make 100x more pizzas. But they are still taking orders from 100x more people. Not only are they not getting enough pizzas delivered that they took orders for, but in their rush to make and deliver more pizzas, they set the kitchen on fire, which makes an even longer wait for pizzas.

      When the pizza you ordered doesn't get delivered, do you have sympathy for the restaurant? Or do you tell them to stop taking orders they can't fill and try not to set the kitchen on fire?

      Now consider the pizza restaurant has 21 billion dollars in cash, is taking your money, and not giving you pizza.

    • theideaofcoffee 1 hour ago
      Give me a break. Sympathy? For microsoft? That might have flown when github was like seven people, but they have nearly unlimited resources to make it better. They're just choosing not to. Let's talk contracts and money before we pull the sympathy card.

      I used to be on-call in a high-traffic environment where single customers pushed more bits than entire nations. I chose the role. I didn't want people's sympathy, if anything, I wanted them to complain to management.

      If it gets too bad they can quit. Maybe that would be for the best, just wear the thing down until it outright fails and no one wants to touch it. One less bullshit service sucking all of the oxygen out.

  • xyzzy_plugh 6 minutes ago
    Is GitHub Cooked? Decidedly yes.

    Migrating to something else has been raised as a concern during every engagement with clients and prospective clients in the past year.

    "Do you use GitHub?"

    "Yes though we'd like to move to something else, we just don't know what yet."

    As soon as the next big thing shows up, they're done. And they know it.

  • joshuahedlund 20 minutes ago
    Given that the outages are caused by record traffic, so far they are only cooked in the Yogi Berra sense of “Nobody goes there anymore, it’s too crowded”
  • _heimdall 15 minutes ago
    I feel for the team having to deal with these problems at github today. Its also beyond me how those in charge for years wouldn't have seen this coming.

    Github was bought by Microsoft, whether they wanted to acknowledge that internally or not. Microsoft went deep on LLMs, and specifically on LLMs for coding use cases. They must have recognized that LLM generated code and PRs would effectively DDoS GitHub.

    I can only assume they simply didn't care, likely driven by greed.

  • JeremyHerrman 2 hours ago
    > "GitHub has had 1125 incidents since February 2016, implying a monthly incident rate of 24"

    1125 incidents / 126 months ≈ 8.9 incidents per month, not 24

    still terrible, but why such an obvious error in the first sentence...

    • graypegg 1 hour ago
      Ahhh, I think the author mixed up two values here. That value seems to actually be the average over the past 3 months.

          const incidents = e.detail.incidents;
          
          // ...snip...
          
          const now = new Date();
          const threeMonthsAgo = new Date(now);
          threeMonthsAgo.setMonth(threeMonthsAgo.getMonth() - 3);
          
          // ...snip...
          
          var currentFreq = recentIncidents.length / 3; // <- We out here, smoking these guns with our homeboy Claude
          
          // ...snip...
          
          var earliest = null;
          for (var j = 0; j < incidents.length; j++) {
            var dd = new Date(incidents[j].started_at); // <- eventually incidents[j].started_at is "2016-03-01T07:07:37.000Z"
            if (!earliest || dd < earliest) earliest = dd;
          }
          
          // ...snip...
          
          document.getElementById('n-since').textContent = earliest
            ? earliest.toLocaleDateString('en-US', { month: 'long', year: 'numeric' })
            : '?';
          document.getElementById('n-rate').textContent = Math.round(currentFreq * 10) / 10;
      
      
      #n-since is going to be either march or feburary. It'll change depending on your timezone because JS's Date object always shifts the date around to match the same instant but in the system's timezone.

      #n-rate has nothing to do with the #n-since month, it's just the last trailing 3 months. And even then, it's sort of underbaked? It's moving the date back by 3 calendar months not taking into account differing numbers of days, so it'll under-report short months.

      I wouldn't trust the stats here.

      Edit: whoops, author updated the template while I was writing this! It now says "Over the last 3 months", though that's still calendar months.

    • 6LLvveMx2koXfwn 1 hour ago
      Not sure whether it has been updated since your comment, but the sentence now reads:

          GitHub has had 1125 incidents since March 2016. Over the last 3 months, they've averaged 24 incidents per month
      
      edit: although they also have 1.2 days of downtime (in a day) for their 'worst days' of downtime table, which suggests some auto number crunching is not working as expected.
      • gen220 1 hour ago
        Yes I tweaked it! The number and copy were mismatched and are no longer!

        That worst day is likely an overlapping incidents accounting issue; I tried to account for overlapping incidents in another view but probably failed to port it over there.

        Should be fixed soon!

    • stevage 1 hour ago
      > GitHub has had 1125 incidents since March 2016. Over the last 3 months, they've averaged 24 incidents per month (↓ 5% vs prev 3mo).

      Looks like they fixed it already

    • 4petesake 1 hour ago
      Prob used Copilot to write the excel formula...
  • Fuzzwah 53 minutes ago
    Near the end of the 8.5 years that I worked at GitHub as an enterprise support engineer, I asked in an all hands if a "GitHub Classic" product had been considered. Much like World of Warcraft Classic, I imagined it would be a rewrite focused on matching the simpler feature set of the past.

    I was basically given the same response that blizzard gave that question; "you think you want that but you don't".

  • bushbaba 3 hours ago
    could have been a page with a static 'Yes' and a significant portion of time it'd be accurate.
  • 404mm 2 hours ago
    If backend GitHub services are anything like GHES then I’m surprised it even managed to scale this much.
  • nightpool 2 hours ago
    Getting rid of Actions and Copilot and other secondary services almost halves Github's incident rate: https://i.imgur.com/XPcMIFr.png

    I'm a big fan of Github Actions and I think people are often a little too harsh on it, but it's clear that it's sad that it's come at such a high cost to the platform's stability

    • Normal_gaussian 1 hour ago
      Actions aren't secondary to most paid GH users; and if they are down it usually means no deployments and no tests, which can often halt work.
    • isityettime 1 hour ago
      GitHub Actions is really, really badly designed. The security model is fundamentally broken, the YAML hell is as bad as any, the log streaming lags like hell, they charge self-hosted runners for using their coordination plane, jobs queuing is really slow, and their software for actually running jobs is cursed and designed in a way that is practically hostile to self-hosting.

      GitHub Actions is god-awful. Have you ever used any other CI tools?

  • djieidj283 1 hour ago
    It’s ironic that the SCM that 2026 software engineers landed on, is one with a complicated distributed usage model and is slow/down because of a centralised service
  • rvz 2 hours ago
    They were cooked the moment they got acquired by Microsoft.

    This is why I foresaw that centralizing everything to GitHub was just generally a bad idea 6 years ago. [0]

    Now that there is no CEO of GitHub, there is no point to GitHub improving.

    [0] https://news.ycombinator.com/item?id=22867803

  • fenio 2 hours ago
  • arlattimore 3 hours ago
    I literally chuckled out loud when I saw the headline :D
  • GrumpySciGuy 1 hour ago
    Yes, but I did not need to look at the tracker to know that.....
  • kevmo 3 hours ago
    An important thing to consider is how much of their uptime without incidents is not the normal working hours. Their incident-free uptime on 9-5 EST, Mon-Fri, is probably like 60%.
    • thombles 2 hours ago
      As a daily GitHub user in Australia I still haven’t figured out why everyone’s complaining about uptime. :)
    • perfectstorm 2 hours ago
      what's normal working hours? very US centric comment IMO. Europe, India, China, Latam etc. don't fall into your 9-5 EST normal working hour bucket.
      • padjo 2 hours ago
        A service like GH will still show a daily usage pattern, often with a peak somewhere around 16:00 UTC when most of the US and Europe are at work.
      • boredatoms 1 hour ago
        Im fairly certain that SWEs only exist on the US west coast
    • CoastalCoder 2 hours ago
      > Their incident-free uptime on 9-5 EST, Mon-Fri, is probably like 60%.

      And it may be even worse in EDT, which is currently in effect!

  • johnea 3 hours ago
    Oh, I thought it said "Github Outrage Tracker".

    I was ready to click...

    • HPsquared 3 hours ago
      That one would be like iscaliforniaonfire.com
  • jakub_g 3 hours ago
    Possibly inspired by:

    https://red-squares.cian.lol/

    • gen220 2 hours ago
      Hey! isgithubcooked.com is my site; I made this site in February '26, I think

      The contribution graph as an outage calendar idea is a commonly recurring one :). I definitely saw it somewhere else as a static asset first before I made this site.

    • ChrisArchitect 3 hours ago