GitHub Puts "Availability First" as It Rethinks Scaling Priorities Following Recent Outages

The company has also apologized for recent outages


GitHub has provided an update following two recent outages and admitted that none of them was acceptable. In an update from earlier today, GitHub’s CTO, Vlad Fedorov, even apologized saying, “we are sorry for the impact they had on you.”

As you may know, the company is currently experiencing an increased demand due to AI-driven/agentic development workflows. In fact, GitHub initially thought that the 10x increase would be enough to keep up with demands; however, it seems like a much larger increase will be necessary.

GitHub reveals plan for huge scale-up

GitHub says that it began preparations for infrastructure scaling as early as October 2025. However, the company understood that it may require building systems able to process up to 30x more loads by February 2026. As activity in repositories, pull requests, and automation keeps increasing rapidly, such a scale-up is required.

Image credit: GitHub

As mentioned above, the company sees the issue not only in volume but in complexity as well. For example, a single pull request affects a number of systems including storage, APIs, search, and background services. When the volume becomes high, such inefficiencies may result in serious issues very quickly.

Outages revealed some weaknesses

The first outage occurred on April 23 and concerned the functionality of the merge queue feature. According to GitHub, some pull requests generated merge commits incorrectly if the changes had been grouped into one merge commit. Over 650 repositories and 2,000 pull requests were involved in the problem, but there was no loss of data.

The second outage occurred on April 27 and affected the search features of the service. The malfunction was associated with an overloaded Elasticsearch cluster, which resulted in partial failures to display results. However, the core operations of Git did not suffer in any way.

Having said that, now GitHub is actively working on its system’s reliability. For example, the company looks to ensure isolation of critical services, minimize risks of a single point of failure, and switch to a multi-cloud strategy. In addition, it is expected to increase transparency in case of future outages.

With all that in mind, it is yet to be seen whether these efforts will suffice as GitHub moves toward further scalability of its platform. As a reminder, GitHub recently announced a major change as to how Copilot will be priced and billed, as it is moving the product from its current request-based model toward full usage-based billing starting June 1, 2026.

Article feature image source: GitHub

More about the topics: AI, Github, microsoft

Readers help support Windows Report. We may get a commission if you buy through our links. Tooltip Icon

Read our disclosure page to find out how can you help Windows Report sustain the editorial team. Read more

User forum

0 messages