GitHub 目前遇到故障。
GitHub Is Experiencing Difficulties

原始链接: https://www.githubstatus.com/?d=2026/08/06

7月24日,GitHub 其中一个可用区出现网络连接故障,导致数据包丢失,影响了该可用区 25% 的网络容量。此次中断发生在世界标准时间(UTC)16:04 至 17:36 之间,起因是计算集群的脊柱交换机与聚合层之间的链路故障。 该事件导致多个服务出现间歇性性能问题和报错,包括 GitHub Actions、Issues、Copilot、身份验证及 Git 操作。GitHub 通过将流量重新路由至备用光纤路径缓解了影响,并于世界标准时间 17:36 完全恢复了服务稳定性。 为防止此类事件再次发生,GitHub 正在加速推进一项计划中的基础设施升级,将网络接口从 100Gbps 提升至 400Gbps。此次升级将增加交换网络结构的带宽,从而确保在未来面对设备或链路丢失时具备更强的韧性。

GitHub 目前正经历服务中断,这引发了 Hacker News 上关于该平台近期可靠性的不满讨论。 用户报告称,GitHub Actions、CI/CD 流水线以及下载发布资源的功能出现了严重中断。对于许多开发人员而言,这些反复出现的故障已成为重大的工作流程瓶颈,阻碍了代码部署和修复安全漏洞等关键任务。 此次讨论凸显了一个典型的供应商锁定案例:尽管许多工程师对 GitHub 频繁的宕机表示深感不满和沮丧,但大多数人指出,迁移所带来的高昂成本(工程时间和运营复杂度)超过了这些中断带来的不便。 一些贡献者建议转向自托管方案以重新获得控制权,而另一些人则认为,尽管最近出现了问题,但 GitHub 仍然比管理自己的基础设施更可靠。归根结底,这一讨论反映了人们对 GitHub 稳定性的日益担忧,并呼吁在服务故障时提供更主动的沟通和状态报告。
相关文章

原文
Resolved - On July 24th at 16:04 UTC, a loss of connectivity occurred in network paths in one of our three physical data center availability zones (AZs). This resulted in packet loss due to the remaining active paths becoming saturated. Our data centers use a leaf-spine switch fabric in each compute cage, and an aggregation layer interconnecting the spines from each cage within each AZ. The loss of connectivity affected links between one cage’s spine switches and the aggregation layer within that specific AZ.

Workloads depending on compute resources in this cage became degraded due to packet loss, and exhibited intermittent errors:

- Actions saw 10% of jobs fail during the impact window, and 5% of jobs succeeded but with delayed starts.
- 27% of GitHub issues interactions saw slow requests or timeouts.
- 4% of GitHub Copilot requests experienced errors, though most automatically retry.
- 4% of git push operations saw impacts during the affected window.
- Authentication requests saw increased latency during the affected window, but error rates, while elevated, were < 1% in all cases.

We were able to mitigate the outage by re-routing affected connections to available fiber paths that were allocated for future capacity upgrades. Sufficient network capacity to eliminate packet loss was restored at 17:07, with most services showing full recovery by 17:16. All paths were restored and services healthy at 17:36.

This incident affected 25% of available network interconnect capacity. Older cages utilize a 100Gbps network interface standard. To remove risk of reoccurrence, a planned upgrade to 400Gbps interfaces is being accelerated as much as possible, ensuring increased bandwidth available at all layers of the switch fabric for resiliency to path or device loss.


Jul 24, 17:36 UTC

Update - We are seeing recovery across all services
Jul 24, 17:24 UTC

Update - The degradation affecting API Requests, Actions, Copilot, Issues, Pages and Pull Requests has been mitigated. We are monitoring to ensure stability.
Jul 24, 17:16 UTC

Update - Actions is experiencing degraded performance. We are continuing to investigate.
Jul 24, 16:41 UTC

Update - We have applied a mitigation and are monitoring for recovery
Jul 24, 16:40 UTC

Update - Actions is experiencing degraded availability. We are continuing to investigate.
Jul 24, 16:28 UTC

Update - Pages is experiencing degraded performance. We are continuing to investigate.
Jul 24, 16:27 UTC

Update - Copilot is experiencing degraded performance. We are continuing to investigate.
Jul 24, 16:26 UTC

Update - We are investigating timeouts to some GitHub services
Jul 24, 16:22 UTC

Update - Pull Requests is experiencing degraded performance. We are continuing to investigate.
Jul 24, 16:20 UTC

Update - Actions is experiencing degraded performance. We are continuing to investigate.
Jul 24, 16:19 UTC

Investigating - We are investigating reports of degraded performance for API Requests and Issues
Jul 24, 16:17 UTC

联系我们 contact @ memedata.com