A technical issue is the cause of OpenAI services being blocked for hours | ChatGPT login | Chat OpenAI | OpenAI ChatGPT | Turtles AI
OpenAI attributes a lengthy outage to a glitch in a new telemetry system that overloaded the Kubernetes infrastructure that powers its core products.
Key Points:
- Problem caused by a new telemetry service implemented by OpenAI.
- Outage lasted about three hours, impacting ChatGPT, Sora and the API.
- Overloaded Kubernetes servers, a key system for managing the infrastructure.
- OpenAI plans measures to avoid similar situations from happening again.
On Wednesday, one of the most significant outages OpenAI has ever faced hit its AI platform. Key services like ChatGPT, Sora, and the developer API were caught in a disruption that began around 3 p.m. Pacific time and took about three hours to resolve. The company identified the root cause as a new telemetry system designed to collect metrics related to Kubernetes, an open-source technology essential for managing containers. Containers, which act as isolated units for running software, are a key pillar of OpenAI’s infrastructure.
In a report released the next day, OpenAI clarified that the outage was not caused by a security issue or a recent product update. Instead, the new telemetry system triggered an unexpected load on Kubernetes operations, specifically on the API servers used to control and manage the main Kubernetes cluster. This overload destabilized the system, including compromising a critical component: the DNS resolution service. This system translates IP addresses into domain names, allowing you to access websites by typing in easy-to-remember names instead of numeric sequences, for example.
DNS caching also added to the problem, delaying diagnosis. When a DNS temporarily stores address information, such as recently visited websites, it can provide outdated data, making it harder to pinpoint the extent of an error. In this case, the telemetry system was fully implemented before the impact was fully understood.
Although OpenAI had identified the problem minutes before it became apparent to users, resolution was slow and cumbersome. The already stressed Kubernetes servers did not allow for timely interventions, forcing engineers to work on workarounds. According to OpenAI, the problem was the result of a combination of factors: infrastructure tests that did not detect the risk of overload, internal processes that failed simultaneously, and interactions between various systems that were not anticipated.
To prevent future incidents of this type, OpenAI announced a series of corrective actions. These include increased attention to monitoring changes made to the infrastructure, more robust mechanisms to ensure continuous access to Kubernetes servers, and better management of phased deployments. In a closing note, the company apologized for the inconvenience caused to users, including developers and businesses that rely on its products, noting that the incident fell short of the standards OpenAI aims to maintain.
OpenAI reinforces its commitment to reliable infrastructure management, turning the critical experience into an opportunity for improvement for the future.
