Salesforce suffered a widespread Core Service outage on September 16, disrupting customers across multiple regions with severe delays, intermittent errors and, in some cases, an inability to access services.
The disruption began at 07:50 UTC, according to Salesforce’s incident record. During its investigation, the company found that requests were stalling while waiting for an internal login service, consuming available server resources. Salesforce later said increased load on a core system component had limited its capacity to process requests.
Salesforce eventually deployed a fix across its affected fleet and restored services. Independent monitoring by Cisco ThousandEyes detected HTTP 503 responses and timeouts from locations around the world during the incident, while finding no evidence that general network degradation was responsible.
Incident status
Resolved. The disruption began at 07:50 UTC on September 16 and affected Salesforce Core Service instances across multiple regions.
- Salesforce Reports Problems Across All Regions
- Internal Requests Began Consuming Server Resources
- The Fix Was Rolled Out Fleetwide
- Independent Monitoring Points to an Application-Layer Failure
- Some Hyperforce Instances Needed Additional Recovery
- Why the Outage Matters for Cloud Infrastructure
- Current Status
Salesforce Reports Problems Across All Regions
Salesforce first reported a service disruption on its Trust status page, later saying it was investigating multiple instances across all regions.
Customers could experience severe delays, intermittent errors or an inability to access some services. The incident also interfered with Salesforce’s own support infrastructure, with some customers unable to create new support cases through the Help portal.
The Register reported that hundreds of instances were affected worldwide, including infrastructure serving customers in the United States, Japan, India, the United Kingdom, France and Germany.
Internal Requests Began Consuming Server Resources
Salesforce’s investigation initially focused on an internal login service. At 09:10 UTC, the company reported that requests were stalling while waiting for responses from that service, using up available server resources.
Engineers initially attempted a rolling restart on an affected instance. Salesforce later abandoned restarts as the primary remediation path while continuing its investigation.
The company subsequently examined whether an external dependency failure was affecting a legacy login server and contacted its third-party infrastructure provider. According to Salesforce, the provider confirmed that its infrastructure was operating normally.
By 10:18 UTC, Salesforce had identified increased load on a core system component as the condition limiting its ability to process requests and was working on a fix.
The Fix Was Rolled Out Fleetwide
Salesforce tested its remediation on a test instance before beginning a fleetwide deployment. At 10:56 UTC, it said the test had been successfully validated and that the fix would be rolled out to affected instances.
The rollout then proceeded region by region. Salesforce reported improving service, although recovery was not immediate across every environment. Some instances required additional remediation, including targeted restarts and manual recovery where the automated fix had not fully resolved the problem.
- 07:50 UTC — Service disruption began.
- 09:10 UTC — Salesforce identified requests stalling on an internal login service and consuming server resources.
- 10:18 UTC — Increased load on a core system component was identified as limiting request-processing capacity.
- 10:56 UTC — Salesforce validated a fix on a test instance and began fleetwide deployment.
- Later September 16 — Salesforce completed additional recovery work and reported the incident resolved.
Independent Monitoring Points to an Application-Layer Failure
Cisco ThousandEyes independently observed Salesforce availability problems beginning at approximately 07:50 UTC from monitoring locations around the world.
Its systems recorded HTTP 503 Service Unavailable responses as well as request timeouts. ThousandEyes said the behavior was consistent with an application-layer problem and that it saw no signs of network degradation during the incident.
ThousandEyes observed broad service recovery by approximately 12:30 UTC. Salesforce continued remediation after that point on instances where the initial rollout had not fully restored normal operation.
Some Hyperforce Instances Needed Additional Recovery
The initial deployment did not completely resolve the problem everywhere. Salesforce later narrowed the remaining disruption to a subset of Hyperforce instances, according to updates reported by The Register.
Some customers who had regained access also reported scheduled jobs not running as expected. Salesforce re-applied its fix where necessary and used targeted restarts and manual recovery for environments that remained affected.
Salesforce said its first-party environments were unaffected during this later stage of the incident. GovCloud customers were also confirmed to be outside the remaining impact.
Why the Outage Matters for Cloud Infrastructure
The incident shows how a capacity problem inside a shared application component can propagate across a geographically distributed SaaS platform without an underlying Internet or third-party infrastructure failure.
That distinction matters for infrastructure teams diagnosing large cloud incidents. External symptoms such as timeouts and HTTP 503 responses can resemble connectivity or upstream provider failures even when the fault sits higher in the application stack.
In this case, independent network observations and Salesforce’s own investigation pointed toward the service layer: a core component experienced increased load, requests stalled around an internal login service, and server resources were consumed while waiting for responses.
Current Status
Salesforce reported the September 16 incident resolved after completing its remediation work. The company said the deployed fix restored services after the disruption affected customers across its global infrastructure.
The incident record establishes the immediate technical condition behind the outage, but Salesforce had not published a detailed post-incident root-cause analysis at the time of reporting. The distinction is important: increased component load and resource exhaustion are confirmed elements of the incident, while the deeper trigger responsible for creating that condition should not be inferred until Salesforce publishes further technical findings.







