Equinix plans to make its new Inference Exchange available from the first quarter of 2027, combining NVIDIA infrastructure with Together AI’s model-serving platform and connectivity through Equinix Fabric.
Announced on September 2, the program targets enterprises deploying AI inference across multiple locations. Its intended role is to put model execution nearer the applications and data it serves, rather than treating the choice of data center as separate from the inference platform.
Three Layers in One Deployment Program
The partners divide responsibility across the infrastructure stack. Equinix supplies facilities, power, cooling and ongoing operations; NVIDIA contributes Enterprise Reference Architectures; Together AI supplies the inference platform, including support for shared multitenant and dedicated single-tenant environments.
Equinix Fabric will connect the offering to cloud, network and AI providers. The announcement describes Together AI’s platform as supporting more than 200 open-source models, but that platform-level figure should not be confused with a guaranteed launch catalog at every location.
Equinix’s product notice describes the announcement as a preview and says further information about deployment locations and customer access will follow. Enterprises therefore have a target launch window, not a confirmed local service endpoint they can use today.
Location Is Part of the Inference Decision
The program is designed around three scenarios: bringing inference closer to users at the metropolitan edge, moving from proprietary models to open alternatives, and locating workloads where data-residency requirements can be addressed. These are intended use cases, not evidence that every deployment automatically satisfies a particular regulatory requirement.
Equinix says the approach brings infrastructure, interconnection and inference capabilities together. The operational implication is that model-serving choices and connectivity choices can be evaluated as a combined deployment, including the network path to enterprise data and the placement of dedicated capacity.
That distinction matters for applications that make repeated calls to a model. A nearer inference service may reduce network travel, but the announcement does not establish measured end-to-end application latency. Model size, request queues and the data path will still affect the experience.
A Global Footprint Is Not a Launch-Site List
The announcement cites Equinix’s network of more than 280 data centers in 77 metropolitan markets. That describes the company’s wider footprint, not a commitment to introduce Inference Exchange in all those markets at once.
For infrastructure buyers, the next useful disclosures are the initial locations, access conditions and the capacity available in each. Those details will determine whether the offering fits an existing application’s topology or requires changes to data placement and connectivity.
The immediate change is a new announced route to distributed enterprise inference. Production availability remains ahead: Equinix’s stated starting point is Q1 2027, and its preview notice leaves location and customer-access details for later updates.







