IBM and Together AI Sign $240 Million Multi-Year Agreement for Nvidia-Powered Inference Cluster on IBM Cloud
Key Takeaways
- •IBM committed $240 million over multiple years to provide Together AI with a dedicated Nvidia HGX B300 inference cluster hosted on IBM Cloud.
- •The B300 systems, configured with Nvidia Spectrum-X Ethernet networking, are scheduled to begin operations in the first quarter of 2027.
- •Together AI currently processes 400 trillion tokens per month through its inference product and supports more than one million developers.
- •Together AI recently completed an $800 million Series C funding round at an $8.3 billion valuation to expand its AI Native Cloud platform.
- •The agreement reflects IBM Cloud's strategy to compete with larger hyperscale providers by offering dedicated GPU capacity tied to specific high-growth AI customers.

IBM has committed to a multi-year, $240 million agreement with Together AI to build a dedicated Nvidia inference cluster on IBM Cloud, the two companies announced on Tuesday. The deal provides the open-source AI provider with a substantial block of enterprise-grade GPU capacity as it expands its reach into large corporate accounts. The agreement also reflects a broader shift in AI infrastructure spending, as demand for inference — the process of running trained models to generate outputs — grows alongside the deployment of AI applications in production environments.
HGX B300 Systems to Come Online in Early 2027
Under the signed agreement, IBM will deploy a large cluster of Nvidia HGX B300 systems within IBM Cloud. The hardware is expected to come online in the first quarter of 2027, according to IBM's newsroom release. The multi-year lead time underscores the extended planning horizons now common for securing next-generation GPU capacity, as cloud providers and AI companies lock in hardware well before it ships.
Together AI will use this capacity to serve inference on open-source models. IBM stated that this is the first dedicated cluster of this scale built for inference on IBM Cloud around the B300, configured with Nvidia's Spectrum-X Ethernet networking.
Nvidia claimed that the configuration can generate 30 times more "AI factory output" than its previous generation, though this figure has not been independently verified.
Together AI's 400 Trillion Monthly Tokens to Run on IBM
Together AI recently completed an $800 million Series C funding round at an $8.3 billion valuation, aimed at expanding what it calls its AI Native Cloud. Founded in 2022, the company rents cloud infrastructure to enterprises seeking to build on open, modular AI stacks rather than closed proprietary models. That positioning places Together AI in a competitive landscape that includes both proprietary API providers such as OpenAI and Anthropic and other open-source-focused inference platforms, as enterprises weigh cost, control, and performance tradeoffs.
According to IBM's press release, Together AI currently channels 400 trillion tokens per month through its inference product and supports more than one million developers. The IBM cluster will provide the infrastructure platform for Together AI's planned expansion.
Together AI CEO Vipul Ved Prakash said the agreement enables the company to reach more customers without requiring them to pay frontier-model prices.
"Enterprises want the performance of the best frontier models without the closed-model price tag, and that only works if the infrastructure underneath is fast and reliable at scale," Prakash said in the release. He added that the cluster represents "a big step in our push to make open-source AI the obvious choice for enterprises."
IBM Deepens Nvidia Partnership to Serve AI Startups
For IBM, the agreement with Together AI adds another layer to its growing partnership with Nvidia. Alan Peacock, general manager of IBM Cloud, said the deployment responds to corporate demand for agentic AI, noting that IBM and Nvidia are "delivering scalable, economical, enterprise-grade AI infrastructure" to accelerate Together AI's operations. The deal signals IBM Cloud's effort to compete for AI infrastructure workloads against larger hyperscale cloud providers by offering dedicated GPU capacity tied to specific high-growth AI customers.
Dion Harris, a senior director for HPC and AI infrastructure at Nvidia, described AI factories as infrastructure becoming as fundamental as electricity or telecommunications.
IBM also noted that the two companies have been collaborating on GPU-native data analytics and unstructured data extraction. The partnership encompasses consulting services, with the Together AI cluster positioned as part of a broader strategy to sell AI capacity to both large enterprises and startups.