IBM has signed a $240 million agreement with Together AI to use an inference cluster equipped with about 2,000 Nvidia Blackwell 300 chips.

Inference is the work performed when a trained model answers, classifies, searches or executes a task. In heavily used products, that phase can consume more cumulative capacity than initial training.

Together AI specialises in serving open models. IBM wants to offer customers an architecture where they can choose and change models according to cost, privacy and capability.

Flexibility is not automatic. Applications, evaluations and data must be designed to compare outcomes. Changing models without measuring quality can replace dependency with uncertainty.

AdvertisementARQUITHEAArchitecture for seeing more clearlyIdeas, buildings and tools for understanding the city through real questions.Follow @arquithea_

The fact. IBM is reserving capacity before the service begins. Enterprise competition is shifting towards whoever lets customers combine models without being locked into one.

The question

Why would a company want several models instead of one excellent one?

Because classification, contracts, images and code do not carry the same cost or risk. Choice reserves the expensive model for difficult work and uses cheaper systems where sufficient quality is enough.