Together AI is expanding its GPU cluster offerings with discounted preemptible compute instances, targeting price-sensitive and interruptible AI workload segments.
Together AI has launched preemptible compute options for its GPU Clusters, offering a flat 50% discount to on-demand pricing. According to the company's recent LinkedIn announcement, these preemptible nodes target interruption-tolerant workloads including short experiments, evaluations, fine-tuning tasks, batch jobs, and inference bursts.
When capacity is reclaimed, workloads receive up to five minutes to checkpoint and exit gracefully. The cluster automatically refills toward the requested target capacity as resources become available. The feature is currently in public preview, suggesting Together AI is validating demand and performance before broader rollout.
The offering reflects Together AI's strategic focus on price-efficient GPU infrastructure scaling in a highly competitive AI compute market. By monetizing idle or flexible capacity while appealing to developers running non-critical workloads, the company can expand its user base while maintaining margins—a positioning that strengthens its competitive standing against larger cloud providers offering comparable discounted compute tiers.