Monday, September 14, 2026
AI Infrastructure · News & Analysis
HomeChips & HardwareReport
Chips & Hardware · Report

Together AI has optimized ThunderKittens, an HPC tensor library, for performance on NVIDIA's Vera Rubin GPU, enhancing workload efficiency on the new accelerator.

Software-hardware co-optimization by a GPU cloud provider shows competitive differentiation through compiler work as NVIDIA's architecture portfolio expands.
Trade pressSlicast · September 12, 2026 · US · Source: TipRanks
importance 25

Together AI has updated its ThunderKittens GPU kernel framework to run on NVIDIA's new Vera Rubin NVL72 platform. Having gained early access to the architecture, the company's kernels team explored the new instruction set and implemented NVFP4 and FP8 GEMM support.

Through optimization techniques including wider matrix multiply-accumulate steps, larger tensor and shared memory, deeper pipelines, and revised collector behavior, the NVFP4 kernel achieved over 22 PFLOPS of performance. According to the company's LinkedIn post, this performance level brings ThunderKittens into competitive range with NVIDIA's cuBLAS and the CuTe domain-specific language for key linear algebra workloads.

Together AI has documented the technical differences between NVIDIA's Blackwell and Vera Rubin architectures and the specific optimizations required to reach these performance levels. The work positions Together AI's software stack to align closely with next-generation NVIDIA hardware, potentially strengthening the company's role in high-performance AI infrastructure. For enterprise customers optimizing workloads on cutting-edge GPU clusters, competitive performance parity with NVIDIA's own libraries could expand Together AI's market opportunity in efficient model training and inference.

Read the original