Faster LLM Inference on AMD Requires Rethinking All-Reduce

Extracting maximum ROI from AMD AI clusters requires moving beyond out-of-the-box software bottlenecks.

Download

While AMD Instinct GPUs offer massive compute density, standard all-reduce communication patterns leave GPUs waiting during tensor parallelism. Explore how DriveNets’ deep-stack software optimization bypasses standard RCCL limitations to lower latency and accelerate Time to First Token (TTFT) under high concurrency.

This white paper details DriveNets’ hierarchical all-reduce optimization tailored for AMD Instinct GPUs (specifically the MI355X). Learn how tuning the communication stack transforms LLM inference performance:

  • How traditional ring and tree all-reduce topologies introduce synchronization delays during tensor parallelism​
  • DriveNets’ custom hierarchical kernel that keeps data movement local and enables simultaneous cross-node exchange​
  • Real-world benchmark results, demonstrating up to 4.3x lower TTFT latency and consistent speedups across common message sizes

Continue reading

DriveNets Advances Open Backend Networking for AI with AMD

Industry Events

DriveNets Advances Open Backend Networking for AI with AMD

DriveNets works with AMD to deliver an open, full stack backend networking solution for AI advanced by AMD Instinct™ M ...

Read more

Collateral

AMD and DriveNets Reference Architecture

Building high-performance AI clusters is no longer a matter of adding GPUs.

Read more
Ultrafast inter-site networking from WhiteFiber, DriveNets, and WEKA

Videos

Ultrafast inter-site networking from WhiteFiber, DriveNets, and WEKA

The Industry’s first commercial deployment of a long-distance, scale-across AI supercluster

Read more