System-Level Optimization of AI Infrastructure with AMD Instinct GPUs

Extracting maximum value from an AI cluster requires an uncompromising, end-to-end system-level optimization journey.

Download

We share our hands-on experience optimizing an AI cluster built on AMD Instinct MI355X GPUs, and how we first validated the optimal host configuration. Explore how we tuned network parameters to ensure that the entire system runs smoothly and delivers industry-leading results.

This white paper showcases DriveNets’ hands-on methodology for optimizing an AI cluster built on AMD Instinct MI355X GPUs. This includes: host configuration, GPU firmware, BIOS settings, operating system parameters, drivers, network behavior, congestion control, switch settings, and workload benchmarking.

  • The methodology is reflected in the DriveNets verified Reference Architecture with AMD, designed to help customers move from GPU deployment to production-grade AI infrastructure​
  • Following this E2E approach, DriveNets is able to improve inference performance across single-node and multi-node environments​
  • Demonstrated higher throughput per GPU, faster time to first token, and strong interactivity underload

Continue reading

DriveNets Advances Open Backend Networking for AI with AMD

Industry Events

DriveNets Advances Open Backend Networking for AI with AMD

DriveNets works with AMD to deliver an open, full stack backend networking solution for AI advanced by AMD Instinct™ M ...

Read more

Collateral

AMD and DriveNets Reference Architecture

Building high-performance AI clusters is no longer a matter of adding GPUs.

Read more

White Papers

Faster LLM Inference on AMD Requires Rethinking All-Reduce

Extracting maximum ROI from AMD AI clusters requires moving beyond out-of-the-box software bottlenecks.

Read more