FeedSources
⌘K
AR

AMD ROCm Blog

38 articles total

Go to source

Explore AMD's latest technical deep dives, software optimization guides, and community insights for the ROCm™ platform.

Go to source
  •  Scaling RL with verl on AMD Instinct MI355X: Async Walkthrough and Sync Benchmark
  •  Exploring XGBoost: A Deep Dive
  •  Memory Instruction Scheduling for Lock-Stepped Kernels on AMD Instinct™ MI300X: Introducing the Series
  •  Bring Claude Code On‑Prem with AMD Instinct GPUs
  •  Production-Ready MXFP4 Online Rotation with Fused Kernels on AMD Instinct™ MI355X
  •  Using ODC to Accelerate AMD SFT Training
  •  Quark Support for HuggingFace Diffusers and SVDQuant
  •  AUP Learning Cloud: Streamlining AI Education on AMD
  •  VSA: Accelerating Video Diffusion Inference with Sparse Attention on AMD GPUs
  •  Reverse-Engineering hipBLASLt TensileLite Kernels: From Solution Name to a Tuning Config
  •  Introducing AMD CDNA™ 5 and the AMD Helios™ Rackscale Solution
  •  Closing the GPU Cluster Validation Gap: A Kubernetes-Native Approach with CVF
  •  AMD GPU Operator v1.5.0: DRA Support, Automated GPU Node Recovery, and Expanded Kubernetes Infrastructure Control
  •  Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide
  •  Enabling Language-specific Reasoning in Multilingual Models with Reinforcement Learning
  •  Introducing Instella-MoE: A State-of-the-Art Fully Open Mixture-of-Experts Language Model
  •  Hyperloom - Autonomous Agentic Inference Optimization for AMD GPUs
  •  Introducing AMD ROCm™ Infera: Scaling Goodput for Agentic AI with Distributed Inference Orchestration
  •  Serve Kimi-K2.5-MXFP4 on MI355X with ATOM
  •  Onboard and Deploy Custom Models in AMD AI Workbench