FeedSources
⌘K
AR

AMD ROCm Blog

48 articles total

Go to source

Explore AMD's latest technical deep dives, software optimization guides, and community insights for the ROCm™ platform.

Go to source
  •  veRL on AMD: Production-Ready RL Post-Training on ROCm
  •  Efficiently Serving NVFP4 Models on AMD Instinct™ MI350X/MI355X Accelerators via Online NVFP4 to Quark MXFP4 Requantization
  •  Enabling DeepSeek-V4-Flash Training on AMD Instinct MI355X GPUs with Primus
  •  Optimizing ATOM and vLLM-ATOM for High-Interactivity Inference
  •  4-bit KV Caching in LMCache: Offloading Quantized KV Beyond HBM for Context-Heavy Agents on AMD MI355X
  •  A Deep Dive into LDS Optimizations on AMD Instinct MI450 GPUs
  •  ROCm 10.0: A Decade of Open Compute, Built for the Age of Agentic AI
  •  Enabling Physical AI Agents with Lemonade
  •  Serving 64Mi-Token Contexts on One AMD Instinct™ MI355X Node
  •  DI Series: Scaling GLM-5.1-FP8 to 64 MI300X GPUs
  •  Scaling RL with verl on AMD Instinct MI355X: Async Walkthrough and Sync Benchmark
  •  Exploring XGBoost: A Deep Dive
  •  Memory Instruction Scheduling for Lock-Stepped Kernels on AMD Instinct™ MI300X: Introducing the Series
  •  Bring Claude Code On‑Prem with AMD Instinct GPUs
  •  Production-Ready MXFP4 Online Rotation with Fused Kernels on AMD Instinct™ MI355X
  •  Using ODC to Accelerate AMD SFT Training
  •  Quark Support for HuggingFace Diffusers and SVDQuant
  •  AUP Learning Cloud: Streamlining AI Education on AMD
  •  VSA: Accelerating Video Diffusion Inference with Sparse Attention on AMD GPUs
  •  Reverse-Engineering hipBLASLt TensileLite Kernels: From Solution Name to a Tuning Config