FeedSourcesModels
PrivacyImpressum
⌘K
AR

AMD ROCm Blog

68 articles total

Go to source

Explore AMD's latest technical deep dives, software optimization guides, and community insights for the ROCm™ platform.

Go to source
  •  Verl 0.9.0 on ROCm™ 10.0: Next-Generation RL Post-Training on AMD Instinct™ GPUs
  •  Local Quantization and Multi-Backend Deployment with AMD Quark on Strix Halo
  •  Model Weight Profiles: Where Do the Parameters Go?
  •  UltraQuant on AMD Instinct: More Efficient Agentic Serving for Qwen3.8-MXFP4
  •  Debugging Logprob Mismatches in LLM Reinforcement Learning
  •  Performance Profiling on AMD GPUs - Part 6: Advanced Thread Trace (ATT) - The Microscope for Your Application
  •  Zebra-HyLo: Upcycling Transformers into Long-Context Hybrid LLMs on AMD Instinct™ GPUs
  •  Benchmarking Kimi-K3 Across vLLM, SGLang, and ATOM on MI350X
  •  Hyperloom: A Multi-Agent Harness for Autonomous Inference Optimization on AMD GPUs
  •  Serving GLM-5.2-MXFP4 on AMD Instinct™ MI355X: When Prefill Context Parallelism Pays
  •  Reproducing AMD MLPerf Inference v6.1 Submission Results
  •  Technical Dive into AMD MLPerf Inference v6.1 Submission
  •  Implementing a High-Performance Custom Diffusion Attention Kernel with FlyDSL
  •  DFlash Speculative Decoding on AMD Instinct MI355X: Up to 5× Faster Qwen3.5 Inference
  •  Thread Trace Part 1: ROCprof Compute Viewer
  •  Knowledge Graph Integration With Poro2 For Enriching Medical Text Processing
  •  An Educational GEMM Ladder for Helios GPUs
  •  Iteratively Tuning hipBLASLt TensileLite Kernels: A Smaller Search, a Faster Kernel
  •  veRL on AMD: Production-Ready RL Post-Training on ROCm
  •  Efficiently Serving NVFP4 Models on AMD Instinct™ MI350X/MI355X Accelerators via Online NVFP4 to Quark MXFP4 Requantization