FeedSources
⌘K
BB

Baseten Blog

29 articles total

Go to source

Stories, updates, and other resources from Baseten.

Go to source
  •  Inference engineering for DeepSeek V4 Pro 0813
  •  Baseten delivers open-source inference for You.com
  •  Qwen3.8-Max: Alibaba's new frontier reasoning model
  •  Introducing NVIDIA Nemotron 3.5 Lightning
  •  Introducing NVIDIA Nemotron 3.5 ASR Streaming
  •  Laguna S 2.1 goes Greek: a repository-scale game transformation
  •  Fine-tuning Qwen3-TTS for high-quality voice cloning
  •  22,580: GPT-2 to Kimi K3, explained
  •  How to run Kimi K3 in any harness: routing with Baseten Switch
  •  Announcing Baseten for Model Labs
  •  Welcome, Mani Parkhe!
  •  Making Kimi K3 tokenization 18x faster for million-token agentic workloads
  •  How to build a day-0 API for Kimi K3
  •  How we built the new fastest API for GLM-5.2
  •  Introducing GLM 5.2 Fast
  •  How to choose an AI model: lessons from Notion and Gamma
  •  How to optimize LLM inference speed and reduce costs in production
  •  H100 vs. H200 GPUs
  •  GLM 5.2 With Vision
  •  Real-time video generation inference on Baseten