Today we’re launching 𝗠𝗶𝗹𝗲𝘀 𝘃𝟬.𝟭, an open-source RL framework for LLMs and multimodal models. RadixArk’s mission is to make AI infrastructure open and accessible. RL is becoming a critical part of modern AI development, but training remains easy to start and hard to debug. Miles helps teams ensure runs are correct, use hardware efficiently, and keep RL running reliably at scale. Over the past nine months, 72 contributors have landed 1,326 commits and built 85 GPU E2E tests, battle-testing Miles on frontier open models including Kimi (Moonshot AI) K3, DeepSeek AI V4, Qwen 3.8, Z.ai GLM 5.2, Thinking Machines Lab Inkling, and MiniMax H3. Miles already powers frontier model development and production RL workloads at humans&, Periodic Labs, Modal, Decagon, Eigent AI, Nebius, IBM, and more, across both NVIDIA and AMD hardware. Check out our brand-new landing page for customer stories, supported models, and production features 👉 miles.radixark.com
About us
- Website
-
https://radixark.ai
External link for RadixArk
- Industry
- Technology, Information and Internet
- Company size
- 11-50 employees
- Type
- Privately Held
Employees at RadixArk
Updates
-
DSpark for Kimi K3 just got an upgrade: we further improved the model for long context and agentic use cases. - Accept length 4.2 in RULER v2 1M context length - Accept length 4.66 on SWE-rebench - Improvements across most of the benchmarks we tested (one small regression on GSM8K) And a milestone worth celebrating: DSpark crossed 140k downloads in the first few days after launch. Seeing a model we trained get picked up this quickly means a lot to us Thank you to everyone who tried it, filed issues, and pushed it into real workloads. Latest checkpoint: https://lnkd.in/gsvgQWMn
-
-
Very excited to partner with Google on TPUs! We're committed to delivering fast, native TPU inference for major open models and support for advanced features. More news to come 🚀
Exciting news for the AI developer community! We’re thrilled to announce a new partnership between Google and RadixArk to bring the full power of SGLang to Google Cloud TPUs. Our goal is simple: provide developers with ultimate flexibility when scaling high-performance serving. Currently, you can run SGLang on TPUs through the SGL-JAX repository. This release enables SOTA performance with features like Radix Cache, speculative decoding, and P/D disaggregation. And we’re just getting started. Later this year, RadixArk will roll out SGL-torchtpu, a PyTorch-native backend that brings seamless TPU support to your existing PyTorch workflows. Check out the full announcement and start building with SGL-JAX today: https://goo.gle/3U2JVOC #GoogleCloud #TPU #SGLang #RadixArk #OpenSourceAI #MachineLearning #GenerativeAI
-
-
RadixArk and Google Cloud are joining forces with the SGLang community to make TPU a drop-in, cost-efficient path to frontier inference. SGL-JAX already serves the major open model families on the latest TPU generations: Gemma, Qwen, DeepSeek, GLM, Kimi, Ling, MiniMax, MiMo, Grok, plus Wan and Flux for video and image generation. With this partnership, TPU unlocks SGLang's production feature set: 5D parallelism, Radix Cache, HiCache, quantization, speculative decoding, on TPU Pallas kernels co-built by Google, RadixArk, and the SGLang community. SGL-torchtpu opens a PyTorch-native path to all of it. SGLang on TPU comes with the same API and the same features, so for teams already running it, getting started is simple. Going forward, TPU support for new open models will land alongside the rest, and that carries forward to each new TPU generation. Full announcement and SGL-JAX repo 👇 The blog: https://lnkd.in/gnkSQrZX Run models on the latest TPU generations with SGL-JAX: https://lnkd.in/gTB5iCgZ
-
-
We trained a DSpark speculator for Inkling NVFP4, built end-to-end with SpecForge on a live SGLang target engine. - 1.89x decode throughput over non-spec at bs=64 (8xB200, TP8) - 7–14% faster than Inkling's built-in MTP at the same batch size - 3.66 mean accept length, up to 4.76 on GSM8K Under the hood: 1️⃣ 5-layer Qwen3-style DFlash parallel-draft backbone 2️⃣ Rank-256 Markov logit-bias head for intra-block dependencies 3️⃣ Per-position confidence head predicts acceptance 4️⃣ Distilled online from 400K Inkling regenerations Hugging Face: https://lnkd.in/gkCvT-W2 Try it out now! Serving command in the comments.
-
-
A great write-up from DigitalOcean! Reliable MI350X capacity + joint engineering with AMD (HIP graphs = ~10x throughput) is how we got DeepSeek V4 to 3.5K+ tok/s/GPU on SGLang. Open models, open hardware, one codebase. Read the full article: https://lnkd.in/g6CuZvPU
The team that maintains SGLang, the engine serving trillions of tokens a day, brought DeepSeek AI V4 to production on DigitalOcean and AMD. 🤝 RadixArk moved to AMD Instinct™ MI350X GPU Droplets and co-engineered a ~10x throughput gain with the AMD team, from debugging on borrowed hardware to production-grade inference.
-
We'll be there on July 21st! Come talk inference, open source, and what's next for SGLang on AMD. See you at the kickoff 🍻 Register: https://luma.com/zz40nkob
Kickoff AMD Advancing AI early. 🍻 We’re hosting a happy hour with AMD, RadixArk, and friends in SF. Valuemaxxing your inference and infra? We'll have lightning talks and live demos from teams doing exactly that. If you're building or backing AI-native companies, you won't want to miss this. 🎟️ https://do.co/4vdCMZv
-
-
RadixArk reposted this
A few weeks ago, I shared that an opportunity I’d been preparing for fell through unexpectedly, just days before my start date. I never imagined that post would reach so many people. To everyone who reached out, sent kind words, made introductions, or simply checked in — thank you. Every conversation, referral, and word of encouragement meant more than you know during a really uncertain time. Today, I'm happy to share that I've joined RadixArk as a Designer! A huge thank-you to Mingyi Lu and Haoguang Cai, who reached out after seeing my last post. From our very first conversation, the process was thoughtful and smooth. The team took real time to understand my background, how I think as a designer, and where I could contribute most. I'm especially excited about the breadth of the Designer role and the opportunity to contribute across many areas of design as we grow. Everyone I met at RadixArk is deeply passionate about what they're building, and just as welcoming to new people. I am beyond thrilled to join this exceptional, fast-growing team behind SGLang and Miles, open-source AI infrastructure used by developers and enterprises around the world. Once again, thank you so much to everyone who supported me over the past few weeks! I genuinely appreciate all the help from tech and design communities, and I hope I’ll have the chance to pay it forward someday. 🤍
-
-
Miles is now featured on the PyTorch Foundation blog. As models grow, shift from dense to MoE, and span more specialized hardware, RL post-training is no longer just about the algorithm. It is a distributed systems problem. Miles is our open-source RL training framework, built for exactly that. It comprises four systems behind a small, pluggable trainer: SGLang (SGLang) for rollout, Megatron-LM (NVIDIA) for training, Ray (Anyscale) for orchestration, and PyTorch (PyTorch) as the common layer for models and numerics. Out of the box, you also get MoE-aware rollout/training alignment, a unified BF16/FP8/MXFP8/NVFP4/INT4-QAT pipeline, fast NCCL/RDMA weight sync, fault tolerance, and ready-to-run recipes for frontier models like DeepSeek V4, GLM 5.2, Qwen3.6, Kimi K2.6, and Nemotron 3 Ultra. Our goal is simple: make frontier-scale LLM RL easier to reproduce, extend, and operate. Thank you, PyTorch Foundation, and everyone who got Miles here, especially the legendary slime team!
Built on PyTorch, Ray, SGLang, and NVIDIA Megatron-LM, Miles is an open source framework from RadixArk for large-scale LLM reinforcement learning post-training. Miles uses PyTorch for models, numerics, profiling, and extensibility; Ray for orchestration; SGLang for rollout generation; and Megatron-LM for distributed training. The framework supports asynchronous rollout and training, NCCL/RDMA weight synchronization, MoE-aware rollout/training alignment, low-precision recipes, LoRA, fault tolerance, observability, and extension points for custom algorithms and model architectures. 🔗 Read more in our latest blog from the Miles Team: https://lnkd.in/ec3kGd_n
-