PinnedWei Lu·Mar 20, 2025Local DeepSeek-R1 671B on $800 configurationsUpdate on April 6,2025: Llama 4 has been released, and along with DeepSeek V3/R1, it marks a step into the top-tier large model platform…A response icon18A response icon18
PinnedWei Lu·May 24, 2024Intel Core Ultra 5 vs. Apple M1 on LLM inferenceI bought an Intel “AI PC” equipped with a Core Ultra 5 125H and 96GB of memory, which means a graphics card with 48GB of VRAM. This mini PC…A response icon6A response icon6
Wei Lu·2d agoUnified Memory for Local MoE Inference: Strix Halo against an RTX 3060 Laptop PC, and What It…A 128 GB APU looks like the obvious machine for a large Mixture-of-Experts model: the whole model fits in memory the GPU can read, with no…A response icon1A response icon1
Wei Lu·5d agoEJB, Spring, Kubernetes: The Java History Behind Claude Code Mods and DeepSeek HarnessWho owns the extension points, and where the container boundary sits — from stateless session beans to Claude Code plugins, DeepSeek…
Wei Lu·Oct 1One draft model, two verdicts: DFlash 2 takes Qwen3.8–27B2026–09–30 · Follow-up to Four-bit floats hit the same wall and The small model gets its revenge · Measurements from 2026–09–30, same…
Wei Lu·Sep 30From vibe coding to the software factory: the five states in between2026–09–29 · A survey of practitioner reports, vendor engineering blogs and a handful of preprints, 2025–2026
Wei Lu·Sep 28Four-bit floats hit the same wall: FreeToken NVFP4 vs llama.cpp2026–09–24 · Follow-up to The small model gets its revenge: Qwen3.8–27B on an RTX 5090 in the same Windows box · Measurements from…
Wei Lu·Sep 24The small model gets its revenge: Qwen3.8–27B on an RTX 5090 in the same Windows boxFollow-up to The smaller model is not the faster one: Qwen3.8–27B dense vs 125B-A6B Flash-Next on the same Strix Halo · Measurements from…
Wei Lu·Sep 11Undo is easy. Redo is where coding agents differ.How Claude Code, Codex CLI, Pi, opencode and qwen-code actually store, branch and rewind a session — read from the source
Wei Lu·Sep 9The smaller model is not the faster one: Qwen3.8–27BFollow-up to Running Qwen3.8-Flash-Next on Windows with a Ryzen AI MAX+ 395 and Chasing the last 30% · Measurements from 2026–09–03 and…