Local Open-Weight LLMs in Coding Harnesses
Read OriginalThis article evaluates local open-weight large language models (LLMs) such as Qwen-Code, Codex, and Claude Code in various coding harnesses. It highlights that 30B Mixture-of-Experts models offer a sweet spot, solving challenging problems at roughly 40 tok/sec on a Mac or DGX Spark, comparable to GPT 5.5. The author compares harness choices, noting Claude Code uses twice as many tokens as Codex, and includes Gemma 4 E2B as a reference to show smaller models struggle. A full write-up is available, with a chart comparing token use and task success across five local-agent tasks.
Comments
No comments yet
Be the first to share your thoughts!
Browser Extension
Get instant access to AllDevBlogs from your browser
Top of the Week
No top articles yet