Sebastian Raschka 6/26/2026

Local Open-Weight LLMs in Coding Harnesses

Read Original

This article evaluates local open-weight large language models (LLMs) such as Qwen-Code, Codex, and Claude Code in various coding harnesses. It highlights that 30B Mixture-of-Experts models offer a sweet spot, solving challenging problems at roughly 40 tok/sec on a Mac or DGX Spark, comparable to GPT 5.5. The author compares harness choices, noting Claude Code uses twice as many tokens as Codex, and includes Gemma 4 E2B as a reference to show smaller models struggle. A full write-up is available, with a chart comparing token use and task success across five local-agent tasks.

Local Open-Weight LLMs in Coding Harnesses

Comments

No comments yet

Be the first to share your thoughts!

Browser Extension

Get instant access to AllDevBlogs from your browser

Top of the Week

No top articles yet