🧠 Someone Asked Me to Add a 975B Model to a LocaLLMs, here’s Why we Can’t (Yet)
Explains why a 975B-parameter MoE multimodal model can't run locally in .NET via ONNX Runtime GenAI due to size, architecture, and conversion limits.
Explains why a 975B-parameter MoE multimodal model can't run locally in .NET via ONNX Runtime GenAI due to size, architecture, and conversion limits.
Martin Fowler shares fragments from the second Future of Software Development Retreat, focusing on agentic development and its impact on software engineering.
Explores the trend of smart model routing to reduce AI costs by automatically selecting the best LLM for each task.
Tutorial on setting up a local coding agent using open-weight LLMs and open-source tools as an alternative to Claude Code and Codex.
Analysis of VibeThinker-3B, a small AI model achieving strong coding and reasoning results through post-training on Qwen2.5-Coder-3B.
A critical take on AI, arguing it's not a person and shouldn't be anthropomorphized.
A plain English guide explaining how LLMs like ChatGPT actually work, covering prediction, vectors, embeddings, and common misconceptions.
Explores whether learning Machine Learning is still valuable in 2026 despite the rise of LLMs, highlighting data control, cost, and human oversight.
Explores how relying on AI for coding without active learning degrades skills, backed by research showing cognitive debt.
Explores 'the inversion' concept from YouTube engineers in 2013, where bot traffic threatened to surpass human views, challenging algorithm design.
A reflection on the Mythos AI model, arguing it's just the next step in rapid AI evolution, not a shocking breakthrough.
An expert discusses the overhyped risks and data limitations of applying AI in life sciences, using examples like AlphaFold.
Explores the balance between model simplicity and precision in system identification for control engineering and machine learning.
Argues for using general SOTA AI models over custom, specialized ones, predicting cheaper, open-source general models will dominate.
A critique of how quantitative benchmarking and evaluation culture shapes and potentially distorts progress in machine learning research.
Explores how autonomous AI agents (autoresearch) can optimize small language models by running hundreds of experiments overnight, improving performance without human intervention.
Analyzes the hidden costs and skill erosion of using AI for coding, emphasizing the need for human oversight.
Exploring the UMAP-MLX project, which achieves up to 46x speedups for UMAP using Apple's MLX, with performance benchmarks.
A skeptic's detailed journey using AI coding agents for complex projects, including porting scikit-learn to Rust, showcasing their surprising capabilities.
A skeptic's detailed journey using AI coding agents to port scikit-learn to Rust, showcasing the surprising capabilities of modern LLMs.