Energy use of AI inference – estimates and efficiency opportunities
Read OriginalThis article examines the energy consumption of AI inference, highlighting the difficulty in comparing estimates from major companies like OpenAI, Google, and Mistral due to differing boundaries and workloads. It introduces Microsoft's new paper (Oviedo et al., 2026) which proposes a bottom-up framework for estimating inference energy under production-like conditions, including optimized serving, batching, and concurrency. Key findings show that standard frontier-model inference uses less energy than public estimates suggest, but long reasoning and agentic workloads can increase energy by up to 13x. The article discusses the impact of output tokens, model size, and the hidden expansion of agentic tasks, providing a nuanced view of AI's energy footprint.
Comments
No comments yet
Be the first to share your thoughts!
Browser Extension
Get instant access to AllDevBlogs from your browser