Ollama - Package Updates
-
[1.13.6]
- Update ollama to 0.30.8
- Full Changelog
- Fixed
ollama launchselecting the wrong provider in some cases - Improved prompt caching by decoupling it from context shift for better KV cache reuse
- More stable MLX inference with hardened linear and embedding layers
- MLX runner now creates snapshots during prompt processing and speculative decoding for improved reliability
- Improved recurrent model support with per-boundary states from the gated-delta kernels
- Full Changelog: https://github.com/ollama/ollama/compare/v0.30.7...v0.30.8
-
[1.13.7]
- Update ollama to 0.30.9
- Full Changelog
- Support for Cohere2Moe architecture
- Fixed LFM2 parser/render for cases where thinking was not emitted
- Fixed issue where
ollama launch claudeand other coding agent or assistant use cases would only output one token - Ollama will now return an error if a single message is larger than the current context window
- Full Changelog: https://github.com/ollama/ollama/compare/v0.30.8...v0.30.9-rc1
-
[1.13.8]
- Update ollama to 0.30.10
- Full Changelog
- models: add Cohere2MoE model by @jmorganca in #16670
- llama: update llama.cpp to b9672 by @pdevine in #16775
-
[1.13.9]
- Update ollama to 0.30.11
- Full Changelog
- launch: add thinking capability detection to opencode by @hoyyeva in #15434
- launch: auto-install Claude Code by @hoyyeva in #16802
- launch: auto-install opencode when missing by @hoyyeva in #16806
- discover: fix inverted iGPU/dGPU Vulkan classification on Windows hybrid graphics by @Sahil170595 in #16669
- mlxrunner: unify and tune speculative decoding by @jessegross in #16791
- launch/codex: detect model drift when Codex App UI switches by @BruceMacD in #16864
- llama: add sm_86 architecture to cuda_v13_windows preset by @anishesg in #16834
- llm: size mmproj offload by projector memory by @dhiltgen in #16866
- llm: preserve generation headroom for shifted prompts by @ParthSareen in #16856
- llm: fix ollama ps double-counting mmap'd weights on partial offload by @discobot in #16709
-
[1.14.0]
- Update ollama to 0.31.1
- Full Changelog
- Tightened Gemma 4 MoE model loading in the MLX engine
- Updated the MLX engine to the latest version, including a new small-batch matmul kernel
- Updated the underlying llama.cpp engine to build 9840
- Improved Gemma 4 multi-token prediction (MTP) performance
-
[1.15.0]
- Update ollama to 0.32.0
- Full Changelog
- New interactive agent experience: running
ollamanow launches an agent to help you code and delegate work - Renamed the Codex App integration to ChatGPT: use ollama launch chatgpt (and --restore to return to your usual ChatGPT profile)
- Simplified integration selection: the ollama launch menu now only offers the most popular integrations (other integrations can be accessed through
ollama launch - Warns before launching older agent models: CodeLlama, Qwen2.5(-coder), Llama 3.x, Mistral, StarCoder, and the base DeepSeek-R1 tags now prompt a deprecation warning before ollama launch continues
- Full Changelog: https://github.com/ollama/ollama/compare/v0.31.2...v0.32.0
-
[1.15.1]
- Update ollama to 0.32.1
- Full Changelog
- Improved Gemma 4 tool calling and multi-turn reasoning, including more reliable tool-response continuations
- Fixed a recurrent MLX model cache leak that could increase memory use across requests, and improved cache snapshot performance
- MLX text model loading now respects
OLLAMA_LOAD_TIMEOUT - Agent web search and fetch now tell users to run
ollama signinwhen authentication is required - The interactive agent now receives the current working directory for better project context
- Fixed
ollama launchso choosing Pick another model for a deprecated model passed with--modelopens the model picker - Updated VS Code setup documentation for the official Ollama extension
- Full Changelog: https://github.com/ollama/ollama/compare/v0.32.0...v0.32.1-rc0
-
[1.15.2]
- Update ollama to 0.32.3
- Full Changelog
- Fixed model downloads that stall before sending data.
- Improved integrations: restored Claude Code Channels, fixed Anthropic thinking streams, and made Hermes Desktop respect
--force-build. - Expanded GPU support with CUDA on Windows ARM64, B200 support through CUDA 12, and lower memory use on Linux CUDA/ROCm iGPUs.
- Added chat, thinking, and tool calling support for Laguna 2.1 models, including a Metal inference fix.
- Fixed GLM tool calls being silently dropped at the end of generation.
- Updated the MLX and llama.cpp engines.
- Full Changelog: https://github.com/ollama/ollama/compare/v0.32.1...v0.32.3
-
[1.15.3]
- Update ollama to 0.32.4
- Full Changelog
- Support Laguna on Apple GPUs via the MLX engine
- Quantize draft-model output heads at the requested type when creating speculative-decoding drafts.
- Fixed Qwen3 MoE decoding for differently-quantized experts, plus faster packed gate/up projection (~49% on M5 Max).
- Full Changelog: https://github.com/ollama/ollama/compare/v0.32.3...v0.32.4
-
[1.15.4]
- Update ollama to 0.32.5
- Full Changelog
- Fixed an MLX Metal bug that could reduce output quality for NVFP4 models, particularly Laguna.
Hello! It looks like you're interested in this conversation, but you don't have an account yet.
Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.
With your input, this post could be even better 💗
Register Login