Compare commits

...
1 Commits
Author SHA1 Message Date
bhethermanandClaude Sonnet 5 966140c3ef Give the MTP llama-server 4 parallel slots instead of 1
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Canceled after 0s
Frontend Build / Format & Build (push) Canceled after 0s
Frontend Build / Unit Tests (push) Canceled after 0s
Release to PyPI / release (push) Canceled after 0s
Release / publish (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Canceled after 0s
Create and publish Docker images with specific build args / notify-helm-charts (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Canceled after 0s
With -np 1, open-webui's browser chat and the voice assistant (and
anything else hitting this OpenAI-compatible endpoint) contended for
a single slot -- whichever touched it last evicted the other's live
KV cache, forcing a multi-second full prompt reprocess on the next
request from the loser.

n_ctx_slot = n_ctx / n_parallel, so ctx is quadrupled to 131072 (this
model's native n_ctx_train) to keep each of the 4 slots at the same
32768 budget as before. Measured KV cache cost is only 271 MiB per
slot, so this only costs ~810 MiB more VRAM than the old config.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 13:40:50 -04:00
+15 -2
View File
@@ -25,8 +25,21 @@ services:
# was raised to match the README since Gemma-4's hybrid SWA # was raised to match the README since Gemma-4's hybrid SWA
# architecture (most layers use a small fixed attention window, only # architecture (most layers use a small fixed attention window, only
# every 5th layer is full-context) keeps KV cache growth with n_ctx # every 5th layer is full-context) keeps KV cache growth with n_ctx
# much cheaper than a plain transformer's. # much cheaper than a plain transformer's -- measured at only 271 MiB
- 'MTP_CTX_SIZE=32768' # of KV cache (target + draft) per 32768-token slot.
#
# -np 4 gives open-webui's browser chat and the voice assistant (and
# anything else hitting this OpenAI-compatible endpoint) their own
# slot each instead of one contending for a single shared slot --
# with -np 1, whichever of them touched the slot last would evict the
# other's live KV cache, forcing a full prompt reprocess (multi-second
# stall) on the next request from the loser. n_ctx_slot = n_ctx /
# n_parallel, so total ctx is quadrupled to keep each slot at the same
# 32768 budget as before -- 131072 matches this model's native
# n_ctx_train exactly, and only costs ~810 MiB more KV cache VRAM
# than the old -np 1 config (still well under the ~7GB that's free).
- 'MTP_CTX_SIZE=131072'
- 'MTP_N_PARALLEL=4'
- 'MTP_BATCH_SIZE=512' - 'MTP_BATCH_SIZE=512'
# Server-side default -- only applies when the client doesn't send its # Server-side default -- only applies when the client doesn't send its
# own temperature (Open WebUI won't unless you set it in the model's # own temperature (Open WebUI won't unless you set it in the model's