Commit Graph
3 Commits
Author SHA1 Message Date
bhethermanandClaude Sonnet 5 966140c3ef Give the MTP llama-server 4 parallel slots instead of 1
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Canceled after 0s
Frontend Build / Format & Build (push) Canceled after 0s
Frontend Build / Unit Tests (push) Canceled after 0s
Release to PyPI / release (push) Canceled after 0s
Release / publish (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Canceled after 0s
Create and publish Docker images with specific build args / notify-helm-charts (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Canceled after 0s
With -np 1, open-webui's browser chat and the voice assistant (and
anything else hitting this OpenAI-compatible endpoint) contended for
a single slot -- whichever touched it last evicted the other's live
KV cache, forcing a multi-second full prompt reprocess on the next
request from the loser.

n_ctx_slot = n_ctx / n_parallel, so ctx is quadrupled to 131072 (this
model's native n_ctx_train) to keep each of the 4 slots at the same
32768 budget as before. Measured KV cache cost is only 271 MiB per
slot, so this only costs ~810 MiB more VRAM than the old config.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 13:40:50 -04:00
bhethermanandClaude Sonnet 5 6f583deff9 Add relaxed-MTP llama-server as a new connection, surface generation speed
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Canceled after 0s
Frontend Build / Format & Build (push) Canceled after 0s
Frontend Build / Unit Tests (push) Canceled after 0s
Release to PyPI / release (push) Canceled after 0s
Release / publish (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Canceled after 0s
Create and publish Docker images with specific build args / notify-helm-charts (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Canceled after 0s
Wires the custom relaxed-acceptance llama-server (see
../mtp-relaxed-decoding) into the stack as a new llama-mtp service,
registered as an additional OpenAI-compatible connection alongside
Ollama. ollama-auth now proxies to llama-mtp instead of the now-empty
Ollama, so LAN clients (e.g. Home Assistant's voice pipeline) keep
working against the same URL/token with no reconfiguration.

Also merges llama.cpp's `timings` extension (dropped by the generic
OpenAI response schema) into message.usage and surfaces it as a
labeled tok/s badge, so relaxed-MTP responses get the same visible
generation-speed info Ollama responses already had.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-04 11:33:12 -04:00
bhetherman 7eff527d16 Add Authentik OIDC login, ollama-auth LAN proxy, and asr/tts sidecar wiring
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Canceled after 0s
Create and publish Docker images with specific build args / notify-helm-charts (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Canceled after 0s
Frontend Build / Format & Build (push) Canceled after 0s
Frontend Build / Unit Tests (push) Canceled after 0s
Release to PyPI / release (push) Canceled after 0s
Release / publish (push) Canceled after 0s
- open-webui: OIDC-only auth via Authentik (password/local login disabled),
  account linking by email
- ollama: GPU reservation, disable idle unload, read-only GGUF import mount
- ollama-auth: nginx bearer-token proxy in front of ollama for LAN access
- docker-compose.audio.yaml + docker-stack.sh: wrap base/gpu/audio compose
  files so asr/tts sidecars (../asr-server, ../tts-server) always come up
  together with the GPU override
2026-08-31 00:46:48 -04:00