Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Canceled after 0s
Frontend Build / Format & Build (push) Canceled after 0s
Frontend Build / Unit Tests (push) Canceled after 0s
Release to PyPI / release (push) Canceled after 0s
Release / publish (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Canceled after 0s
Create and publish Docker images with specific build args / notify-helm-charts (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Canceled after 0s
Wires the custom relaxed-acceptance llama-server (see ../mtp-relaxed-decoding) into the stack as a new llama-mtp service, registered as an additional OpenAI-compatible connection alongside Ollama. ollama-auth now proxies to llama-mtp instead of the now-empty Ollama, so LAN clients (e.g. Home Assistant's voice pipeline) keep working against the same URL/token with no reconfiguration. Also merges llama.cpp's `timings` extension (dropped by the generic OpenAI response schema) into message.usage and surfaces it as a labeled tok/s badge, so relaxed-MTP responses get the same visible generation-speed info Ollama responses already had. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
40 lines
1.4 KiB
Plaintext
40 lines
1.4 KiB
Plaintext
map_hash_bucket_size 128;
|
|
|
|
map $http_authorization $ollama_auth_ok {
|
|
"Bearer ${OLLAMA_AUTH_TOKEN}" 1;
|
|
default 0;
|
|
}
|
|
|
|
server {
|
|
listen 11434;
|
|
|
|
location / {
|
|
if ($ollama_auth_ok = 0) {
|
|
return 401;
|
|
}
|
|
|
|
# Ollama itself now serves no models -- this proxy targets
|
|
# llama-mtp (the relaxed-MTP server) instead, so LAN clients that
|
|
# already trust this URL/token (e.g. Home Assistant's voice
|
|
# pipeline) don't need reconfiguring. Docker's embedded DNS
|
|
# (127.0.0.11) can reassign a container hostname to a new IP on
|
|
# restart; proxy_pass to a literal upstream caches that IP for the
|
|
# container's lifetime, so route through a variable to force
|
|
# nginx to re-resolve via this resolver on each request instead.
|
|
resolver 127.0.0.11 valid=10s;
|
|
set $mtp_upstream llama-mtp;
|
|
proxy_pass http://$mtp_upstream:8030;
|
|
proxy_http_version 1.1;
|
|
proxy_set_header Connection "";
|
|
proxy_set_header Host $host;
|
|
# Client presents OLLAMA_AUTH_TOKEN (validated above); llama-server
|
|
# behind this proxy expects its own MTP_AUTH_TOKEN instead, so swap
|
|
# it here rather than requiring every client to be reconfigured.
|
|
proxy_set_header Authorization "Bearer ${MTP_AUTH_TOKEN}";
|
|
proxy_buffering off;
|
|
proxy_read_timeout 600s;
|
|
proxy_send_timeout 600s;
|
|
client_max_body_size 0;
|
|
}
|
|
}
|