Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true
USE_CUDA_VER=cu126
free_disk:true name:cuda126 suffix:-cuda126]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Canceled after 0s
Frontend Build / Format & Build (push) Canceled after 0s
Frontend Build / Unit Tests (push) Canceled after 0s
Release to PyPI / release (push) Canceled after 0s
Release / publish (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Canceled after 0s
Create and publish Docker images with specific build args / notify-helm-charts (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Canceled after 0s
Wires the custom relaxed-acceptance llama-server (see ../mtp-relaxed-decoding) into the stack as a new llama-mtp service, registered as an additional OpenAI-compatible connection alongside Ollama. ollama-auth now proxies to llama-mtp instead of the now-empty Ollama, so LAN clients (e.g. Home Assistant's voice pipeline) keep working against the same URL/token with no reconfiguration. Also merges llama.cpp's `timings` extension (dropped by the generic OpenAI response schema) into message.usage and surfaces it as a labeled tok/s badge, so relaxed-MTP responses get the same visible generation-speed info Ollama responses already had. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
60 lines
2.3 KiB
YAML
60 lines
2.3 KiB
YAML
services:
|
|
llama-mtp:
|
|
build:
|
|
context: ../mtp-relaxed-decoding
|
|
container_name: llama-mtp
|
|
restart: unless-stopped
|
|
ports:
|
|
- '8030:8030'
|
|
volumes:
|
|
# Same host dir ollama's /gguf-import mount points at -- target +
|
|
# MTP drafter already live there.
|
|
- /home/brian-llm/models/gguf:/models:ro
|
|
environment:
|
|
- 'MTP_AUTH_TOKEN=${MTP_AUTH_TOKEN}'
|
|
- 'MTP_TARGET_GGUF=/models/gemma-4-E2B-it-Q4_K_M.gguf'
|
|
- 'MTP_DRAFT_GGUF=/models/mtp-gemma-4-E2B-it.gguf'
|
|
# n_max beyond ~2 wasn't worth it for this drafter, and relaxed_top_n
|
|
# 2-20 all performed similarly -- see ../mtp-relaxed-decoding/README.md
|
|
# §6. Benchmarked on an RTX 5090 there; re-check on this Titan Xp
|
|
# before trusting these as tuned rather than just reasonable defaults.
|
|
- 'MTP_N_MAX=2'
|
|
- 'MTP_RELAXED_TOP_N=5'
|
|
# Titan Xp has 12GB VRAM shared with asr/tts/ollama -- batch size kept
|
|
# below the README's RTX-5090 sizing (1024) for headroom, but context
|
|
# was raised to match the README since Gemma-4's hybrid SWA
|
|
# architecture (most layers use a small fixed attention window, only
|
|
# every 5th layer is full-context) keeps KV cache growth with n_ctx
|
|
# much cheaper than a plain transformer's.
|
|
- 'MTP_CTX_SIZE=32768'
|
|
- 'MTP_BATCH_SIZE=512'
|
|
# Server-side default -- only applies when the client doesn't send its
|
|
# own temperature (Open WebUI won't unless you set it in the model's
|
|
# Advanced Params).
|
|
- 'MTP_TEMP=0.6'
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
count: 1
|
|
capabilities: [gpu]
|
|
|
|
open-webui:
|
|
depends_on:
|
|
- llama-mtp
|
|
environment:
|
|
# Registers this alongside (not instead of) the Ollama connection --
|
|
# Open WebUI lists both in the model picker.
|
|
- 'OPENAI_API_BASE_URLS=http://llama-mtp:8030/v1'
|
|
- 'OPENAI_API_KEYS=${MTP_AUTH_TOKEN}'
|
|
|
|
ollama-auth:
|
|
depends_on:
|
|
- llama-mtp
|
|
environment:
|
|
# ollama-auth/default.conf.template now proxies to llama-mtp (Ollama
|
|
# itself serves no models anymore) and swaps the client's
|
|
# OLLAMA_AUTH_TOKEN for this before forwarding upstream.
|
|
- 'MTP_AUTH_TOKEN=${MTP_AUTH_TOKEN}'
|