Add relaxed-MTP llama-server as a new connection, surface generation speed
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args: free_disk:false name:main suffix:]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/amd64 runner:ubuntu-latest], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args: free_disk:false name:main suffix:]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true USE_CUDA_VER=cu126 free_disk:true name:cuda126 suffix:-cuda126]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_CUDA=true free_disk:true name:cuda suffix:-cuda]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_OLLAMA=true free_disk:false name:ollama suffix:-ollama]) (push) Canceled after 0s
Create and publish Docker images with specific build args / build (map[arch:linux/arm64 runner:ubuntu-24.04-arm], map[build_args:USE_SLIM=true free_disk:false name:slim suffix:-slim]) (push) Canceled after 0s
Frontend Build / Format & Build (push) Canceled after 0s
Frontend Build / Unit Tests (push) Canceled after 0s
Release to PyPI / release (push) Canceled after 0s
Release / publish (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:cuda suffix:-cuda]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:cuda126 suffix:-cuda126]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:main suffix:]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:ollama suffix:-ollama]) (push) Canceled after 0s
Create and publish Docker images with specific build args / merge (map[name:slim suffix:-slim]) (push) Canceled after 0s
Create and publish Docker images with specific build args / notify-helm-charts (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (, main) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda, cuda) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-cuda126, cuda126) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-ollama, ollama) (push) Canceled after 0s
Create and publish Docker images with specific build args / copy-to-dockerhub (-slim, slim) (push) Canceled after 0s

Wires the custom relaxed-acceptance llama-server (see
../mtp-relaxed-decoding) into the stack as a new llama-mtp service,
registered as an additional OpenAI-compatible connection alongside
Ollama. ollama-auth now proxies to llama-mtp instead of the now-empty
Ollama, so LAN clients (e.g. Home Assistant's voice pipeline) keep
working against the same URL/token with no reconfiguration.

Also merges llama.cpp's `timings` extension (dropped by the generic
OpenAI response schema) into message.usage and surfaces it as a
labeled tok/s badge, so relaxed-MTP responses get the same visible
generation-speed info Ollama responses already had.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
2026-09-04 11:33:12 -04:00
co-authored by Claude Sonnet 5
parent 7eff527d16
commit 6f583deff9
6 changed files with 113 additions and 13 deletions
+8 -3
View File
@@ -2744,7 +2744,8 @@
};
const chatCompletionEventHandler = async (data, message, chatId) => {
const { id, done, choices, content, output, sources, selected_model_id, error, usage } = data;
const { id, done, choices, content, output, sources, selected_model_id, error, usage, timings } =
data;
// Store raw OR-aligned output items from backend
if (output) {
@@ -2797,8 +2798,12 @@
message.arena = true;
}
if (usage) {
message.usage = usage;
if (usage || timings) {
// llama.cpp-compatible servers (e.g. our relaxed-MTP llama-server) send
// generation speed as a sibling `timings` object rather than inside
// `usage` -- merge it in so it surfaces in the same response-info
// tooltip Ollama's eval_count/eval_duration fields already populate.
message.usage = { ...(usage || {}), ...(timings || {}) };
}
history.messages[message.id] = message;
@@ -144,6 +144,23 @@
}
}
// Generation speed, normalized across backends: Ollama reports
// eval_count/eval_duration (nanoseconds), llama.cpp-compatible servers
// (e.g. our relaxed-MTP llama-server) report timings.predicted_per_second
// directly -- note that field name is llama.cpp's generic term for "the
// main generation loop's output," i.e. the actual response stream, NOT
// the MTP drafter in isolation (which is separately broken out as
// draft_n/draft_n_accepted).
const getTokensPerSecond = (usage: Record<string, unknown> | undefined | null) => {
if (!usage) return null;
if (typeof usage.predicted_per_second === 'number') return usage.predicted_per_second;
if (typeof usage.eval_count === 'number' && typeof usage.eval_duration === 'number' && usage.eval_duration > 0) {
return (usage.eval_count / usage.eval_duration) * 1e9;
}
return null;
};
$: tokensPerSecond = getTokensPerSecond(message.usage);
export let siblings;
export let setInputText: Function = () => {};
@@ -1189,6 +1206,14 @@
</Tooltip>
{/if}
{#if tokensPerSecond}
<span
class="text-xs text-gray-400 dark:text-gray-500 px-1 self-center whitespace-nowrap"
>
{tokensPerSecond.toFixed(1)} tok/s
</span>
{/if}
{#if message.usage}
<Tooltip
content={message.usage