entrypoint.sh now honors MTP_N_PARALLEL (default 1, unchanged) for -np instead of hardcoding 1, so open-webui's docker-compose.mtp.yaml can request multiple llama-server slots. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Relaxed acceptance for MTP speculative decoding — full guide
This documents exactly how we built a custom llama.cpp with a "relaxed" speculative
decoding acceptance criterion, and used it to get a real throughput win from Gemma 4
E2B's MTP (multi-token-prediction) drafter — something the stock config never achieved.
End result: n_max=2 + our new --spec-draft-relaxed-top-n 5 (roughly 2-20 all work
similarly) gave ~260-275 tok/s, beating both Ollama's stock MTP defaults (~205-220
tok/s) and even the plain non-speculative baseline (~250-260 tok/s). This is the only
configuration in the whole investigation where MTP paid for itself.
Everything below assumes Ubuntu 24.04 (this was done on WSL2) with an NVIDIA GPU. Adjust package/arch specifics for other distros.
0. Why this exists (skip if you just want the steps)
Ollama's own MTP support (DRAFT directive in a Modelfile) works, but:
- It hardcodes
--spec-draft-n-max 4with no way to change it from Ollama. - The underlying llama.cpp acceptance rule is strict: a drafted token is only accepted if it exactly equals a fresh sample drawn from the target model's own distribution at that position. One mismatch anywhere and the whole draft batch beyond that point is thrown away.
- For our target/drafter pair (
unsloth/gemma-4-E2B-it-GGUF, targetQ4_K_XL, draftermtp-gemma-4-E2B-it-Q8_0.gguf), that strict rule gave only ~18-24% acceptance, so MTP was slower than not using it at all.
We patched llama.cpp to add a genuinely relaxed rule — accept a drafted token if it ranks in the target's top-N candidates by probability, not just if it's the literal one token the target happened to sample. That's not exposed anywhere in stock llama.cpp or Ollama; it required a source patch.
1. Prerequisites
1.1 Ollama with a target+drafter model already built
You need Ollama running with a target model and its MTP drafter combined via a
Modelfile's DRAFT directive. Two ways to get the GGUF files, both from the same
public (non-gated) Hugging Face repo:
mkdir -p ~/mtp-build && cd ~/mtp-build
curl -L -o target.gguf \
"https://huggingface.co/unsloth/gemma-4-E2B-it-GGUF/resolve/main/gemma-4-E2B-it-UD-Q4_K_XL.gguf"
curl -L -o draft.gguf \
"https://huggingface.co/unsloth/gemma-4-E2B-it-GGUF/resolve/main/MTP/mtp-gemma-4-E2B-it-Q8_0.gguf"
cat > Modelfile <<'EOF'
FROM ./target.gguf
DRAFT ./draft.gguf
EOF
ollama create gemma4:e2b-mtp -f Modelfile
No --experimental flag and no gated Hugging Face access needed — that's only required
if you build from raw safetensors instead of a pre-quantized GGUF. Swap in whatever
target/drafter GGUF pair you actually want to test; the pairing mechanism is the same.
Sanity check it actually engages speculative decoding:
journalctl -u ollama --no-pager -o cat | grep -i "draft-mtp\|draft acceptance"
You should see lines like:
spec common_specu: adding speculative implementation 'draft-mtp'
slot print_timing: id 0 | task N | draft acceptance = 0.XXXXX (...)
1.2 Find the exact llama.cpp commit your Ollama vendors
This matters — you want to patch the exact source your Ollama's llama-server was
built from, not upstream main (APIs drift).
/usr/local/lib/ollama/llama-server --version
Output looks like:
version: 0.3.0-dev (build 1, commit d222767c7)
That commit hash is a real commit in ggml-org/llama.cpp — Ollama vendors upstream
llama.cpp with some patches on top, and this hash identifies the base. Confirm it
resolves:
curl -s "https://api.github.com/search/commits?q=hash:<commit>+repo:ggml-org/llama.cpp" \
-H "Accept: application/vnd.github.cloak-preview" | python3 -m json.tool
For this run it was d222767c7a6516559a3f49e7721b6c6b1acc87b4
("kleidiai: Rework KleidiAI Build System/Integration (#26077)").
1.3 Install the CUDA build toolchain
Ubuntu's own nvidia-cuda-toolkit package (CUDA 12.0) is too old for recent GPU
architectures (e.g. Blackwell/RTX 50-series needs CUDA 12.8+). Use NVIDIA's own repo:
cd /tmp
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2404/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt-get update
# Toolkit only — NOT the "cuda" or "cuda-drivers" meta-package. On WSL2 the GPU driver
# is Windows-side; installing a Linux driver package would conflict with the passthrough
# that's already working. On bare-metal Linux this distinction may not matter, but the
# toolkit-only package is still what you want (you already have a working driver).
sudo apt-get install -y build-essential cmake git cuda-toolkit-12-8
Verify:
export PATH=/usr/local/cuda/bin:$PATH
nvcc --version # should report release 12.8
g++ --version # 13.x is fine; nvcc 12.8 supports up to gcc 13
cmake --version # 3.24+ needed for CMAKE_CUDA_ARCHITECTURES=native
Gotcha: if apt-get update/install fails with a dpkg/apt lock error, check what's
holding it before doing anything drastic:
ps -p <pid> -o pid,ppid,etime,cmd
If it's apt-daily.service stuck "activating" for an unreasonable amount of time (check
systemctl status apt-daily.service), that's a known, safe-to-clear stuck state:
sudo systemctl stop apt-daily.service apt-daily.timer
Then retry. If it's a fresh, short-lived unattended-upgrades process, just wait for it
to finish on its own — don't kill it.
2. Get the source and apply the patch
mkdir -p ~/llama-relaxed && cd ~/llama-relaxed
git init -q
git remote add origin https://github.com/ggml-org/llama.cpp.git
git fetch --depth 1 origin <exact-commit-hash-from-1.2>
git checkout -q FETCH_HEAD
Apply the patch (saved alongside this guide as relaxed-top-n.patch):
git apply /path/to/relaxed-top-n.patch
If git apply fails (source drifted from what we patched), see "The patch by hand"
below and apply the five changes manually — they're small and self-contained.
What the patch does
One new flag: --spec-draft-relaxed-top-n N (default 0 = disabled, original strict
behavior unchanged). When N > 0, a drafted token is accepted if it ranks among the
target's top-N candidates by probability at that position — not just if it's the one
token the target actually sampled.
Five files change:
| file | change |
|---|---|
common/common.h |
new int32_t relaxed_top_n = 0; field on common_params_speculative_draft |
common/sampling.h |
both common_sampler_sample_and_accept_n overloads gain a trailing int32_t relaxed_top_n = 0 param |
common/sampling.cpp |
the actual logic — see below |
common/arg.cpp |
new CLI flag, following the exact pattern of the existing --spec-draft-p-min |
tools/server/server-context.cpp |
the one call site, passes params_base.speculative.draft.relaxed_top_n through |
The core change, in common_sampler_sample_and_accept_n (common/sampling.cpp):
// Returns true if `token` ranks within the top `relaxed_top_n` candidates of `cur_p`
// by probability. O(cur_p.size) and correct regardless of `cur_p.sorted`, since it
// counts strictly-higher-probability entries rather than assuming sort order.
static bool relaxed_top_n_contains(const llama_token_data_array & cur_p, llama_token token, int32_t relaxed_top_n) {
float token_p = -1.0f;
for (size_t k = 0; k < cur_p.size; k++) {
if (cur_p.data[k].id == token) {
token_p = cur_p.data[k].p;
break;
}
}
if (token_p < 0.0f) {
return false; // token isn't even in the candidate set (e.g. filtered out by top_k/top_p/min_p)
}
int32_t rank = 0;
for (size_t k = 0; k < cur_p.size; k++) {
if (cur_p.data[k].p > token_p) {
rank++;
if (rank >= relaxed_top_n) {
return false;
}
}
}
return true;
}
...used inside the accept loop:
const bool exact = (draft[i] == id);
bool accept = exact;
if (!accept && relaxed_top_n > 0) {
accept = relaxed_top_n_contains(gsmpl->cur_p, draft[i], relaxed_top_n);
if (getenv("RELAXED_VERIFY_LOG")) {
fprintf(stderr, "[RELAXED_VERIFY] pos=%zu relaxed_top_n=%d cur_p.size=%zu cur_p.sorted=%d draft='%s'(id=%d) target_sample='%s'(id=%d) -> %s\n",
i, relaxed_top_n, gsmpl->cur_p.size, (int) gsmpl->cur_p.sorted,
common_token_to_piece(ctx, draft[i]).c_str(), draft[i],
common_token_to_piece(ctx, id).c_str(), id,
accept ? "ACCEPT(relaxed)" : "reject");
}
}
// if the draft token is accepted under the relaxed rule, emit *that* token
// (not `id`) — server-context.cpp fed the whole drafted sequence into the
// target as one batch, so downstream KV state and idxs[i+1]'s logits already
// assume draft[i] was the real token at this position.
const llama_token emitted = accept ? draft[i] : id;
common_sampler_accept(gsmpl, emitted, true);
result.push_back(emitted);
if (!accept) {
break;
}
The RELAXED_VERIFY_LOG env-gated debug print is left in intentionally (harmless when
unset) — it's how we verified the patch was actually doing something (see §5).
Correctness note worth understanding: in batched speculative verification, the
target processes the whole drafted sequence as one forward pass, so position i's
candidate distribution already reflects "what comes after draft[0..i-1]" regardless of
whether draft[i] itself gets accepted. That means accepting draft[i] under the
relaxed rule is self-consistent — as long as you push draft[i] (not the target's own
id) into the result and into common_sampler_accept, which is what the patch does.
gsmpl->cur_p is always populated before this point (via set_logits(), called
unconditionally at the top of common_sampler_sample()), including when
--spec-draft-backend-sampling is on — so the patch works with the fast GPU-sampling
path, not just the CPU fallback.
3. Build
cd ~/llama-relaxed
export PATH=/usr/local/cuda/bin:$PATH
cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=native -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=OFF
cmake --build build --config Release -j"$(nproc)" --target llama-server
CMAKE_CUDA_ARCHITECTURES=native auto-detects your installed GPU (needs CMake 3.24+).
Watch the configure output for a line like:
-- Replacing 120-real in CMAKE_CUDA_ARCHITECTURES_NATIVE with 120a-real
-- Using CMAKE_CUDA_ARCHITECTURES=120a-real ...
confirming it targeted your actual card (this was 120a-real for an RTX 5090 / Blackwell).
Build artifacts land in build/bin/, including llama-server and every shared library
it needs (libggml-cuda.so, libllama.so, etc.) co-located in the same directory.
This matters: ggml's dynamic backend loader scans the executable's own directory for
backend .so files, so having everything in one place means CUDA is auto-detected with
no extra environment variables. (If you ever need to point at a separate CUDA backend
.so, e.g. reusing Ollama's own bundled one instead of building your own, the relevant
var is GGML_BACKEND_PATH=/path/to/libggml-cuda.so — it must be the exact .so file
path, not a directory. Not needed for a self-built binary.)
Verify:
build/bin/llama-server --help | grep -A2 "relaxed-top-n"
Should show your new flag.
4. Run it standalone
Important: this flag only exists in your custom build. Ollama's own ollama serve/
ollama run will never pass it through — Ollama hardcodes its own speculative-decoding
flags with no user override for this parameter. To use the patch, run llama-server
directly, bypassing Ollama's process entirely (Ollama can keep running in the
background; they're independent processes and can coexist on the same GPU as long as
there's VRAM headroom).
#!/bin/bash
# run_relaxed.sh
BUILT_DIR=~/llama-relaxed/build/bin
TARGET=~/mtp-build/target.gguf # or wherever your target GGUF lives
DRAFT=~/mtp-build/draft.gguf # or wherever your drafter GGUF lives
export LD_LIBRARY_PATH="$BUILT_DIR:$LD_LIBRARY_PATH"
exec "$BUILT_DIR/llama-server" \
--model "$TARGET" \
--port "$PORT" --host 127.0.0.1 --no-webui --offline -c 32768 -np 1 \
--log-verbosity 4 --no-log-prefix --no-log-timestamps --no-jinja --chat-template chatml \
--spec-type draft-mtp --spec-draft-n-max "$NMAX" --spec-draft-relaxed-top-n "$RELAXED" \
--spec-draft-backend-sampling \
--spec-draft-model "$DRAFT" -ngl 999 --spec-draft-ngl 999 \
--flash-attn auto -b 1024 -ub 1024 --context-shift --keep 4
chmod +x run_relaxed.sh
PORT=38700 NMAX=2 RELAXED=5 ./run_relaxed.sh
Wait for GET /health to return {"status":"ok"}, then hit its native /completion
endpoint directly (this is llama.cpp's own API, not Ollama's /api/generate):
curl -s http://127.0.0.1:38700/completion -d '{
"prompt": "Write a short poem about code.",
"n_predict": 128,
"temperature": 0.6,
"top_k": 64,
"top_p": 0.95,
"min_p": 0.01,
"cache_prompt": false
}'
If you reuse Ollama's already-downloaded blob files instead of your own copies (they
live under /usr/share/ollama/.ollama/models/blobs/sha256-<hash>, filenames = the
manifest hashes from ollama show <model> --modelfile or captured from a running
process's cmdline), you'll hit a permission error — those files are owned by the
ollama user/group. Two options: copy your own files instead (simplest), or run via
sg ollama -c ./run_relaxed.sh if your user is in the ollama group (check with
getent group ollama; newly-added group membership needs a fresh login shell, sg lets
you pick it up without one). Don't nest sg calls (sg a -c "sg b -c ...") — it
tends to trigger an interactive password prompt that hangs non-interactive invocations.
5. Verify it's actually working — don't just trust the throughput numbers
Throughput/acceptance-rate changes alone don't prove the patch is doing what you think. Verify directly:
Control test: run with RELAXED=0 on your new build and confirm behavior matches
your original (unpatched) binary's strict-mode numbers. This rules out "the rebuild
changed something unrelated." Also confirm zero RELAXED_VERIFY_LOG lines appear
(proves the new code path is genuinely inert when disabled — it's gated behind
relaxed_top_n > 0).
Positive test: run with RELAXED_VERIFY_LOG=1 RELAXED=<something> and grep the
output:
RELAXED_VERIFY_LOG=1 PORT=38700 NMAX=2 RELAXED=5 ./run_relaxed.sh > server.log 2>&1 &
# ...make some completion requests...
grep "RELAXED_VERIFY" server.log
grep -c "ACCEPT(relaxed)" server.log
grep -c "reject" server.log
You should see a genuine mix of accepts and rejects (not all-or-nothing), and the
specific token pairs should look semantically sensible when you eyeball them — e.g.
draft=' Code' target_sample=' Algorithm' -> ACCEPT is a plausible substitution;
draft=' Journey' target_sample=' Whisper' -> reject is a sensible rejection.
Stress test the extremes: set relaxed_top_n absurdly high (e.g. 300000, bigger
than the vocab) and disable min_p in the request ("min_p": 0.0 — see the gotcha
below) to force every draft token accepted. You should see draft acceptance = 1.00000
in the logs and — critically — coherent-but-degenerate output (ours fell into
repetition loops), not crashes or garbled/invalid tokens. A clean failure mode there is
good evidence the KV-state handling (emitting draft[i] correctly) is sound; a crash or
corrupted text would mean the patch has a real bug.
Read the actual completion text, not just the stats, at every step. Coherent
prose = the relaxed criterion isn't silently corrupting output; a still-passable poem
at relaxed_top_n=1000+ is a meaningful signal, not just noise.
Gotcha: min_p silently caps your candidate pool
The sampler chain's default min_p (0.05 in llama-server) aggressively prunes
gsmpl->cur_p down to a handful of candidates (we observed sizes of 1-5) before the
relaxed check ever sees it — regardless of top_k/top_p on the request, and
regardless of --spec-draft-backend-sampling on/off. If you're testing what
relaxed_top_n values actually do, explicitly set "min_p": 0.01 (or 0.0) in every
request — otherwise you're comparing relaxed_top_n values against a candidate pool
that's already artificially tiny, which was a real methodological mistake we made and
had to redo. Add a size/sorted field to the debug log (as above) if you ever need to
double-check the candidate pool size directly instead of inferring it.
6. What we found (this specific target/drafter pair, RTX 5090)
All numbers: gemma-4-E2B-it-UD-Q4_K_XL.gguf target + mtp-gemma-4-E2B-it-Q8_0.gguf
drafter, temperature 0.6-1.0 (didn't matter much), min_p=0.01, 128-token completions,
5-run averages after a warm-up request.
| config | tok/s | acceptance |
|---|---|---|
| plain target, no speculative decoding | ~250-260 | — |
n_max=4 strict (Ollama's hardcoded default) |
~205-220 | ~18-24% |
n_max=2 strict |
~230-245 | ~30-36% |
n_max=16, p_min=0.8 (a value recommended online for a different setup) |
~124-140 | ~55-70% (but slow — bigger verify batches cost more than they save) |
n_max=2 + relaxed_top_n=2-20 (this patch) |
~260-275 | ~38-53% |
n_max=4/8/16 + relaxed_top_n |
progressively worse (~230→130) | mean accepted length stays flat ~2.3-2.7 regardless — extra draft depth is wasted compute |
Takeaways that seemed to hold up across every variant we tried (temperature 0/0.1/0.6/1.0,
min_p 0.0/0.01/0.05, p_min, target quantization Q4 vs Q8):
n_maxbeyond ~2 is not worth it for this drafter — its useful predictive depth caps out around position 2-3 regardless of acceptance rule, temperature, or quantization. Biggern_maxjust makes verification batches more expensive without a proportional increase in useful accepted tokens.- The relaxed criterion is the actual lever that made MTP worthwhile here. Every strict-mode config we tried underperformed the non-speculative baseline; every reasonable relaxed config beat it.
- There's a plateau, not a sharp peak, across
relaxed_top_n ≈ 2-20— differences within that range are close to sample noise at n=5 runs. Don't over-index on the exact best value from a small sweep; re-run a few times before trusting a "winner." - Quantization match between target/drafter (Q4 vs Q8 target) and temperature both turned out to be second-order effects, not the dominant factor, contrary to our initial (wrong) hypotheses — worth re-testing on your own hardware/model rather than assuming these transfer.
Appendix: files in this folder
relaxed-top-n.patch— the exact diff against llama.cpp commitd222767c7, apply withgit apply relaxed-top-n.patchfrom a clean checkout of that commit.README.md— this guide.