Files
bhethermanandClaude Sonnet 5 0ba2ea45be Add MTP_N_PARALLEL to entrypoint.sh, bump submodule pointers
entrypoint.sh now honors MTP_N_PARALLEL (default 1, unchanged) for
-np instead of hardcoding 1, so open-webui's docker-compose.mtp.yaml
can request multiple llama-server slots.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-05 13:40:59 -04:00
..

Relaxed acceptance for MTP speculative decoding — full guide

This documents exactly how we built a custom llama.cpp with a "relaxed" speculative decoding acceptance criterion, and used it to get a real throughput win from Gemma 4 E2B's MTP (multi-token-prediction) drafter — something the stock config never achieved.

End result: n_max=2 + our new --spec-draft-relaxed-top-n 5 (roughly 2-20 all work similarly) gave ~260-275 tok/s, beating both Ollama's stock MTP defaults (~205-220 tok/s) and even the plain non-speculative baseline (~250-260 tok/s). This is the only configuration in the whole investigation where MTP paid for itself.

Everything below assumes Ubuntu 24.04 (this was done on WSL2) with an NVIDIA GPU. Adjust package/arch specifics for other distros.


0. Why this exists (skip if you just want the steps)

Ollama's own MTP support (DRAFT directive in a Modelfile) works, but:

  • It hardcodes --spec-draft-n-max 4 with no way to change it from Ollama.
  • The underlying llama.cpp acceptance rule is strict: a drafted token is only accepted if it exactly equals a fresh sample drawn from the target model's own distribution at that position. One mismatch anywhere and the whole draft batch beyond that point is thrown away.
  • For our target/drafter pair (unsloth/gemma-4-E2B-it-GGUF, target Q4_K_XL, drafter mtp-gemma-4-E2B-it-Q8_0.gguf), that strict rule gave only ~18-24% acceptance, so MTP was slower than not using it at all.

We patched llama.cpp to add a genuinely relaxed rule — accept a drafted token if it ranks in the target's top-N candidates by probability, not just if it's the literal one token the target happened to sample. That's not exposed anywhere in stock llama.cpp or Ollama; it required a source patch.


1. Prerequisites

1.1 Ollama with a target+drafter model already built

You need Ollama running with a target model and its MTP drafter combined via a Modelfile's DRAFT directive. Two ways to get the GGUF files, both from the same public (non-gated) Hugging Face repo:

mkdir -p ~/mtp-build && cd ~/mtp-build
curl -L -o target.gguf \
  "https://huggingface.co/unsloth/gemma-4-E2B-it-GGUF/resolve/main/gemma-4-E2B-it-UD-Q4_K_XL.gguf"
curl -L -o draft.gguf \
  "https://huggingface.co/unsloth/gemma-4-E2B-it-GGUF/resolve/main/MTP/mtp-gemma-4-E2B-it-Q8_0.gguf"

cat > Modelfile <<'EOF'
FROM ./target.gguf
DRAFT ./draft.gguf
EOF

ollama create gemma4:e2b-mtp -f Modelfile

No --experimental flag and no gated Hugging Face access needed — that's only required if you build from raw safetensors instead of a pre-quantized GGUF. Swap in whatever target/drafter GGUF pair you actually want to test; the pairing mechanism is the same.

Sanity check it actually engages speculative decoding:

journalctl -u ollama --no-pager -o cat | grep -i "draft-mtp\|draft acceptance"

You should see lines like:

spec common_specu: adding speculative implementation 'draft-mtp'
slot print_timing: id 0 | task N | draft acceptance = 0.XXXXX (...)

1.2 Find the exact llama.cpp commit your Ollama vendors

This matters — you want to patch the exact source your Ollama's llama-server was built from, not upstream main (APIs drift).

/usr/local/lib/ollama/llama-server --version

Output looks like:

version: 0.3.0-dev (build 1, commit d222767c7)

That commit hash is a real commit in ggml-org/llama.cpp — Ollama vendors upstream llama.cpp with some patches on top, and this hash identifies the base. Confirm it resolves:

curl -s "https://api.github.com/search/commits?q=hash:<commit>+repo:ggml-org/llama.cpp" \
  -H "Accept: application/vnd.github.cloak-preview" | python3 -m json.tool

For this run it was d222767c7a6516559a3f49e7721b6c6b1acc87b4 ("kleidiai: Rework KleidiAI Build System/Integration (#26077)").

1.3 Install the CUDA build toolchain

Ubuntu's own nvidia-cuda-toolkit package (CUDA 12.0) is too old for recent GPU architectures (e.g. Blackwell/RTX 50-series needs CUDA 12.8+). Use NVIDIA's own repo:

cd /tmp
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2404/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt-get update

# Toolkit only — NOT the "cuda" or "cuda-drivers" meta-package. On WSL2 the GPU driver
# is Windows-side; installing a Linux driver package would conflict with the passthrough
# that's already working. On bare-metal Linux this distinction may not matter, but the
# toolkit-only package is still what you want (you already have a working driver).
sudo apt-get install -y build-essential cmake git cuda-toolkit-12-8

Verify:

export PATH=/usr/local/cuda/bin:$PATH
nvcc --version    # should report release 12.8
g++ --version      # 13.x is fine; nvcc 12.8 supports up to gcc 13
cmake --version    # 3.24+ needed for CMAKE_CUDA_ARCHITECTURES=native

Gotcha: if apt-get update/install fails with a dpkg/apt lock error, check what's holding it before doing anything drastic:

ps -p <pid> -o pid,ppid,etime,cmd

If it's apt-daily.service stuck "activating" for an unreasonable amount of time (check systemctl status apt-daily.service), that's a known, safe-to-clear stuck state:

sudo systemctl stop apt-daily.service apt-daily.timer

Then retry. If it's a fresh, short-lived unattended-upgrades process, just wait for it to finish on its own — don't kill it.


2. Get the source and apply the patch

mkdir -p ~/llama-relaxed && cd ~/llama-relaxed
git init -q
git remote add origin https://github.com/ggml-org/llama.cpp.git
git fetch --depth 1 origin <exact-commit-hash-from-1.2>
git checkout -q FETCH_HEAD

Apply the patch (saved alongside this guide as relaxed-top-n.patch):

git apply /path/to/relaxed-top-n.patch

If git apply fails (source drifted from what we patched), see "The patch by hand" below and apply the five changes manually — they're small and self-contained.

What the patch does

One new flag: --spec-draft-relaxed-top-n N (default 0 = disabled, original strict behavior unchanged). When N > 0, a drafted token is accepted if it ranks among the target's top-N candidates by probability at that position — not just if it's the one token the target actually sampled.

Five files change:

file change
common/common.h new int32_t relaxed_top_n = 0; field on common_params_speculative_draft
common/sampling.h both common_sampler_sample_and_accept_n overloads gain a trailing int32_t relaxed_top_n = 0 param
common/sampling.cpp the actual logic — see below
common/arg.cpp new CLI flag, following the exact pattern of the existing --spec-draft-p-min
tools/server/server-context.cpp the one call site, passes params_base.speculative.draft.relaxed_top_n through

The core change, in common_sampler_sample_and_accept_n (common/sampling.cpp):

// Returns true if `token` ranks within the top `relaxed_top_n` candidates of `cur_p`
// by probability. O(cur_p.size) and correct regardless of `cur_p.sorted`, since it
// counts strictly-higher-probability entries rather than assuming sort order.
static bool relaxed_top_n_contains(const llama_token_data_array & cur_p, llama_token token, int32_t relaxed_top_n) {
    float token_p = -1.0f;
    for (size_t k = 0; k < cur_p.size; k++) {
        if (cur_p.data[k].id == token) {
            token_p = cur_p.data[k].p;
            break;
        }
    }
    if (token_p < 0.0f) {
        return false; // token isn't even in the candidate set (e.g. filtered out by top_k/top_p/min_p)
    }

    int32_t rank = 0;
    for (size_t k = 0; k < cur_p.size; k++) {
        if (cur_p.data[k].p > token_p) {
            rank++;
            if (rank >= relaxed_top_n) {
                return false;
            }
        }
    }
    return true;
}

...used inside the accept loop:

const bool exact = (draft[i] == id);
bool accept = exact;
if (!accept && relaxed_top_n > 0) {
    accept = relaxed_top_n_contains(gsmpl->cur_p, draft[i], relaxed_top_n);
    if (getenv("RELAXED_VERIFY_LOG")) {
        fprintf(stderr, "[RELAXED_VERIFY] pos=%zu relaxed_top_n=%d cur_p.size=%zu cur_p.sorted=%d draft='%s'(id=%d) target_sample='%s'(id=%d) -> %s\n",
                i, relaxed_top_n, gsmpl->cur_p.size, (int) gsmpl->cur_p.sorted,
                common_token_to_piece(ctx, draft[i]).c_str(), draft[i],
                common_token_to_piece(ctx, id).c_str(), id,
                accept ? "ACCEPT(relaxed)" : "reject");
    }
}

// if the draft token is accepted under the relaxed rule, emit *that* token
// (not `id`) — server-context.cpp fed the whole drafted sequence into the
// target as one batch, so downstream KV state and idxs[i+1]'s logits already
// assume draft[i] was the real token at this position.
const llama_token emitted = accept ? draft[i] : id;

common_sampler_accept(gsmpl, emitted, true);
result.push_back(emitted);

if (!accept) {
    break;
}

The RELAXED_VERIFY_LOG env-gated debug print is left in intentionally (harmless when unset) — it's how we verified the patch was actually doing something (see §5).

Correctness note worth understanding: in batched speculative verification, the target processes the whole drafted sequence as one forward pass, so position i's candidate distribution already reflects "what comes after draft[0..i-1]" regardless of whether draft[i] itself gets accepted. That means accepting draft[i] under the relaxed rule is self-consistent — as long as you push draft[i] (not the target's own id) into the result and into common_sampler_accept, which is what the patch does.

gsmpl->cur_p is always populated before this point (via set_logits(), called unconditionally at the top of common_sampler_sample()), including when --spec-draft-backend-sampling is on — so the patch works with the fast GPU-sampling path, not just the CPU fallback.


3. Build

cd ~/llama-relaxed
export PATH=/usr/local/cuda/bin:$PATH
cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=native -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=OFF
cmake --build build --config Release -j"$(nproc)" --target llama-server

CMAKE_CUDA_ARCHITECTURES=native auto-detects your installed GPU (needs CMake 3.24+). Watch the configure output for a line like:

-- Replacing 120-real in CMAKE_CUDA_ARCHITECTURES_NATIVE with 120a-real
-- Using CMAKE_CUDA_ARCHITECTURES=120a-real ...

confirming it targeted your actual card (this was 120a-real for an RTX 5090 / Blackwell).

Build artifacts land in build/bin/, including llama-server and every shared library it needs (libggml-cuda.so, libllama.so, etc.) co-located in the same directory. This matters: ggml's dynamic backend loader scans the executable's own directory for backend .so files, so having everything in one place means CUDA is auto-detected with no extra environment variables. (If you ever need to point at a separate CUDA backend .so, e.g. reusing Ollama's own bundled one instead of building your own, the relevant var is GGML_BACKEND_PATH=/path/to/libggml-cuda.so — it must be the exact .so file path, not a directory. Not needed for a self-built binary.)

Verify:

build/bin/llama-server --help | grep -A2 "relaxed-top-n"

Should show your new flag.


4. Run it standalone

Important: this flag only exists in your custom build. Ollama's own ollama serve/ ollama run will never pass it through — Ollama hardcodes its own speculative-decoding flags with no user override for this parameter. To use the patch, run llama-server directly, bypassing Ollama's process entirely (Ollama can keep running in the background; they're independent processes and can coexist on the same GPU as long as there's VRAM headroom).

#!/bin/bash
# run_relaxed.sh
BUILT_DIR=~/llama-relaxed/build/bin
TARGET=~/mtp-build/target.gguf      # or wherever your target GGUF lives
DRAFT=~/mtp-build/draft.gguf        # or wherever your drafter GGUF lives
export LD_LIBRARY_PATH="$BUILT_DIR:$LD_LIBRARY_PATH"
exec "$BUILT_DIR/llama-server" \
  --model "$TARGET" \
  --port "$PORT" --host 127.0.0.1 --no-webui --offline -c 32768 -np 1 \
  --log-verbosity 4 --no-log-prefix --no-log-timestamps --no-jinja --chat-template chatml \
  --spec-type draft-mtp --spec-draft-n-max "$NMAX" --spec-draft-relaxed-top-n "$RELAXED" \
  --spec-draft-backend-sampling \
  --spec-draft-model "$DRAFT" -ngl 999 --spec-draft-ngl 999 \
  --flash-attn auto -b 1024 -ub 1024 --context-shift --keep 4
chmod +x run_relaxed.sh
PORT=38700 NMAX=2 RELAXED=5 ./run_relaxed.sh

Wait for GET /health to return {"status":"ok"}, then hit its native /completion endpoint directly (this is llama.cpp's own API, not Ollama's /api/generate):

curl -s http://127.0.0.1:38700/completion -d '{
  "prompt": "Write a short poem about code.",
  "n_predict": 128,
  "temperature": 0.6,
  "top_k": 64,
  "top_p": 0.95,
  "min_p": 0.01,
  "cache_prompt": false
}'

If you reuse Ollama's already-downloaded blob files instead of your own copies (they live under /usr/share/ollama/.ollama/models/blobs/sha256-<hash>, filenames = the manifest hashes from ollama show <model> --modelfile or captured from a running process's cmdline), you'll hit a permission error — those files are owned by the ollama user/group. Two options: copy your own files instead (simplest), or run via sg ollama -c ./run_relaxed.sh if your user is in the ollama group (check with getent group ollama; newly-added group membership needs a fresh login shell, sg lets you pick it up without one). Don't nest sg calls (sg a -c "sg b -c ...") — it tends to trigger an interactive password prompt that hangs non-interactive invocations.


5. Verify it's actually working — don't just trust the throughput numbers

Throughput/acceptance-rate changes alone don't prove the patch is doing what you think. Verify directly:

Control test: run with RELAXED=0 on your new build and confirm behavior matches your original (unpatched) binary's strict-mode numbers. This rules out "the rebuild changed something unrelated." Also confirm zero RELAXED_VERIFY_LOG lines appear (proves the new code path is genuinely inert when disabled — it's gated behind relaxed_top_n > 0).

Positive test: run with RELAXED_VERIFY_LOG=1 RELAXED=<something> and grep the output:

RELAXED_VERIFY_LOG=1 PORT=38700 NMAX=2 RELAXED=5 ./run_relaxed.sh > server.log 2>&1 &
# ...make some completion requests...
grep "RELAXED_VERIFY" server.log
grep -c "ACCEPT(relaxed)" server.log
grep -c "reject" server.log

You should see a genuine mix of accepts and rejects (not all-or-nothing), and the specific token pairs should look semantically sensible when you eyeball them — e.g. draft=' Code' target_sample=' Algorithm' -> ACCEPT is a plausible substitution; draft=' Journey' target_sample=' Whisper' -> reject is a sensible rejection.

Stress test the extremes: set relaxed_top_n absurdly high (e.g. 300000, bigger than the vocab) and disable min_p in the request ("min_p": 0.0 — see the gotcha below) to force every draft token accepted. You should see draft acceptance = 1.00000 in the logs and — critically — coherent-but-degenerate output (ours fell into repetition loops), not crashes or garbled/invalid tokens. A clean failure mode there is good evidence the KV-state handling (emitting draft[i] correctly) is sound; a crash or corrupted text would mean the patch has a real bug.

Read the actual completion text, not just the stats, at every step. Coherent prose = the relaxed criterion isn't silently corrupting output; a still-passable poem at relaxed_top_n=1000+ is a meaningful signal, not just noise.

Gotcha: min_p silently caps your candidate pool

The sampler chain's default min_p (0.05 in llama-server) aggressively prunes gsmpl->cur_p down to a handful of candidates (we observed sizes of 1-5) before the relaxed check ever sees it — regardless of top_k/top_p on the request, and regardless of --spec-draft-backend-sampling on/off. If you're testing what relaxed_top_n values actually do, explicitly set "min_p": 0.01 (or 0.0) in every request — otherwise you're comparing relaxed_top_n values against a candidate pool that's already artificially tiny, which was a real methodological mistake we made and had to redo. Add a size/sorted field to the debug log (as above) if you ever need to double-check the candidate pool size directly instead of inferring it.


6. What we found (this specific target/drafter pair, RTX 5090)

All numbers: gemma-4-E2B-it-UD-Q4_K_XL.gguf target + mtp-gemma-4-E2B-it-Q8_0.gguf drafter, temperature 0.6-1.0 (didn't matter much), min_p=0.01, 128-token completions, 5-run averages after a warm-up request.

config tok/s acceptance
plain target, no speculative decoding ~250-260 —
n_max=4 strict (Ollama's hardcoded default) ~205-220 ~18-24%
n_max=2 strict ~230-245 ~30-36%
n_max=16, p_min=0.8 (a value recommended online for a different setup) ~124-140 ~55-70% (but slow — bigger verify batches cost more than they save)
n_max=2 + relaxed_top_n=2-20 (this patch) ~260-275 ~38-53%
n_max=4/8/16 + relaxed_top_n progressively worse (~230→130) mean accepted length stays flat ~2.3-2.7 regardless — extra draft depth is wasted compute

Takeaways that seemed to hold up across every variant we tried (temperature 0/0.1/0.6/1.0, min_p 0.0/0.01/0.05, p_min, target quantization Q4 vs Q8):

  • n_max beyond ~2 is not worth it for this drafter — its useful predictive depth caps out around position 2-3 regardless of acceptance rule, temperature, or quantization. Bigger n_max just makes verification batches more expensive without a proportional increase in useful accepted tokens.
  • The relaxed criterion is the actual lever that made MTP worthwhile here. Every strict-mode config we tried underperformed the non-speculative baseline; every reasonable relaxed config beat it.
  • There's a plateau, not a sharp peak, across relaxed_top_n ≈ 2-20 — differences within that range are close to sample noise at n=5 runs. Don't over-index on the exact best value from a small sweep; re-run a few times before trusting a "winner."
  • Quantization match between target/drafter (Q4 vs Q8 target) and temperature both turned out to be second-order effects, not the dominant factor, contrary to our initial (wrong) hypotheses — worth re-testing on your own hardware/model rather than assuming these transfer.

Appendix: files in this folder

  • relaxed-top-n.patch — the exact diff against llama.cpp commit d222767c7, apply with git apply relaxed-top-n.patch from a clean checkout of that commit.
  • README.md — this guide.