# Relaxed acceptance for MTP speculative decoding — full guide This documents exactly how we built a custom `llama.cpp` with a "relaxed" speculative decoding acceptance criterion, and used it to get a real throughput win from Gemma 4 E2B's MTP (multi-token-prediction) drafter — something the stock config never achieved. **End result**: `n_max=2` + our new `--spec-draft-relaxed-top-n 5` (roughly 2-20 all work similarly) gave **~260-275 tok/s**, beating both Ollama's stock MTP defaults (~205-220 tok/s) and even the plain non-speculative baseline (~250-260 tok/s). This is the only configuration in the whole investigation where MTP paid for itself. Everything below assumes Ubuntu 24.04 (this was done on WSL2) with an NVIDIA GPU. Adjust package/arch specifics for other distros. --- ## 0. Why this exists (skip if you just want the steps) Ollama's own MTP support (`DRAFT` directive in a Modelfile) works, but: - It hardcodes `--spec-draft-n-max 4` with no way to change it from Ollama. - The underlying llama.cpp acceptance rule is **strict**: a drafted token is only accepted if it exactly equals a fresh sample drawn from the target model's own distribution at that position. One mismatch anywhere and the whole draft batch beyond that point is thrown away. - For our target/drafter pair (`unsloth/gemma-4-E2B-it-GGUF`, target `Q4_K_XL`, drafter `mtp-gemma-4-E2B-it-Q8_0.gguf`), that strict rule gave only ~18-24% acceptance, so MTP was *slower* than not using it at all. We patched llama.cpp to add a genuinely relaxed rule — accept a drafted token if it ranks in the target's top-N candidates by probability, not just if it's the literal one token the target happened to sample. That's not exposed anywhere in stock llama.cpp or Ollama; it required a source patch. --- ## 1. Prerequisites ### 1.1 Ollama with a target+drafter model already built You need Ollama running with a target model and its MTP drafter combined via a Modelfile's `DRAFT` directive. Two ways to get the GGUF files, both from the same public (non-gated) Hugging Face repo: ```bash mkdir -p ~/mtp-build && cd ~/mtp-build curl -L -o target.gguf \ "https://huggingface.co/unsloth/gemma-4-E2B-it-GGUF/resolve/main/gemma-4-E2B-it-UD-Q4_K_XL.gguf" curl -L -o draft.gguf \ "https://huggingface.co/unsloth/gemma-4-E2B-it-GGUF/resolve/main/MTP/mtp-gemma-4-E2B-it-Q8_0.gguf" cat > Modelfile <<'EOF' FROM ./target.gguf DRAFT ./draft.gguf EOF ollama create gemma4:e2b-mtp -f Modelfile ``` No `--experimental` flag and no gated Hugging Face access needed — that's only required if you build from raw safetensors instead of a pre-quantized GGUF. Swap in whatever target/drafter GGUF pair you actually want to test; the pairing mechanism is the same. Sanity check it actually engages speculative decoding: ```bash journalctl -u ollama --no-pager -o cat | grep -i "draft-mtp\|draft acceptance" ``` You should see lines like: ``` spec common_specu: adding speculative implementation 'draft-mtp' slot print_timing: id 0 | task N | draft acceptance = 0.XXXXX (...) ``` ### 1.2 Find the exact llama.cpp commit your Ollama vendors This matters — you want to patch the *exact* source your Ollama's `llama-server` was built from, not upstream `main` (APIs drift). ```bash /usr/local/lib/ollama/llama-server --version ``` Output looks like: ``` version: 0.3.0-dev (build 1, commit d222767c7) ``` That commit hash is a real commit in `ggml-org/llama.cpp` — Ollama vendors upstream llama.cpp with some patches on top, and this hash identifies the base. Confirm it resolves: ```bash curl -s "https://api.github.com/search/commits?q=hash:+repo:ggml-org/llama.cpp" \ -H "Accept: application/vnd.github.cloak-preview" | python3 -m json.tool ``` For this run it was `d222767c7a6516559a3f49e7721b6c6b1acc87b4` ("kleidiai: Rework KleidiAI Build System/Integration (#26077)"). ### 1.3 Install the CUDA build toolchain Ubuntu's own `nvidia-cuda-toolkit` package (CUDA 12.0) is too old for recent GPU architectures (e.g. Blackwell/RTX 50-series needs CUDA 12.8+). Use NVIDIA's own repo: ```bash cd /tmp wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2404/x86_64/cuda-keyring_1.1-1_all.deb sudo dpkg -i cuda-keyring_1.1-1_all.deb sudo apt-get update # Toolkit only — NOT the "cuda" or "cuda-drivers" meta-package. On WSL2 the GPU driver # is Windows-side; installing a Linux driver package would conflict with the passthrough # that's already working. On bare-metal Linux this distinction may not matter, but the # toolkit-only package is still what you want (you already have a working driver). sudo apt-get install -y build-essential cmake git cuda-toolkit-12-8 ``` Verify: ```bash export PATH=/usr/local/cuda/bin:$PATH nvcc --version # should report release 12.8 g++ --version # 13.x is fine; nvcc 12.8 supports up to gcc 13 cmake --version # 3.24+ needed for CMAKE_CUDA_ARCHITECTURES=native ``` **Gotcha**: if `apt-get update`/`install` fails with a dpkg/apt lock error, check what's holding it before doing anything drastic: ```bash ps -p -o pid,ppid,etime,cmd ``` If it's `apt-daily.service` stuck "activating" for an unreasonable amount of time (check `systemctl status apt-daily.service`), that's a known, safe-to-clear stuck state: ```bash sudo systemctl stop apt-daily.service apt-daily.timer ``` Then retry. If it's a fresh, short-lived `unattended-upgrades` process, just wait for it to finish on its own — don't kill it. --- ## 2. Get the source and apply the patch ```bash mkdir -p ~/llama-relaxed && cd ~/llama-relaxed git init -q git remote add origin https://github.com/ggml-org/llama.cpp.git git fetch --depth 1 origin git checkout -q FETCH_HEAD ``` Apply the patch (saved alongside this guide as `relaxed-top-n.patch`): ```bash git apply /path/to/relaxed-top-n.patch ``` If `git apply` fails (source drifted from what we patched), see "The patch by hand" below and apply the five changes manually — they're small and self-contained. ### What the patch does One new flag: `--spec-draft-relaxed-top-n N` (default `0` = disabled, original strict behavior unchanged). When `N > 0`, a drafted token is accepted if it ranks among the target's top-`N` candidates by probability at that position — not just if it's the one token the target actually sampled. Five files change: | file | change | |---|---| | `common/common.h` | new `int32_t relaxed_top_n = 0;` field on `common_params_speculative_draft` | | `common/sampling.h` | both `common_sampler_sample_and_accept_n` overloads gain a trailing `int32_t relaxed_top_n = 0` param | | `common/sampling.cpp` | the actual logic — see below | | `common/arg.cpp` | new CLI flag, following the exact pattern of the existing `--spec-draft-p-min` | | `tools/server/server-context.cpp` | the one call site, passes `params_base.speculative.draft.relaxed_top_n` through | The core change, in `common_sampler_sample_and_accept_n` (`common/sampling.cpp`): ```cpp // Returns true if `token` ranks within the top `relaxed_top_n` candidates of `cur_p` // by probability. O(cur_p.size) and correct regardless of `cur_p.sorted`, since it // counts strictly-higher-probability entries rather than assuming sort order. static bool relaxed_top_n_contains(const llama_token_data_array & cur_p, llama_token token, int32_t relaxed_top_n) { float token_p = -1.0f; for (size_t k = 0; k < cur_p.size; k++) { if (cur_p.data[k].id == token) { token_p = cur_p.data[k].p; break; } } if (token_p < 0.0f) { return false; // token isn't even in the candidate set (e.g. filtered out by top_k/top_p/min_p) } int32_t rank = 0; for (size_t k = 0; k < cur_p.size; k++) { if (cur_p.data[k].p > token_p) { rank++; if (rank >= relaxed_top_n) { return false; } } } return true; } ``` ...used inside the accept loop: ```cpp const bool exact = (draft[i] == id); bool accept = exact; if (!accept && relaxed_top_n > 0) { accept = relaxed_top_n_contains(gsmpl->cur_p, draft[i], relaxed_top_n); if (getenv("RELAXED_VERIFY_LOG")) { fprintf(stderr, "[RELAXED_VERIFY] pos=%zu relaxed_top_n=%d cur_p.size=%zu cur_p.sorted=%d draft='%s'(id=%d) target_sample='%s'(id=%d) -> %s\n", i, relaxed_top_n, gsmpl->cur_p.size, (int) gsmpl->cur_p.sorted, common_token_to_piece(ctx, draft[i]).c_str(), draft[i], common_token_to_piece(ctx, id).c_str(), id, accept ? "ACCEPT(relaxed)" : "reject"); } } // if the draft token is accepted under the relaxed rule, emit *that* token // (not `id`) — server-context.cpp fed the whole drafted sequence into the // target as one batch, so downstream KV state and idxs[i+1]'s logits already // assume draft[i] was the real token at this position. const llama_token emitted = accept ? draft[i] : id; common_sampler_accept(gsmpl, emitted, true); result.push_back(emitted); if (!accept) { break; } ``` The `RELAXED_VERIFY_LOG` env-gated debug print is left in intentionally (harmless when unset) — it's how we verified the patch was actually doing something (see §5). **Correctness note worth understanding**: in batched speculative verification, the target processes the *whole* drafted sequence as one forward pass, so position `i`'s candidate distribution already reflects "what comes after `draft[0..i-1]`" regardless of whether `draft[i]` itself gets accepted. That means accepting `draft[i]` under the relaxed rule is self-consistent — as long as you push `draft[i]` (not the target's own `id`) into the result and into `common_sampler_accept`, which is what the patch does. `gsmpl->cur_p` is always populated before this point (via `set_logits()`, called unconditionally at the top of `common_sampler_sample()`), including when `--spec-draft-backend-sampling` is on — so the patch works with the fast GPU-sampling path, not just the CPU fallback. --- ## 3. Build ```bash cd ~/llama-relaxed export PATH=/usr/local/cuda/bin:$PATH cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=native -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=OFF cmake --build build --config Release -j"$(nproc)" --target llama-server ``` `CMAKE_CUDA_ARCHITECTURES=native` auto-detects your installed GPU (needs CMake 3.24+). Watch the configure output for a line like: ``` -- Replacing 120-real in CMAKE_CUDA_ARCHITECTURES_NATIVE with 120a-real -- Using CMAKE_CUDA_ARCHITECTURES=120a-real ... ``` confirming it targeted your actual card (this was `120a-real` for an RTX 5090 / Blackwell). Build artifacts land in `build/bin/`, including `llama-server` and every shared library it needs (`libggml-cuda.so`, `libllama.so`, etc.) **co-located in the same directory**. This matters: ggml's dynamic backend loader scans the executable's own directory for backend `.so` files, so having everything in one place means CUDA is auto-detected with no extra environment variables. (If you ever need to point at a *separate* CUDA backend `.so`, e.g. reusing Ollama's own bundled one instead of building your own, the relevant var is `GGML_BACKEND_PATH=/path/to/libggml-cuda.so` — it must be the exact `.so` file path, not a directory. Not needed for a self-built binary.) Verify: ```bash build/bin/llama-server --help | grep -A2 "relaxed-top-n" ``` Should show your new flag. --- ## 4. Run it standalone **Important**: this flag only exists in your custom build. Ollama's own `ollama serve`/ `ollama run` will never pass it through — Ollama hardcodes its own speculative-decoding flags with no user override for this parameter. To use the patch, run `llama-server` directly, bypassing Ollama's process entirely (Ollama can keep running in the background; they're independent processes and can coexist on the same GPU as long as there's VRAM headroom). ```bash #!/bin/bash # run_relaxed.sh BUILT_DIR=~/llama-relaxed/build/bin TARGET=~/mtp-build/target.gguf # or wherever your target GGUF lives DRAFT=~/mtp-build/draft.gguf # or wherever your drafter GGUF lives export LD_LIBRARY_PATH="$BUILT_DIR:$LD_LIBRARY_PATH" exec "$BUILT_DIR/llama-server" \ --model "$TARGET" \ --port "$PORT" --host 127.0.0.1 --no-webui --offline -c 32768 -np 1 \ --log-verbosity 4 --no-log-prefix --no-log-timestamps --no-jinja --chat-template chatml \ --spec-type draft-mtp --spec-draft-n-max "$NMAX" --spec-draft-relaxed-top-n "$RELAXED" \ --spec-draft-backend-sampling \ --spec-draft-model "$DRAFT" -ngl 999 --spec-draft-ngl 999 \ --flash-attn auto -b 1024 -ub 1024 --context-shift --keep 4 ``` ```bash chmod +x run_relaxed.sh PORT=38700 NMAX=2 RELAXED=5 ./run_relaxed.sh ``` Wait for `GET /health` to return `{"status":"ok"}`, then hit its native `/completion` endpoint directly (this is llama.cpp's own API, not Ollama's `/api/generate`): ```bash curl -s http://127.0.0.1:38700/completion -d '{ "prompt": "Write a short poem about code.", "n_predict": 128, "temperature": 0.6, "top_k": 64, "top_p": 0.95, "min_p": 0.01, "cache_prompt": false }' ``` **If you reuse Ollama's already-downloaded blob files instead of your own copies** (they live under `/usr/share/ollama/.ollama/models/blobs/sha256-`, filenames = the manifest hashes from `ollama show --modelfile` or captured from a running process's cmdline), you'll hit a permission error — those files are owned by the `ollama` user/group. Two options: copy your own files instead (simplest), or run via `sg ollama -c ./run_relaxed.sh` if your user is in the `ollama` group (check with `getent group ollama`; newly-added group membership needs a fresh login shell, `sg` lets you pick it up without one). **Don't nest `sg` calls** (`sg a -c "sg b -c ..."`) — it tends to trigger an interactive password prompt that hangs non-interactive invocations. --- ## 5. Verify it's actually working — don't just trust the throughput numbers Throughput/acceptance-rate changes alone don't prove the patch is doing what you think. Verify directly: **Control test**: run with `RELAXED=0` on your *new* build and confirm behavior matches your original (unpatched) binary's strict-mode numbers. This rules out "the rebuild changed something unrelated." Also confirm zero `RELAXED_VERIFY_LOG` lines appear (proves the new code path is genuinely inert when disabled — it's gated behind `relaxed_top_n > 0`). **Positive test**: run with `RELAXED_VERIFY_LOG=1 RELAXED=` and grep the output: ```bash RELAXED_VERIFY_LOG=1 PORT=38700 NMAX=2 RELAXED=5 ./run_relaxed.sh > server.log 2>&1 & # ...make some completion requests... grep "RELAXED_VERIFY" server.log grep -c "ACCEPT(relaxed)" server.log grep -c "reject" server.log ``` You should see a genuine mix of accepts and rejects (not all-or-nothing), and the specific token pairs should look semantically sensible when you eyeball them — e.g. `draft=' Code' target_sample=' Algorithm' -> ACCEPT` is a plausible substitution; `draft=' Journey' target_sample=' Whisper' -> reject` is a sensible rejection. **Stress test the extremes**: set `relaxed_top_n` absurdly high (e.g. `300000`, bigger than the vocab) *and* disable `min_p` in the request (`"min_p": 0.0` — see the gotcha below) to force every draft token accepted. You should see `draft acceptance = 1.00000` in the logs and — critically — **coherent-but-degenerate** output (ours fell into repetition loops), not crashes or garbled/invalid tokens. A clean failure mode there is good evidence the KV-state handling (emitting `draft[i]` correctly) is sound; a crash or corrupted text would mean the patch has a real bug. **Read the actual completion text**, not just the stats, at every step. Coherent prose = the relaxed criterion isn't silently corrupting output; a still-passable poem at `relaxed_top_n=1000+` is a meaningful signal, not just noise. ### Gotcha: `min_p` silently caps your candidate pool The sampler chain's default `min_p` (0.05 in llama-server) aggressively prunes `gsmpl->cur_p` down to a handful of candidates (we observed sizes of 1-5) *before* the relaxed check ever sees it — regardless of `top_k`/`top_p` on the request, and regardless of `--spec-draft-backend-sampling` on/off. If you're testing what `relaxed_top_n` values actually do, **explicitly set `"min_p": 0.01` (or `0.0`) in every request** — otherwise you're comparing `relaxed_top_n` values against a candidate pool that's already artificially tiny, which was a real methodological mistake we made and had to redo. Add a `size`/`sorted` field to the debug log (as above) if you ever need to double-check the candidate pool size directly instead of inferring it. --- ## 6. What we found (this specific target/drafter pair, RTX 5090) All numbers: `gemma-4-E2B-it-UD-Q4_K_XL.gguf` target + `mtp-gemma-4-E2B-it-Q8_0.gguf` drafter, `temperature` 0.6-1.0 (didn't matter much), `min_p=0.01`, 128-token completions, 5-run averages after a warm-up request. | config | tok/s | acceptance | |---|---:|---:| | plain target, no speculative decoding | ~250-260 | — | | `n_max=4` strict (Ollama's hardcoded default) | ~205-220 | ~18-24% | | `n_max=2` strict | ~230-245 | ~30-36% | | `n_max=16, p_min=0.8` (a value recommended online for a different setup) | ~124-140 | ~55-70% (but slow — bigger verify batches cost more than they save) | | `n_max=2` + **relaxed_top_n=2-20** (this patch) | **~260-275** | ~38-53% | | `n_max=4/8/16` + relaxed_top_n | progressively worse (~230→130) | mean accepted length stays flat ~2.3-2.7 regardless — extra draft depth is wasted compute | Takeaways that seemed to hold up across every variant we tried (temperature 0/0.1/0.6/1.0, `min_p` 0.0/0.01/0.05, `p_min`, target quantization Q4 vs Q8): - **`n_max` beyond ~2 is not worth it** for this drafter — its useful predictive depth caps out around position 2-3 regardless of acceptance rule, temperature, or quantization. Bigger `n_max` just makes verification batches more expensive without a proportional increase in useful accepted tokens. - **The relaxed criterion is the actual lever that made MTP worthwhile here.** Every strict-mode config we tried underperformed the non-speculative baseline; every reasonable relaxed config beat it. - **There's a plateau, not a sharp peak**, across `relaxed_top_n ≈ 2-20` — differences within that range are close to sample noise at n=5 runs. Don't over-index on the exact best value from a small sweep; re-run a few times before trusting a "winner." - Quantization match between target/drafter (Q4 vs Q8 target) and temperature both turned out to be second-order effects, not the dominant factor, contrary to our initial (wrong) hypotheses — worth re-testing on your own hardware/model rather than assuming these transfer. --- ## Appendix: files in this folder - `relaxed-top-n.patch` — the exact diff against llama.cpp commit `d222767c7`, apply with `git apply relaxed-top-n.patch` from a clean checkout of that commit. - `README.md` — this guide.