mtp-relaxed-decoding/ holds the patch, build, and guide for a custom llama-server with relaxed-acceptance MTP speculative decoding (see its README for the full writeup). Bumps the open-webui submodule to the commit that adds it as a new service, connection, and the ollama-auth proxy swap. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
439 lines
19 KiB
Markdown
439 lines
19 KiB
Markdown
# Relaxed acceptance for MTP speculative decoding — full guide
|
|
|
|
This documents exactly how we built a custom `llama.cpp` with a "relaxed" speculative
|
|
decoding acceptance criterion, and used it to get a real throughput win from Gemma 4
|
|
E2B's MTP (multi-token-prediction) drafter — something the stock config never achieved.
|
|
|
|
**End result**: `n_max=2` + our new `--spec-draft-relaxed-top-n 5` (roughly 2-20 all work
|
|
similarly) gave **~260-275 tok/s**, beating both Ollama's stock MTP defaults (~205-220
|
|
tok/s) and even the plain non-speculative baseline (~250-260 tok/s). This is the only
|
|
configuration in the whole investigation where MTP paid for itself.
|
|
|
|
Everything below assumes Ubuntu 24.04 (this was done on WSL2) with an NVIDIA GPU. Adjust
|
|
package/arch specifics for other distros.
|
|
|
|
---
|
|
|
|
## 0. Why this exists (skip if you just want the steps)
|
|
|
|
Ollama's own MTP support (`DRAFT` directive in a Modelfile) works, but:
|
|
|
|
- It hardcodes `--spec-draft-n-max 4` with no way to change it from Ollama.
|
|
- The underlying llama.cpp acceptance rule is **strict**: a drafted token is only
|
|
accepted if it exactly equals a fresh sample drawn from the target model's own
|
|
distribution at that position. One mismatch anywhere and the whole draft batch
|
|
beyond that point is thrown away.
|
|
- For our target/drafter pair (`unsloth/gemma-4-E2B-it-GGUF`, target `Q4_K_XL`, drafter
|
|
`mtp-gemma-4-E2B-it-Q8_0.gguf`), that strict rule gave only ~18-24% acceptance, so MTP
|
|
was *slower* than not using it at all.
|
|
|
|
We patched llama.cpp to add a genuinely relaxed rule — accept a drafted token if it
|
|
ranks in the target's top-N candidates by probability, not just if it's the literal
|
|
one token the target happened to sample. That's not exposed anywhere in stock llama.cpp
|
|
or Ollama; it required a source patch.
|
|
|
|
---
|
|
|
|
## 1. Prerequisites
|
|
|
|
### 1.1 Ollama with a target+drafter model already built
|
|
|
|
You need Ollama running with a target model and its MTP drafter combined via a
|
|
Modelfile's `DRAFT` directive. Two ways to get the GGUF files, both from the same
|
|
public (non-gated) Hugging Face repo:
|
|
|
|
```bash
|
|
mkdir -p ~/mtp-build && cd ~/mtp-build
|
|
curl -L -o target.gguf \
|
|
"https://huggingface.co/unsloth/gemma-4-E2B-it-GGUF/resolve/main/gemma-4-E2B-it-UD-Q4_K_XL.gguf"
|
|
curl -L -o draft.gguf \
|
|
"https://huggingface.co/unsloth/gemma-4-E2B-it-GGUF/resolve/main/MTP/mtp-gemma-4-E2B-it-Q8_0.gguf"
|
|
|
|
cat > Modelfile <<'EOF'
|
|
FROM ./target.gguf
|
|
DRAFT ./draft.gguf
|
|
EOF
|
|
|
|
ollama create gemma4:e2b-mtp -f Modelfile
|
|
```
|
|
|
|
No `--experimental` flag and no gated Hugging Face access needed — that's only required
|
|
if you build from raw safetensors instead of a pre-quantized GGUF. Swap in whatever
|
|
target/drafter GGUF pair you actually want to test; the pairing mechanism is the same.
|
|
|
|
Sanity check it actually engages speculative decoding:
|
|
|
|
```bash
|
|
journalctl -u ollama --no-pager -o cat | grep -i "draft-mtp\|draft acceptance"
|
|
```
|
|
|
|
You should see lines like:
|
|
```
|
|
spec common_specu: adding speculative implementation 'draft-mtp'
|
|
slot print_timing: id 0 | task N | draft acceptance = 0.XXXXX (...)
|
|
```
|
|
|
|
### 1.2 Find the exact llama.cpp commit your Ollama vendors
|
|
|
|
This matters — you want to patch the *exact* source your Ollama's `llama-server` was
|
|
built from, not upstream `main` (APIs drift).
|
|
|
|
```bash
|
|
/usr/local/lib/ollama/llama-server --version
|
|
```
|
|
|
|
Output looks like:
|
|
```
|
|
version: 0.3.0-dev (build 1, commit d222767c7)
|
|
```
|
|
|
|
That commit hash is a real commit in `ggml-org/llama.cpp` — Ollama vendors upstream
|
|
llama.cpp with some patches on top, and this hash identifies the base. Confirm it
|
|
resolves:
|
|
|
|
```bash
|
|
curl -s "https://api.github.com/search/commits?q=hash:<commit>+repo:ggml-org/llama.cpp" \
|
|
-H "Accept: application/vnd.github.cloak-preview" | python3 -m json.tool
|
|
```
|
|
|
|
For this run it was `d222767c7a6516559a3f49e7721b6c6b1acc87b4`
|
|
("kleidiai: Rework KleidiAI Build System/Integration (#26077)").
|
|
|
|
### 1.3 Install the CUDA build toolchain
|
|
|
|
Ubuntu's own `nvidia-cuda-toolkit` package (CUDA 12.0) is too old for recent GPU
|
|
architectures (e.g. Blackwell/RTX 50-series needs CUDA 12.8+). Use NVIDIA's own repo:
|
|
|
|
```bash
|
|
cd /tmp
|
|
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2404/x86_64/cuda-keyring_1.1-1_all.deb
|
|
sudo dpkg -i cuda-keyring_1.1-1_all.deb
|
|
sudo apt-get update
|
|
|
|
# Toolkit only — NOT the "cuda" or "cuda-drivers" meta-package. On WSL2 the GPU driver
|
|
# is Windows-side; installing a Linux driver package would conflict with the passthrough
|
|
# that's already working. On bare-metal Linux this distinction may not matter, but the
|
|
# toolkit-only package is still what you want (you already have a working driver).
|
|
sudo apt-get install -y build-essential cmake git cuda-toolkit-12-8
|
|
```
|
|
|
|
Verify:
|
|
```bash
|
|
export PATH=/usr/local/cuda/bin:$PATH
|
|
nvcc --version # should report release 12.8
|
|
g++ --version # 13.x is fine; nvcc 12.8 supports up to gcc 13
|
|
cmake --version # 3.24+ needed for CMAKE_CUDA_ARCHITECTURES=native
|
|
```
|
|
|
|
**Gotcha**: if `apt-get update`/`install` fails with a dpkg/apt lock error, check what's
|
|
holding it before doing anything drastic:
|
|
```bash
|
|
ps -p <pid> -o pid,ppid,etime,cmd
|
|
```
|
|
If it's `apt-daily.service` stuck "activating" for an unreasonable amount of time (check
|
|
`systemctl status apt-daily.service`), that's a known, safe-to-clear stuck state:
|
|
```bash
|
|
sudo systemctl stop apt-daily.service apt-daily.timer
|
|
```
|
|
Then retry. If it's a fresh, short-lived `unattended-upgrades` process, just wait for it
|
|
to finish on its own — don't kill it.
|
|
|
|
---
|
|
|
|
## 2. Get the source and apply the patch
|
|
|
|
```bash
|
|
mkdir -p ~/llama-relaxed && cd ~/llama-relaxed
|
|
git init -q
|
|
git remote add origin https://github.com/ggml-org/llama.cpp.git
|
|
git fetch --depth 1 origin <exact-commit-hash-from-1.2>
|
|
git checkout -q FETCH_HEAD
|
|
```
|
|
|
|
Apply the patch (saved alongside this guide as `relaxed-top-n.patch`):
|
|
```bash
|
|
git apply /path/to/relaxed-top-n.patch
|
|
```
|
|
|
|
If `git apply` fails (source drifted from what we patched), see "The patch by hand"
|
|
below and apply the five changes manually — they're small and self-contained.
|
|
|
|
### What the patch does
|
|
|
|
One new flag: `--spec-draft-relaxed-top-n N` (default `0` = disabled, original strict
|
|
behavior unchanged). When `N > 0`, a drafted token is accepted if it ranks among the
|
|
target's top-`N` candidates by probability at that position — not just if it's the one
|
|
token the target actually sampled.
|
|
|
|
Five files change:
|
|
|
|
| file | change |
|
|
|---|---|
|
|
| `common/common.h` | new `int32_t relaxed_top_n = 0;` field on `common_params_speculative_draft` |
|
|
| `common/sampling.h` | both `common_sampler_sample_and_accept_n` overloads gain a trailing `int32_t relaxed_top_n = 0` param |
|
|
| `common/sampling.cpp` | the actual logic — see below |
|
|
| `common/arg.cpp` | new CLI flag, following the exact pattern of the existing `--spec-draft-p-min` |
|
|
| `tools/server/server-context.cpp` | the one call site, passes `params_base.speculative.draft.relaxed_top_n` through |
|
|
|
|
The core change, in `common_sampler_sample_and_accept_n` (`common/sampling.cpp`):
|
|
|
|
```cpp
|
|
// Returns true if `token` ranks within the top `relaxed_top_n` candidates of `cur_p`
|
|
// by probability. O(cur_p.size) and correct regardless of `cur_p.sorted`, since it
|
|
// counts strictly-higher-probability entries rather than assuming sort order.
|
|
static bool relaxed_top_n_contains(const llama_token_data_array & cur_p, llama_token token, int32_t relaxed_top_n) {
|
|
float token_p = -1.0f;
|
|
for (size_t k = 0; k < cur_p.size; k++) {
|
|
if (cur_p.data[k].id == token) {
|
|
token_p = cur_p.data[k].p;
|
|
break;
|
|
}
|
|
}
|
|
if (token_p < 0.0f) {
|
|
return false; // token isn't even in the candidate set (e.g. filtered out by top_k/top_p/min_p)
|
|
}
|
|
|
|
int32_t rank = 0;
|
|
for (size_t k = 0; k < cur_p.size; k++) {
|
|
if (cur_p.data[k].p > token_p) {
|
|
rank++;
|
|
if (rank >= relaxed_top_n) {
|
|
return false;
|
|
}
|
|
}
|
|
}
|
|
return true;
|
|
}
|
|
```
|
|
|
|
...used inside the accept loop:
|
|
|
|
```cpp
|
|
const bool exact = (draft[i] == id);
|
|
bool accept = exact;
|
|
if (!accept && relaxed_top_n > 0) {
|
|
accept = relaxed_top_n_contains(gsmpl->cur_p, draft[i], relaxed_top_n);
|
|
if (getenv("RELAXED_VERIFY_LOG")) {
|
|
fprintf(stderr, "[RELAXED_VERIFY] pos=%zu relaxed_top_n=%d cur_p.size=%zu cur_p.sorted=%d draft='%s'(id=%d) target_sample='%s'(id=%d) -> %s\n",
|
|
i, relaxed_top_n, gsmpl->cur_p.size, (int) gsmpl->cur_p.sorted,
|
|
common_token_to_piece(ctx, draft[i]).c_str(), draft[i],
|
|
common_token_to_piece(ctx, id).c_str(), id,
|
|
accept ? "ACCEPT(relaxed)" : "reject");
|
|
}
|
|
}
|
|
|
|
// if the draft token is accepted under the relaxed rule, emit *that* token
|
|
// (not `id`) — server-context.cpp fed the whole drafted sequence into the
|
|
// target as one batch, so downstream KV state and idxs[i+1]'s logits already
|
|
// assume draft[i] was the real token at this position.
|
|
const llama_token emitted = accept ? draft[i] : id;
|
|
|
|
common_sampler_accept(gsmpl, emitted, true);
|
|
result.push_back(emitted);
|
|
|
|
if (!accept) {
|
|
break;
|
|
}
|
|
```
|
|
|
|
The `RELAXED_VERIFY_LOG` env-gated debug print is left in intentionally (harmless when
|
|
unset) — it's how we verified the patch was actually doing something (see §5).
|
|
|
|
**Correctness note worth understanding**: in batched speculative verification, the
|
|
target processes the *whole* drafted sequence as one forward pass, so position `i`'s
|
|
candidate distribution already reflects "what comes after `draft[0..i-1]`" regardless of
|
|
whether `draft[i]` itself gets accepted. That means accepting `draft[i]` under the
|
|
relaxed rule is self-consistent — as long as you push `draft[i]` (not the target's own
|
|
`id`) into the result and into `common_sampler_accept`, which is what the patch does.
|
|
|
|
`gsmpl->cur_p` is always populated before this point (via `set_logits()`, called
|
|
unconditionally at the top of `common_sampler_sample()`), including when
|
|
`--spec-draft-backend-sampling` is on — so the patch works with the fast GPU-sampling
|
|
path, not just the CPU fallback.
|
|
|
|
---
|
|
|
|
## 3. Build
|
|
|
|
```bash
|
|
cd ~/llama-relaxed
|
|
export PATH=/usr/local/cuda/bin:$PATH
|
|
cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=native -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=OFF
|
|
cmake --build build --config Release -j"$(nproc)" --target llama-server
|
|
```
|
|
|
|
`CMAKE_CUDA_ARCHITECTURES=native` auto-detects your installed GPU (needs CMake 3.24+).
|
|
Watch the configure output for a line like:
|
|
```
|
|
-- Replacing 120-real in CMAKE_CUDA_ARCHITECTURES_NATIVE with 120a-real
|
|
-- Using CMAKE_CUDA_ARCHITECTURES=120a-real ...
|
|
```
|
|
confirming it targeted your actual card (this was `120a-real` for an RTX 5090 / Blackwell).
|
|
|
|
Build artifacts land in `build/bin/`, including `llama-server` and every shared library
|
|
it needs (`libggml-cuda.so`, `libllama.so`, etc.) **co-located in the same directory**.
|
|
This matters: ggml's dynamic backend loader scans the executable's own directory for
|
|
backend `.so` files, so having everything in one place means CUDA is auto-detected with
|
|
no extra environment variables. (If you ever need to point at a *separate* CUDA backend
|
|
`.so`, e.g. reusing Ollama's own bundled one instead of building your own, the relevant
|
|
var is `GGML_BACKEND_PATH=/path/to/libggml-cuda.so` — it must be the exact `.so` file
|
|
path, not a directory. Not needed for a self-built binary.)
|
|
|
|
Verify:
|
|
```bash
|
|
build/bin/llama-server --help | grep -A2 "relaxed-top-n"
|
|
```
|
|
Should show your new flag.
|
|
|
|
---
|
|
|
|
## 4. Run it standalone
|
|
|
|
**Important**: this flag only exists in your custom build. Ollama's own `ollama serve`/
|
|
`ollama run` will never pass it through — Ollama hardcodes its own speculative-decoding
|
|
flags with no user override for this parameter. To use the patch, run `llama-server`
|
|
directly, bypassing Ollama's process entirely (Ollama can keep running in the
|
|
background; they're independent processes and can coexist on the same GPU as long as
|
|
there's VRAM headroom).
|
|
|
|
```bash
|
|
#!/bin/bash
|
|
# run_relaxed.sh
|
|
BUILT_DIR=~/llama-relaxed/build/bin
|
|
TARGET=~/mtp-build/target.gguf # or wherever your target GGUF lives
|
|
DRAFT=~/mtp-build/draft.gguf # or wherever your drafter GGUF lives
|
|
export LD_LIBRARY_PATH="$BUILT_DIR:$LD_LIBRARY_PATH"
|
|
exec "$BUILT_DIR/llama-server" \
|
|
--model "$TARGET" \
|
|
--port "$PORT" --host 127.0.0.1 --no-webui --offline -c 32768 -np 1 \
|
|
--log-verbosity 4 --no-log-prefix --no-log-timestamps --no-jinja --chat-template chatml \
|
|
--spec-type draft-mtp --spec-draft-n-max "$NMAX" --spec-draft-relaxed-top-n "$RELAXED" \
|
|
--spec-draft-backend-sampling \
|
|
--spec-draft-model "$DRAFT" -ngl 999 --spec-draft-ngl 999 \
|
|
--flash-attn auto -b 1024 -ub 1024 --context-shift --keep 4
|
|
```
|
|
|
|
```bash
|
|
chmod +x run_relaxed.sh
|
|
PORT=38700 NMAX=2 RELAXED=5 ./run_relaxed.sh
|
|
```
|
|
|
|
Wait for `GET /health` to return `{"status":"ok"}`, then hit its native `/completion`
|
|
endpoint directly (this is llama.cpp's own API, not Ollama's `/api/generate`):
|
|
|
|
```bash
|
|
curl -s http://127.0.0.1:38700/completion -d '{
|
|
"prompt": "Write a short poem about code.",
|
|
"n_predict": 128,
|
|
"temperature": 0.6,
|
|
"top_k": 64,
|
|
"top_p": 0.95,
|
|
"min_p": 0.01,
|
|
"cache_prompt": false
|
|
}'
|
|
```
|
|
|
|
**If you reuse Ollama's already-downloaded blob files instead of your own copies** (they
|
|
live under `/usr/share/ollama/.ollama/models/blobs/sha256-<hash>`, filenames = the
|
|
manifest hashes from `ollama show <model> --modelfile` or captured from a running
|
|
process's cmdline), you'll hit a permission error — those files are owned by the
|
|
`ollama` user/group. Two options: copy your own files instead (simplest), or run via
|
|
`sg ollama -c ./run_relaxed.sh` if your user is in the `ollama` group (check with
|
|
`getent group ollama`; newly-added group membership needs a fresh login shell, `sg` lets
|
|
you pick it up without one). **Don't nest `sg` calls** (`sg a -c "sg b -c ..."`) — it
|
|
tends to trigger an interactive password prompt that hangs non-interactive invocations.
|
|
|
|
---
|
|
|
|
## 5. Verify it's actually working — don't just trust the throughput numbers
|
|
|
|
Throughput/acceptance-rate changes alone don't prove the patch is doing what you think.
|
|
Verify directly:
|
|
|
|
**Control test**: run with `RELAXED=0` on your *new* build and confirm behavior matches
|
|
your original (unpatched) binary's strict-mode numbers. This rules out "the rebuild
|
|
changed something unrelated." Also confirm zero `RELAXED_VERIFY_LOG` lines appear
|
|
(proves the new code path is genuinely inert when disabled — it's gated behind
|
|
`relaxed_top_n > 0`).
|
|
|
|
**Positive test**: run with `RELAXED_VERIFY_LOG=1 RELAXED=<something>` and grep the
|
|
output:
|
|
```bash
|
|
RELAXED_VERIFY_LOG=1 PORT=38700 NMAX=2 RELAXED=5 ./run_relaxed.sh > server.log 2>&1 &
|
|
# ...make some completion requests...
|
|
grep "RELAXED_VERIFY" server.log
|
|
grep -c "ACCEPT(relaxed)" server.log
|
|
grep -c "reject" server.log
|
|
```
|
|
You should see a genuine mix of accepts and rejects (not all-or-nothing), and the
|
|
specific token pairs should look semantically sensible when you eyeball them — e.g.
|
|
`draft=' Code' target_sample=' Algorithm' -> ACCEPT` is a plausible substitution;
|
|
`draft=' Journey' target_sample=' Whisper' -> reject` is a sensible rejection.
|
|
|
|
**Stress test the extremes**: set `relaxed_top_n` absurdly high (e.g. `300000`, bigger
|
|
than the vocab) *and* disable `min_p` in the request (`"min_p": 0.0` — see the gotcha
|
|
below) to force every draft token accepted. You should see `draft acceptance = 1.00000`
|
|
in the logs and — critically — **coherent-but-degenerate** output (ours fell into
|
|
repetition loops), not crashes or garbled/invalid tokens. A clean failure mode there is
|
|
good evidence the KV-state handling (emitting `draft[i]` correctly) is sound; a crash or
|
|
corrupted text would mean the patch has a real bug.
|
|
|
|
**Read the actual completion text**, not just the stats, at every step. Coherent
|
|
prose = the relaxed criterion isn't silently corrupting output; a still-passable poem
|
|
at `relaxed_top_n=1000+` is a meaningful signal, not just noise.
|
|
|
|
### Gotcha: `min_p` silently caps your candidate pool
|
|
|
|
The sampler chain's default `min_p` (0.05 in llama-server) aggressively prunes
|
|
`gsmpl->cur_p` down to a handful of candidates (we observed sizes of 1-5) *before* the
|
|
relaxed check ever sees it — regardless of `top_k`/`top_p` on the request, and
|
|
regardless of `--spec-draft-backend-sampling` on/off. If you're testing what
|
|
`relaxed_top_n` values actually do, **explicitly set `"min_p": 0.01` (or `0.0`) in every
|
|
request** — otherwise you're comparing `relaxed_top_n` values against a candidate pool
|
|
that's already artificially tiny, which was a real methodological mistake we made and
|
|
had to redo. Add a `size`/`sorted` field to the debug log (as above) if you ever need to
|
|
double-check the candidate pool size directly instead of inferring it.
|
|
|
|
---
|
|
|
|
## 6. What we found (this specific target/drafter pair, RTX 5090)
|
|
|
|
All numbers: `gemma-4-E2B-it-UD-Q4_K_XL.gguf` target + `mtp-gemma-4-E2B-it-Q8_0.gguf`
|
|
drafter, `temperature` 0.6-1.0 (didn't matter much), `min_p=0.01`, 128-token completions,
|
|
5-run averages after a warm-up request.
|
|
|
|
| config | tok/s | acceptance |
|
|
|---|---:|---:|
|
|
| plain target, no speculative decoding | ~250-260 | — |
|
|
| `n_max=4` strict (Ollama's hardcoded default) | ~205-220 | ~18-24% |
|
|
| `n_max=2` strict | ~230-245 | ~30-36% |
|
|
| `n_max=16, p_min=0.8` (a value recommended online for a different setup) | ~124-140 | ~55-70% (but slow — bigger verify batches cost more than they save) |
|
|
| `n_max=2` + **relaxed_top_n=2-20** (this patch) | **~260-275** | ~38-53% |
|
|
| `n_max=4/8/16` + relaxed_top_n | progressively worse (~230→130) | mean accepted length stays flat ~2.3-2.7 regardless — extra draft depth is wasted compute |
|
|
|
|
Takeaways that seemed to hold up across every variant we tried (temperature 0/0.1/0.6/1.0,
|
|
`min_p` 0.0/0.01/0.05, `p_min`, target quantization Q4 vs Q8):
|
|
|
|
- **`n_max` beyond ~2 is not worth it** for this drafter — its useful predictive depth
|
|
caps out around position 2-3 regardless of acceptance rule, temperature, or
|
|
quantization. Bigger `n_max` just makes verification batches more expensive without a
|
|
proportional increase in useful accepted tokens.
|
|
- **The relaxed criterion is the actual lever that made MTP worthwhile here.** Every
|
|
strict-mode config we tried underperformed the non-speculative baseline; every
|
|
reasonable relaxed config beat it.
|
|
- **There's a plateau, not a sharp peak**, across `relaxed_top_n ≈ 2-20` — differences
|
|
within that range are close to sample noise at n=5 runs. Don't over-index on the exact
|
|
best value from a small sweep; re-run a few times before trusting a "winner."
|
|
- Quantization match between target/drafter (Q4 vs Q8 target) and temperature both
|
|
turned out to be second-order effects, not the dominant factor, contrary to our
|
|
initial (wrong) hypotheses — worth re-testing on your own hardware/model rather than
|
|
assuming these transfer.
|
|
|
|
---
|
|
|
|
## Appendix: files in this folder
|
|
|
|
- `relaxed-top-n.patch` — the exact diff against llama.cpp commit `d222767c7`, apply
|
|
with `git apply relaxed-top-n.patch` from a clean checkout of that commit.
|
|
- `README.md` — this guide.
|