Getting Qwen3.8-Flash-Next on my 2× CMP 170HX

in «Projects» by Plean

Hardware: 2× NVIDIA CMP 170HX (GA100, sm80, 64 GB HBM2e each via the community cmpunlocker (Go donate to amoghmunikote he deserves every cent) unlock), PCIe Gen2 x4, 64 GB system RAM, Unraid host, llama-swap router, one NVME drive.

NOTE: All statements were made by October 7th 2026, vLLM and model quants change quickly and this may be out of date quickly

What do I want?

From prior testing I was able to find out that for the dual CMP 170HX setup running at PCI Gen 2x4 perform best running with vLLM over llama.cpp and running with pipeline parallelism (PP) over tensor parallelism (TP) (I'll make a little page about this as well to catalogue how this was tested and the findings). Unfortunately, off the shelf vLLM does not support Qwen3.8-Flash-Next running with pipeline parallelism and MTP. To add on top of all of this with limited system memory running the full BF16 PLE table (96GB) offloaded to RAM becomes impossible requiring PLE offload to NVME (Some testing was done with INT4 PLE table replacements to minimze RAM usage but this was not fruitful and will not be talked about here) and as you guessed it vLLM does not support the NVME PLE table offload as well.

This leaves me with the following wants: 1. vLLM - for performance 2. PLE NVME Offload - for RAM limitations 3. Pipeline parallelism - Tested as being more performant than TP on my setup 4. MTP speculative decoding - To go fast

How did I get there?

The simple answer, a lot of random huggingface repos and models and a lot of helpful writeups from people smarter than me. Then working with LLMs to parse all the different information I was able to find that I belived to be helpful to me building out a solution that worked for my setup.

Here is what I found that ended up working the best for me. The huggingface repository and quantization I used is here. Made by the Minachist, this repository is designed for some interesting setups. His focus is primarily multiple consumer grade GPUs using vLLM with pipeline parallelism with MTP to speed up generation. It is a mixed precision model and I ended up going with the 4x3090 setup as it's precision was a bit higher due to having more INT8 weights than the 3x3090 more quantized version of the model. Please go an read his articles as well as they deserve some love with how comprehensive they are Article #1 Article #2.

What needed to be done differently than the Minachist setup

Minachist's setup his system for consumer cards on normal Gen4 x16 slots with a lot of system RAM. My cards are limited to Gen2 x4 link which is about about a sixteenth of the communication bandwidth his GPUs have plus less system RAM. From my understanding his offload architecture exists to solve challenges with VRAM and on my hardware (larger amount of VRAM) several of his tricks could potentially slow me down.

  1. His setups park the QSA KV cache and the routed experts in host RAM because a ~65 GB model has nowhere else to go on 24 GB cards. I have 128 GB of VRAM across the two cards, so the model body plus a full precision KV cache fits with room to spare. Copying his tricks here would route per-token traffic across the slowest PCIe link in the machine and eat system RAM I cannot spare.
  2. His recipe repacks the table because every byte of RAM/VRAM is contested on a 24 GB rig. I don't need the space saved (my INT4 PLE table experiments were not fruitful). For me placing the BF16 table file on NVMe (mmap + page prefetch) keeps it completely off the VRAM/RAM which is what I need.
  3. His docs state the draft head predicts a single token and that wider drafts aren't supported. I measured k=3 stable with ~4.0 tokens accepted per step on so I made the decision to just ignore that note.
  4. His reference setups run the KV cache itself at fp8 (part of how the QSA patch squeezes long contexts onto small cards). With 64 GB per card I do not quantize my KV cache at all.

Here are the final results (warm, measured through llama-swap):

workload tok/s
single-stream, MTP-3, predictable content (code, prose, low-entropy; their Avg generation throughput windows) 197–201
single-stream, MTP-3, thinking-heavy (technical/reasoning) 116–122
prefill ~5.8K t/s
240K-token needle HIT
MTP-3 k=1 doubled output at parity; k=3 adds up to +40% on predictable content (mean acceptance length 4.0; acceptance 0.45–1.0, content-dependent)

For those who are interested in recreating this themselves, below are what I would call the LLM instructions which you can pass to your agent to build this out for your own setup. Good luck!

LLM Instructions

The full recipe, written for an LLM agent to execute on a similar box. Paths are Unraid defaults (/mnt/user/appdata, /mnt/cache/appdata) — substitute your own mounts. No KV quantization anywhere: the cache runs full BF16 (the wrapper passes --kv-cache-dtype bfloat16 explicitly).

Hardware: 2× NVIDIA CMP 170HX (GA100, sm80, 64 GB HBM2e each via the community unlock), PCIe Gen2 x4 interconnects (weaker than the reference rigs' Gen4 x16), 64 GB system RAM, Unraid host, llama-swap router, one fast NVMe (the PLE table must live on NVMe).

How the ~200 t/s is generated: MTP-3 drafts 3 tokens/step at ~99% per-position acceptance on predictable content → 4.0 tokens/step ≈ 201 t/s at the engine's ~50 steps/s.

The architecture decisions that make this work on THIS hardware:

  • PP=2 / TP=1. Tensor parallelism all-reduces every token over the Gen2x4 link and dies there; pipeline passes one activation hop per step. This matches the machine's best topology and the minamism patch set's design.
  • Everything in VRAM except the PLE. The 65–67 GiB body + a full-precision BF16 KV cache fit 128 GiB with room — the final recipe does NOT run the KV in fp8; we deliberately do NOT use the reference rigs' QSA-KV-to-host-RAM or experts-in-host-RAM tricks (they exist for 24 GB cards and would be a PCIe-per- token disaster on Gen2x4 and a RAM disaster on 64 GB).
  • PLE n-gram table on NVMe (mmap + page prefetch). No INT4 repack, no plugin — the table is BF16, built once, read randomly at ~39 KiB/token, LRU page-cached.
  • MTP-3 native (see "MTP" below) — the drafter is what doubles output.

1. Ingredients

piece what where from
base image vllm/vllm-openai@sha256:43f13b4c624ab9e9e6753d0eeb5953268bff334f2a826239e6f4a2197d47bb96 (0.29.1rc1.dev47) docker.io (pinned by the patches)
patches Minachist vllm-patch/ — the 2x24GB set + the 3x24GB+MTP MTP trio in the model repo
model Minachist/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound — main (gs128/INT6) or 4x3090 (gs32/INT8, higher precision); we run 4x3090 huggingface.co (gated: Qwen Community License — needs an HF token with the repo accepted)
MTP overlay make_mtp_int4.py from the MTP build dir in the model repo
serving llama-swap (any recent version) + docker —

2. The model math

  • Body is INT4/INT6/INT8 mixed, compressed-tensors, symmetric group quant: routed experts 4-bit (gs32 on the 4x3090 branch), linear_attn/QSA/shared 6-bit (INT8 on the branch), hyper-connection 8-bit, lm_head/embed/PLE-proj 8-bit, router/vision/PLE table/MTP BF16. Body ≈ 65–67 GiB.
  • 4x3090 branch vs main: same bits on experts but finer scale groups (gs32 vs gs128) and INT8 instead of INT6 on dense projections — better precision, we measured zero decode speed change (116–120 vs 122 single-stream).
  • PLE table = the model-extra-ngram-* shards (BF16 [320001536, 160], ≈ 95.4 GiB). It is flattened to a raw table file once and reused across boots; the table file is shared by both branches (identical BF16 PLE). Once the table file exists the shards are pure duplication — delete them (§6).

3. Downloads

# base image (22 GB; needs docker free space — resize docker.img first if tight)
docker pull vllm/vllm-openai@sha256:43f13b4c624ab9e9e6753d0eeb5953268bff334f2a826239e6f4a2197d47bb96

# model (full manifest, incl. the 95 GiB PLE shards — ~180 GiB; the shards are
# dropped once the first boot has built the table file — see §6)
hf download Minachist/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound \
  --revision 4x3090 \
  --local-dir /mnt/user/appdata/models/Minachist/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound-4x3090

We did the HF pull inside a scratch Ubuntu container whose /models mount landed on NVMe; any tool that can write the dir works.

4. Build the image

Assemble a build dir with: the 16 patch files (12 from vllm-patch/2x24GB/: flash-next-vllm.patch, decode-01..07, ttft-01, mem-01, vision-01, upstream-01; plus mtp-01..03 and make_mtp_int4.py from vllm-patch/3x24GB+MTP/) + this Dockerfile pattern:

ARG BASE=docker.io/vllm/vllm-openai@sha256:43f13b4c...47bb96
FROM ${BASE}
# COPY all 16 patch files to /tmp; COPY make_mtp_int4.py /opt/flash-next/
RUN cd /usr/local/lib/python3.12/dist-packages \
 && for p in /tmp/flash-next-*.patch; do patch -p1 --batch --forward < "$p"; done \
 && rm /tmp/flash-next-*.patch \
 && python3 -m py_compile vllm/models/qwen4_exp/nvidia/model.py \
      vllm/models/qwen4_exp/nvidia/hyperconnection.py \
      vllm/models/qwen4_exp/nvidia/ngram_embedding.py \
      vllm/models/qwen4_exp/nvidia/model_state.py \
      vllm/models/qwen4_exp/nvidia/qsa.py vllm/models/qwen4_exp/nvidia/ops/qsa.py \
      vllm/models/qwen4_exp/nvidia/mtp.py vllm/model_executor/models/config.py \
      vllm/v1/core/kv_cache_utils.py vllm/platforms/interface.py \
      vllm/v1/worker/gpu_worker.py vllm/v1/worker/gpu/spec_decode/utils.py \
      vllm/v1/worker/gpu/model_runner.py vllm/v1/worker/gpu/model_states/mamba_hybrid.py \
      vllm/v1/worker/mamba_utils.py

(We keep the exact file-by-file patch order from Minachist's own Dockerfile — see vllm-patch/2x24GB/Dockerfile — the order matters. The build is ~2 min: pure patch-apply + py_compile, no wheel rebuild.)

docker build -t qwen38-flash-next:2x64 -f Dockerfile .

5. MTP: one-time overlay

The MTP module is BF16; we convert its experts to INT4 once per checkpoint:

mkdir -p /mnt/cache/appdata/models/vllm/flashnext-mtp-int4
docker run --rm --gpus all --entrypoint python3 \
  -v "$MODEL:/model:ro" -v "$MTP_OUT:/out" \
  qwen38-flash-next:2x64 /opt/flash-next/make_mtp_int4.py /model /out
# then merge INTO the model dir (protect originals first):
cp -n "$MODEL/config.json"                 "$MTP_OUT/orig-config.json"
cp -n "$MODEL/model.safetensors.index.json" "$MTP_OUT/orig-index.json"
cp -f  "$MTP_OUT"/{config.json,model.safetensors.index.json,model_extra_tensors.safetensors} "$MODEL/"

Why in-place: docker runc cannot bind a file over a file inside another read-only bind (create mountpoint ... not a directory); podman can, docker can't. Merging the overlay into the model dir avoids the whole problem. Re-run this after any checkpoint change.

6. PLE table (first boot builds it)

The model dir's PLE shards are flattened to a raw table once, at first boot. Pre-create the bind target at its exact size, sparse — the patch treats an existing-but-wrong-size file as an error and a sparse exact-size file as "rebuild me":

truncate -s 102400491520 /mnt/cache/appdata/models/vllm/ngram_table.bin

First boot then writes ~92 GiB into it ("PLE mmap table ... sparse/incomplete; refilling from checkpoint"); later boots map it COW. Give the first boot ~10 min longer than normal. The refill source is the model dir's PLE shards — once the shard drop below has been done, a rebuild first needs the scoped re-fetch described there; a reboot alone no longer rebuilds the table.

Dropping the shards once the table file exists (−96 GB)

The 33 shards exist only to build the table, and the index makes deleting them a two-step job. The patched weight_loader() returns early when the mmap is prefilled, so a valid table file means the PLE tensors are never copied from the safetensors; filter_duplicate_safetensors_files() hard-fails cold start on ANY index-referenced file missing from the dir — so the index gets stripped first, then the files go. Preconditions verified on the live engine before deleting: the gate logged mode=c, prefilled=True; after weight load no process held a handle on any model-extra-ngram-*; and --model /model with HF_HUB_OFFLINE=1 means nothing re-validates against the Hub. Done on a serving engine — only the next cold boot is affected.

cd "$MODEL"
cp -p model.safetensors.index.json model.safetensors.index.json.bak
python3 - <<'EOF'
import json, os
p = 'model.safetensors.index.json'
idx = json.load(open(p))
idx['weight_map'] = {k: v for k, v in idx['weight_map'].items()
                     if 'ngram_embedding.shard_' not in k}   # drops 128 entries
json.dump(idx, open(p + '.tmp', 'w'), indent=2)
os.replace(p + '.tmp', p)
EOF
rm model-extra-ngram-*.safetensors   # 33 files, ~96 GB

The 128 dropped keys are exactly the shards' contents (…layers.1.ple.ple_embedding.ngram_embedding.shard_*.weight). The 13 small ple.* tensors (conv1d, key/value_proj, norm_*, ngram_heads offsets, layer_multipliers) live in model-00002-of-00017 and stay.

metadata.total_size/total_parameters stay stale after the strip — vLLM never reads them; leave them be.

Expected at the next cold boot — that boot is the actual validation of this change and has NOT run yet (the delete above leaves a serving engine untouched): Loading safetensors checkpoint shards counts 18, not 51, and the PLE gate line is unchanged (mode=c, prefilled=True).

If you instead see refilling from checkpoint, the table file is broken and the local rebuild source is needed — a scoped re-fetch, not the full §3 manifest:

cp -p model.safetensors.index.json.bak model.safetensors.index.json
hf download Minachist/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound \
  --revision 6a995be5542081bb65d9b7e8f580f0b8220ce0e7 \
  --include 'model-extra-ngram-*' --local-dir "$MODEL"
# 6a995be… is the snapshot commit recorded at pull time — the key of the cached
# snapshot manifest .cache/huggingface/trees/6a995be…json, which lists all 33
# shards with size + per-file lfs_sha256 (keep that file!). If the commit does
# not resolve, fall back to the mutable branch `4x3090` and verify against the
# manifest hashes. Check the re-pulled shards against lfs_sha256, then:
rm /mnt/cache/appdata/models/vllm/ngram_table.bin  # next boot auto-rebuilds (~10 min extra)

7. llama-swap wiring

Two load-bearing facts about llama-swap's cmd:: 1. It argv-splits and strips every quote character (both ' and "), so any JSON argument (--speculative-config '{...}', --engram-config, --compilation-config) cannot live inline — the JSON arrives mangled ({method:mtp,...} → vLLM parse error). Put the docker run in a tiny wrapper script that bash-quotes properly. 2. The healthCheckTimeout that matters is the top-level one (default 900 s = 15 min — the per-model key is ignored in recent versions). First boots (PLE rebuild + shfs weight load) exceed 15 min; we run 2000.

Keep the wrapper in the config folder so it's versioned with the config (llama-swap mounts it at /config):

/mnt/cache/appdata/llama-swap/config/flashnext-2x64.sh:

#!/usr/bin/env bash
set -euo pipefail
PORT="${1:-8000}"
MODEL=/mnt/cache/appdata/models/Minachist/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound-4x3090
PLE=/mnt/cache/appdata/models/vllm/ngram_table.bin
NAME=flashnext-2x64
docker rm -f "$NAME" 2>/dev/null || true
exec docker run --runtime nvidia --rm --name "$NAME" \
  --gpus 2 \
  --ipc=host --shm-size=96g --cap-add=SYS_PTRACE \
  --ulimit memlock=-1 \
  -p "${PORT}:8000" \
  -v "$MODEL:/model:ro" \
  -v /mnt/cache/appdata/models/vllm/kernel-cache/vllm:/root/.cache/vllm \
  -v /mnt/cache/appdata/models/vllm/kernel-cache/triton:/root/.triton \
  -e HF_HUB_OFFLINE=1 -e VLLM_CACHE_ROOT=/root/.cache/vllm -e TRITON_CACHE_DIR=/root/.triton \
  -e CUDA_DEVICE_ORDER=PCI_BUS_ID -e VLLM_PP_LAYER_PARTITION=24,24 \
  -e VLLM_PLE_MMAP_PATH=/vllm-ngram/ngram_table.bin -v "$PLE:/vllm-ngram/ngram_table.bin" \
  -e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
  -e VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0 -e VLLM_USE_BREAKABLE_CUDAGRAPH=0 \
  qwen38-flash-next:2x64 \
  --model /model --served-model-name Qwen3.8-Flash-Next \
  --pipeline-parallel-size 2 --tensor-parallel-size 1 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
  --per-request-spec-decode-metrics detailed --engram-config '{"cpu_offload":true}' \
  --max-model-len 262144 --max-num-seqs 6 --max-num-batched-tokens 1568 \
  --gpu-memory-utilization 0.96 \
  --kv-cache-dtype bfloat16 \
  --enable-prefix-caching --mamba-cache-mode align \
  --compilation-config '{"mode":0,"cudagraph_mode":"FULL_DECODE_ONLY","cudagraph_capture_sizes":[1,2,4,6,8]}' \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --trust-remote-code --host 0.0.0.0 --port 8000

Notes: --gpus 2 (numeric) is the only form that survives the splitter; kernel caches on NVMe make every boot after the first ~5 min.

config.yaml model entry:

healthCheckTimeout: 2000        # top-level; startup gate for slow first boots
models:
  Qwen3.8-Flash-Next:
    cmdStop: docker stop flashnext-2x64
    proxy: "http://${vllm_proxy_host}:${PORT}"
    cmd: /config/flashnext-2x64.sh ${PORT}
    ttl: 0
    metadata:
      context_length: 262144

Load it by sending any request with "model": "Qwen3.8-Flash-Next" through llama-swap.

8. MTP: why k=3 and when not

minamism documents "the draft head predicts 1 token; wider drafts are not supported", but the engine runs k=3 fine (their patch set is quiet about it). Measured on this rig:

  • k=1: mean acceptance length 2.0 (the k=1 drafter ~doubles output), single 122.
  • k=3: 3.96–4.0 tokens/step on predictable content (197–201 t/s), 2.3–2.9 on thinking-heavy content, single-stream parity at worst (117–122).
  • Tradeoff: with k=3 the padded decode batch is seqs × (1+k) = 4 × seqs, so cudagraph_capture_sizes should cover 4-multiples if you raise concurrency, and 2-stream aggregate varies more on hard content (154–212).

No wider spec (k=5) — the QSA indexer compress-ratio defines a hard ceiling at 3 drafts + 1 bonus for this architecture (established community finding).

9. Boot gates — did it really work?

From docker logs flashnext-2x64:

  • Resolved architecture: Qwen4ExpMTP (MTP registered)
  • PLE table mapped from /vllm-ngram/ngram_table.bin (95.37 GiB) + PLE page prefetch on: ... process_madvise (cap_add SYS_PTRACE present)
  • Using HummingLinearKernel/MarlinLinearKernel ... MARLIN WNA16 MoE
  • GPU KV cache size: <>= 1.8M tokens (7+ requests of 262144)
  • Application startup complete

Then verify: chat smoke → /metrics on the engine port (spec counters rising), 240K-token needle (we use a needle-in-repeated-paragraph harness with a 600+ token budget — this model spends reasoning tokens before answering), and the benches below with a warm engine.

10. Troubleshooting log (traps we hit, in order)

  1. 0-byte PLE file → "is 0 bytes but this model needs 102400491520": per section 6 — pre-size sparse to the EXACT byte count. With the shards dropped, the refill needs section 6's scoped re-fetch first; booting alone can no longer rebuild the table.
  2. fp8 KV is NOT in the final recipe: the minamism base supports it via their QSA patch (--kv-cache-dtype fp8), but we run the KV cache in full BF16 — the unquantized cache still fits (the ~1.8M-token KV gate in section 9) and fp8 KV only adds accuracy risk and edge cases on a rig with this much VRAM.
  3. Boot killed at 15 min: top-level healthCheckTimeout (section 7).
  4. JSON args mangled inline: wrapper script (section 7).
  5. docker "cannot set both Count and DeviceIDs": --gpus 2, not device=0,1 (the latter needs quote chars that llama-swap strips).
  6. Nested file bind failure when mounting the MTP overlay: merge in-place (section 5).
  7. Low wall-TPS on agent turns is prefill+TIME, not decode — see section 10.