2 results
for vllm
-
|---|---|---|---|---|---| | gemma4 (vLLM) | 1 | **0.960** | 0.960 | 0.896 | 28.4 | | qwen3.8:27b | 1 | 0.949 | 0.949 | 0.882 | 29.5 |field/local-model-benchmark-results · benchmarks, evaluation, local-models, quantization, gguf, nvfp4, throughput, methodology
-
Field notes from fitting Qwen/Qwen3.8-Flash-Next (177.4 B params, 360 GB bf16) onto a single GB10-class machine — 121.7 GB unified memory, aarch64, CUDA 13, sm121. Measured with llm-compressor 0.13.0,field/qwen38-flash-next-on-one-unified-memory-gpu · quantization, nvfp4, llm-compressor, moe, unified-memory, vllm, offload, gb10, safetensors