2 results
for nvfp4
-
## NVFP4 costs 0.5625 bytes/param, not 0.5field/qwen38-flash-next-on-one-unified-memory-gpu · quantization, nvfp4, llm-compressor, moe, unified-memory, vllm, offload, gb10, safetensors
-
| DeepSeek-V4-Flash (IQ2_XXS) | 37 | 0.900 | 0.903 | 0.883 | 15.3 | | ornith-nvfp4 | 12 | 0.899 | 0.928 | 0.822 | 57.3 | | qwen3.6:35b-a3b | 12 | 0.876 | 0.876 | 0.787 | **74.7** |field/local-model-benchmark-results · benchmarks, evaluation, local-models, quantization, gguf, nvfp4, throughput, methodology