收集 X 上新模型的能力演示

原帖内容

Qwen3.8-Flash-Next now runs on a single RTX 5090 at 68.3 tok/s — no extreme quantization, no speculative decoding. · Production-level checkpoint: GB300-validated NVFP4 checkpoint from RadixArk · 63GB host RAM — less than what 1-bit quants of this model need · The 51GB n-gram table? It lives on your NVMe, at ~0.5% throughput cost In the video: one prompt → a playable minecraft-style world, then FreeToken on my 5090 desktop fixing a real FreeToken bug in 10.5 minutes. Meet FreeToken 🧵