原帖内容
holy shit. a 3bit quant just built me a working game, driving hermes agent from one sentence, tested it itself, and won me over. that is deepseek v4 flash 0731 IQ3_XXS, dancing on one dgx spark, the smallest 3bit that fits on 128gb. i gave it one line, "build a simple space shooter game i can test right now". and watch it plan for build, wrote it in a single file, ran it, debugged its own test, and verified the result before it ever showed me. it oneshot a working game, tested correctly, and it felt genuinely reliable. a 3bit quant is not supposed to do that. the numbers so you know it is real: 16.5 tok/s decode, 332 prefill, 128k context, no spec decode, every layer on the gpu, 284b total, 13b active. i think you are watching the 284b weights come through, even squeezed to 3bit there is enough model left to actually reason, and you can see it in how it works the problem. you might never run this. watch it anyway. a 284b model on a box on a shelf, taking one sentence and shipping a game, fully local, nothing leaving the network. this is what open weights on hardware you own looks like now.






