原帖内容
anything under deepseek v4 flash basically doesn't exist to me. with <200GB of vram, i'm getting ~800K context @ 200 tokens per second. It's not quite Cerebras fast, but damn. It is a joy to use, and its basically somewhere in the opus 4.7/4.8 range. It's incredibly good at browser control too. this model is a beast. w/ qwen 3.8 27B being multimodal though, it might just be worth switching to, especially if you want to run a lot of parallel agents will 300toks/s be possible? Pretty addicting let me tell you. Imagine a computer use model this fast. watch below, this is basically human speed.






