原帖内容

DeepSeek shipped their first vision model, so I ran it through my form-filling test. Not an electronic form, an image is all the model ever sees. No coordinates, no DOM. V4 Flash Vision probes positions, screenshots itself, places each element and nudges it into place on its own. ✅ every field filled ✅ self-verified before submit ⚠️ ticks land next to the box, not in it ⚠️ character-box drift on longer strings 5m48s · 288k in / 38k out Misses a few boxes, but it got there.