原帖内容

GLM-5.3 Flash (first multimodal in the GLM 5 family) just aced my form filling test. The model can only see the form as an image, it has to work out positions entirely on its own. ✅ One character per box: DOB & ID number land perfectly ✅ Checkboxes ticked dead center ✅ Zero external grounding model, pure native vision. 36K input / 10K output tokens. Small model. Serious eyes.👀