收集 X 上新模型的能力演示

原帖内容

Tested DeepSeek V4 Pro 0813, Kimi K3, and GLM 5.2 through FlappyBench. 3 models, same prompt with the /design command. Scored on gameplay features, UX/UI, and cost. 🔹 DeepSeek V4 Pro 0813 → 8/10 · $0.0005 🔹 Kimi K3 → 9.5/10 · $0.0740 🔹 GLM 5.2 → 9/10 · $0.0480 DeepSeek is 148x cheaper than Kimi and 96x cheaper than GLM, for 84% of the quality. But $0.0005 for a playable one-shot game is insane. Kimi K3 had the best output, with GLM 5.2 close behind.