原帖内容

Just tested @Kimi_Moonshot K3 against GPT-5.6 and Grok 4.5 on the same task: build a Subway Surfers game. Kimi K3 was my favorite. Better UI, smoother gameplay, and the overall experience felt more polished. GPT-5.6 was solid too, while Grok 4.5 fell behind on the UI. The interesting part is that all three needed at least one follow-up prompt. Kimi also took a bit longer to generate, but the extra time paid off. Ran Kimi K3 through @nebiustf. Low inference cost + a really capable open source model is a great combo.