原帖内容

Introducing Kimi K3 for autoresearch One of the hard parts of pointing models at open-ended research is they often prematurely call "Eureka" on claims. However, in our tests Kimi has proven to be a much more grounded research companion. In the task below, we had Kimi reproduce findings from "Self-Distilled RLVR". It correctly reproduces the paper's core mechanism, showing RLSD maintains policy entropy while GRPO's collapses over training, but it doesn't overclaim. It noted that the paper never states how many seeds were run, and that in our LoRA setup RLSD doesn't separate from GRPO by a meaningful margin.