AIBuildAI open-sources PostTrain Agent, tops PostTrainBench with 46.6 points
AIBuildAI released PostTrain Agent, an open-source recursive self-improving agent that autonomously performs LLM post-training, including algorithm design, data collection, coding, and experiments, without human intervention. On the PostTrainBench benchmark, it scored 46.6, ranking first among frontier models and agent systems, close to human expert performance (51.1). The team also open-sourced the code, knowledge base, and technical report.
Coverage timeline
机器之心机器之心
给定一个基座模型、一个目标能力和算力预算,RSI Agent 自己设计后训练算法、收集和整理数据、写代码、跑实验、迭代改进,最终交出一个训练好的模型,全程无人介入。在自主后训练基准 PostTrainBench 上,它以 46.6 分排名第一,超过所有前沿模型与 Agent 系统,接近人类专家表现(51.1)。 近日,AIBuildAI 团队发布了 PostTrain Agent ,一个用于自主后训练大语言模型的递归自我改进(RSI)Agent,并开源了全部代码与配套知识库。 开源地址:https://github.com/aibuildai-inc/aibuildai-llm-posttrain-agent 知识库:https://github.com/aibuildai-inc/aibuildai-knowledge-base 技术报告:https://github.com/aibuildai-inc/aibuildai-llm-posttrain-agent/blob/main/docs/aibuildai-llm-post-train-agent.pdf 为什么需要自动化的后训练 后训练(post-training)是把一个基座模型变成能遵循指令、会推理、会用工具、懂某个领域的模型的过程——也是一家公司把通用模型变成「自己的模型」的方式。它涉及的决策不少:用什么数据、怎么配比;用监督微调(SFT)、偏好优化还是强化学习;学习率、训练步数、最后交付哪个 checkpoint。 这些决策彼此耦合:合适的算法取决于数据配比,有效的数据配比又取决于基座模型。交付一个后训练好的模型,往往要一个有经验的团队花上数周,每次实验还要消耗可观的算力。AIBuildAI PostTrain Agent 研究的问题是:能否把这件事完全交给 AI—— 输入基座模型、目标能力和算力预算,Agent 自己完成后训练的全部流程,输出一个后训练好的模型。 PostTrainBench:给 Agent 的一道考题 PostTrainBench(Rank et al., ICML 2026)把这个问题做成了一个基准:给定基座模型(如 Qwen3-4B-Base)和目标能力(如工具调用),Agent 在单张 H100 上、10 小时内自主设计并执行后训练方案,产出的 checkpoint 在基准
