About Me
I am Huatong Song (宋华彤), a second-year Master’s student at the Gaoling School of Artificial Intelligence (GSAI), Renmin University of China (RUC), supervised by Prof. Xin Zhao. Previously, I received dual Bachelor’s degrees from RUC in July 2025: a Bachelor of Engineering in Artificial Intelligence from the GSAI and a Bachelor’s degree in Finance from the Gaoli Institute.
My recent research focuses on AI Agents, particularly agentic post-training, including search agents, coding agents, and general agents, with the goal of improving their reasoning, tool-use, and long-horizon interaction capabilities. I am also particularly interested in training through complex agent harnesses and recursive self-improvement (RSI).
I am expected to graduate in June 2028. If you are interested in me, please feel free to contact me.
Experience
- 2026.09 – Now · ByteDance Seed Model-Product Posttrain-Work · LLM Research Intern ·
- 2026.04 – 2026.09 · IQuest Research · LLM Research Intern ·
- 2025.11 – 2026.04 · Boss Zhipin Nanbeige · LLM Research Intern ·
- 2025.05 – 2025.11 · ByteDance Seed-Edge · LLM Research Intern ·
Research
(* indicates equal contribution, † indicates corresponding author)
Company Technical Reports
Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model
Technical Report, 2026UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
Technical Report, 2025Seed1.8 Model Card: Towards Generalized Real-World Agency
Model Card / Technical Report, 2026
Selected Publications
Here, I list my work as a first or co-first author. For a complete list, please see my Google Scholar profile.
ClawGym II: Exploring Black-Box RL on Agent Harness
Technical Report, 2026ClawGym: A Scalable Framework for Building Effective Claw Agents
Technical Report, 2026SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training
Technical Report, 2026SWE-World: Building Software Engineering Agents in Docker-Free Environments
Preprint, 2026R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning
EMNLP-Findings, 2025SimpleDeepSearcher: Deep Information Seeking via Web-Powered Reasoning Trajectory Synthesis
EMNLP-Findings, 2025R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
Technical Report, 2025Yulan-mini: An Open Data-Efficient Language Model
Technical Report, 2024YuLan-Mini: Pushing the Limits of Open Data-Efficient Language Model
ACL-Main-Oral, 2025
Awards
- 2026 Linghang Dean’s Scholarship, GSAI
- 2025 Outstanding Undergraduate Graduation Thesis (Top 5% of the GSAI)
- 2024 China National Scholarship (top 1.5%)
- 2024 Linghang Dean’s Scholarship, GSAI
- 2022 First-Class Academic Scholarship, RUC
