I am currently with ByteDance Seed and a Ph.D. candidate at the Institute of Automation, Chinese Academy of Sciences (CASIA) and the University of Chinese Academy of Sciences (UCAS), advised by Prof. Tieniu Tan.
My research focuses on multimodal intelligence. I believe natural, real-time interaction is an important next step for multimodal models. I am particularly interested in voice agents and streaming audio-visual systems that continuously perceive their surroundings and interact with people through full-duplex conversations. Questions that interest me include when an agent should respond, how it should handle interruptions, and how memory and tool use can support sustained interaction.
My previous work spans multimodal model training and evaluation, post-training and reward modeling, and agentic multimodal systems, with 15+ papers as first, co-first, or corresponding author at leading venues. These experiences shape my current interest in building interactive agents that understand people and their environments, respond with appropriate timing, and act reliably.