Yi-Fan Zhang 张一帆ByteDance Seed

Multimodal intelligence / Research

Intelligence
in interaction.

I study how multimodal agents can see, listen, and interact with people in real time.

Explore selected research
Yi-Fan Zhang
Yi-Fan ZhangPh.D. Candidate, CASIA / UCASAdvised by Prof. Tieniu Tan

Natural conversation. Continuous perception.

About Me

I am currently with ByteDance Seed and a Ph.D. candidate at the Institute of Automation, Chinese Academy of Sciences (CASIA) and the University of Chinese Academy of Sciences (UCAS), advised by Prof. Tieniu Tan.

My research focuses on multimodal intelligence. I believe natural, real-time interaction is an important next step for multimodal models. I am particularly interested in voice agents and streaming audio-visual systems that continuously perceive their surroundings and interact with people through full-duplex conversations. Questions that interest me include when an agent should respond, how it should handle interruptions, and how memory and tool use can support sustained interaction.

My previous work spans multimodal model training and evaluation, post-training and reward modeling, and agentic multimodal systems, with 15+ papers as first, co-first, or corresponding author at leading venues. These experiences shape my current interest in building interactive agents that understand people and their environments, respond with appropriate timing, and act reliably.

Total citations
7,000+
Citations on my most-cited first-author paper
2,600+
First / co-first-author papers with 100+ citations
7

Citation snapshot: . Latest figures on Google Scholar.

Open to research discussions and collaborations on real-time multimodal interaction, voice agents, and the training and evaluation of interactive systems.

Previously, I have been fortunate to work with Prof. Jingdong Wang at Microsoft Research Asia and Prof. Rong Jin at Alibaba DAMO Academy. I have also interned at ByteDance, Kuaishou, Skywork, and Squirrel AI.

Research Interests

Current focus

Real-Time Multimodal Interaction

Voice agents, streaming audio-visual perception, and full-duplex dialogue. I am interested in how an agent follows a changing environment and handles turn-taking, pauses, and interruptions with appropriate timing.

Multimodal Agents & Memory

Tool use, active exploration, and memory for sustained, context-aware interaction. My work includes Thyme, Skywork-R1V4, and Omni-DeepSearch. I am interested in how perception and memory inform what an agent does next.

Post-Training & Reward Modeling

Learning from preferences and feedback to improve multimodal behavior, through MM-RLHF, R1-Reward, and BaseReward. I am also interested in rubric-based rewards and self-evolving reward systems.

Earlier interests include continual learning, out-of-distribution generalization, time-series forecasting, AI for education, and content moderation.

News

Earlier news, 2022–2025
  • Sep 2025
    Released Thyme, thinking beyond images with executable code generation.
  • Jul 2025
    Released Kwai Keye-VL.
  • May 2025
    MM-RLHF and DAMO accepted by ICML 2025. Released R1-Reward.
  • Apr 2025
    Released MME-Unify for unified multimodal models.
  • Feb 2025
    Released MM-RLHF, with 120K human preference annotations.
  • Jan 2025
    MME-RealWorld accepted by ICLR 2025.
  • Jun 2024
    Released SliME for high-resolution multimodal models.
  • Mar 2024
    Two papers on in-context learning and symbolic reasoning accepted by NAACL 2024.
  • Oct 2023
    OneNet accepted by NeurIPS 2023.
  • May 2023
    AdaNPC accepted by ICML 2023, DRM accepted by KDD 2023.
  • Jan 2023
    Environment Label Smoothing accepted by ICLR 2023.
  • Apr 2022
    DDG selected for an Oral presentation at CVPR 2022.

Selected Publications

Selected work grouped by research direction. Full publication list on Google Scholar. * Equal contribution. † Corresponding author.

Interaction, Streaming Perception & Multimodal Agents

Selected collaborations and first-author work that connect to my current interests.

Post-Training & Reward Modeling

Multimodal Training & Evaluation

For earlier work on domain generalization, adaptation, and time-series forecasting, see Google Scholar, AdaNPC, and OneNet.

Experience

ByteDance Seed

Current affiliation

Research interests in real-time multimodal interaction and agents.

Previous Research Internships

Kuaishou Technology

Multimodal Large Language Models

Skywork AI

Agentic Multimodal Systems

ByteDance

Research Intern

Squirrel AI

LLMs for Education

Alibaba DAMO Academy

Advised by Prof. Rong Jin

Microsoft Research Asia

Advised by Prof. Jingdong Wang

Selected Awards

Professional Service

Conference Reviewer

ML/AI: ICML (2022–2026), NeurIPS (2022–2026), ICLR (2023–2026), AISTATS (2025), AAAI (2023–2024)
Vision: CVPR (2022–2024), ICCV (2023, 2025), ECCV (2022, 2024)
NLP: ACL (2025), EMNLP (2023–2024), NAACL (2024)

Journal Reviewer

IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)
IEEE Transactions on Image Processing (TIP)
International Journal of Computer Vision (IJCV)
Transactions on Machine Learning Research (TMLR)
IEEE Transactions on Information Forensics & Security (T-IFS)

Workshop Organizer

PC Member for MILETS@PAKDD’23, DMLR@ICML’23