Sunqi Fan

self-image.png

I am a second-year CS PhD student at Tsinghua University, advised by Prof. Shi-Min Hu. I am also an intern at Tencent Hunyuan Foundation Model Department starting from November 2025.

My research interest lies in multimodal LLM / agent and physical AI, with a focus on designing intelligent agents to understand the dynamic real world and act.

Some of my recent works:

(1) Video Agent: Agentic Keyframe Search, Tool-Augmented Video Reasoning [NeurIPS’25], Video-Guided Agentic Tasks [ECCV’26];

(2) GUI Agent: GUICrafter, HyMobileAgent [Technical Report].

Previously, I obtained my Bachelor’s degree in Computer Science from Tsinghua University. You may find my CV here: Sunqi’s Curriculum Vitae. Feel free to reach out to me via email or WeChat if you have any questions or just want to chat!

Email: stephensunqifan@gmail.com

WeChat: fsq159357FSQ

Google Scholar / Github / Twitter / LinkedIn

news

Sep 11, 2026 Attend ECCV 2026 in Malmö, Sweden 🇸🇪. [Photo]
Aug 20, 2026 I give a talk on Video-Guided Agentic Tasks at Zhiyuan.
Jun 30, 2026 We release GUICrafter, a weakly-supervised and self-evolving GUI agent.
Jun 21, 2026 One paper on video understanding is accpeted by ECCV’26! Check out this project page.
Jun 20, 2026 A wonderful trip to Inner Mongolia. [Photo 1] [Photo 2]
Dec 06, 2025 Attend NeurIPS 2025 in San Diego. [Photo 1] [Photo 2] [Photo 3]
Oct 22, 2025 Attend CNCC 2025 in Harbin.
Sep 18, 2025 Our paper on tool-augmented VideoQA is accpted by NeurIPS 2025!
Sep 01, 2025 Begin my journey to pursue a CS PhD at Tsinghua University.
Jun 06, 2025 Attend VALSE 2025 in Zhuhai.

selected publications

  1. GUICrafter.png
    GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots
    Sunqi Fan, Lingshan Chen, Runqi Yin, and 4 more authors
    2026
  2. vg-gui-TASKER.png
    Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction
    Sunqi Fan, Qingle Liu, Runqi Yin, and 2 more authors
    In ECCV, 2026
  3. VideoTool.png
    Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task
    Sunqi Fan, Jiashuo Cui, Meng-Hao Guo, and 1 more author
    In NeurIPS, 2025
  4. oral
    FlexKBQA.png
    FlexKBQA: a flexible LLM-powered framework for few-shot knowledge base question answering
    Zhenyu Li*Sunqi Fan*, Yu Gu, and 5 more authors
    In AAAI, 2024

internships

Tencent Hunyuan
Research Intern, Tencent Hunyuan
Mentor: Han Hu
2025.11 - Now
Megvii Research
Research Intern, Megvii Research
Mentor: Jiajun Liang
2023.4 - 2023.12