News

July, 2026

Our paper received the WAIC Academic 2026 Multi-Modal Agents for Science (MMAS) Workshop Outstanding Paper Award!

April, 2026

Our paper “Design First, Code Later: Aesthetically Pleasing Template-Free Slides Generation” was accepted by the ACL 2026 Findings!

February, 2025

Our paper “Multi-LLM-Agents Debate-Performance, Efficiency, and Scaling Challenges” was accepted by the Blogpost track of ICLR 2025!

... see all News


Hello👋, I’m Zhiyao!

🌱I am currently pursuing a Ph.D. in Artificial Intelligence at Northwestern Polytechnical University (NWPU), advised by Prof. Zhen Wang and Dr. Shuyue Hu. My research focuses on multi-agent systems driven by large language models (LLMs) and vision-language models (VLMs). I completed my undergraduate degree in Information Security at Northwestern Polytechnical University(NWPU) in June 2024.

💡I’m particularly interested in developing robust and practical AI systems capable of effectively managing diverse and dynamic scenarios. This involves enhancing the adaptability, efficiency, and reliability of AI-driven solutions to tackle complex real-world problems.

👯Beyond academics, I hold a black belt in Taekwondo and have achieved the highest level in electronic keyboard. Whether in research or personal life, I enjoy challenges and believe in bringing curiosity and enthusiasm to everything I do.

📫Feel free to connect—I’m always excited to discuss innovative ideas or collaborate on cutting-edge research! For further details, please refer to my resume.





Publications

  1. Design First, Code Later: Aesthetically Pleasing Template-Free Slides Generation
    Cui, Zhiyao and Wang, Chenxu and Hu, Shuyue and Zhang, Yiqun and Shao, Wenqi and Zhang, Qiaosheng and Wang, Zhen
    Findings of the Association for Computational Linguistics: ACL 2026 , 2026
    arXiv GitHub
    @inproceedings{cui2026design, title = {Design First, Code Later: Aesthetically Pleasing Template-Free Slides Generation}, author = {Cui, Zhiyao and Wang, Chenxu and Hu, Shuyue and Zhang, Yiqun and Shao, Wenqi and Zhang, Qiaosheng and Wang, Zhen}, booktitle = {Findings of the Association for Computational Linguistics: ACL 2026}, year = {2026}, keywords = {main}, arxiv = {https://arxiv.org/abs/2605.26451}, image = {designfirst.png}, github = {https://github.com/sxswz213/DeepSlides} }
    Producing presentation slides automatically entails coordinating narrative structure with page-level graphic design under strict spatial constraints. For such structured multimodal tasks, a well-organized design process is essential to ensure the final quality of slides. Existing approaches rely on fixed templates or directly emit executable code, thereby both limiting the creative layout-design capabilities of LLMs and bypassing the essential slide-page design step. To address these limitations, this paper: (1) proposes a hierarchical slides generation workflow DeepSlides that systematically organizes slide design tasks without any predefined template or style, decoupling slide-page design from implementation; (2) introduces SlideDesign, a dataset tailored specifically for slides generation tasks; (3) presents a multi-agent reinforcement learning training paradigm and trains a couple of models SlideQwens for slide design and implementation. Experimental results demonstrate that our proposed framework outperforms baseline methods on evaluated metrics and achieves superior performance in human preference evaluations.
  2. AgentPanel: Toward a New Paradigm for Human–AI Collaboration in Exploring Scientific Questions
    Cui, Zhiyao and Wang, Qianyi and Yan, Haoyang and Zhang, Yiqun and Ren, Siyue and Zhang, Hangfan and Tan, Zelin and Li, Hao and Mu, Chunjiang and Cai, Dexian and Zhang, Shao and Zhang, Chen and Li, Meng and Chai, Jianan and Fan, Yuting and Ye, Zichao and Yang, Xiaolei and Lu, Xinyao and Yu, Yuyang and Lou, Wenjie and Wang, Xiaosong and Ling, Fenghua and Feng, Shiyang and Su, Mao and Zhang, Qiaosheng and Zhang, Bo and Chen, Yang and Bai, Lei and Hu, Shuyue
    2026
    arXiv Link GitHub
    @unpublished{cui2026agentpanel, title = {AgentPanel: Toward a New Paradigm for Human--AI Collaboration in Exploring Scientific Questions}, author = {Cui, Zhiyao and Wang, Qianyi and Yan, Haoyang and Zhang, Yiqun and Ren, Siyue and Zhang, Hangfan and Tan, Zelin and Li, Hao and Mu, Chunjiang and Cai, Dexian and Zhang, Shao and Zhang, Chen and Li, Meng and Chai, Jianan and Fan, Yuting and Ye, Zichao and Yang, Xiaolei and Lu, Xinyao and Yu, Yuyang and Lou, Wenjie and Wang, Xiaosong and Ling, Fenghua and Feng, Shiyang and Su, Mao and Zhang, Qiaosheng and Zhang, Bo and Chen, Yang and Bai, Lei and Hu, Shuyue}, year = {2026}, eprint = {2608.03283}, archiveprefix = {arXiv}, primaryclass = {cs.AI}, arxiv = {https://arxiv.org/abs/2608.03283}, image = {agentpanel.png}, link = {https://agentpanel.cc/}, github = {https://github.com/InternScience/AgentPanel}, keywords = {main} }
    Identifying promising scientific ideas remains an important challenge in research practice. Researchers commonly rely on small-group discussions or one-to-one interactions with a single large language model, yet these approaches often expose them to only a limited range of perspectives and directions. We present AgentPanel, a multi-agent forum for human–AI collaboration in scientific exploration. Heterogeneous agents asynchronously discuss scientific questions in a forum-style environment, while researchers can submit questions, browse and organize candidate ideas, engage agents in follow-up interactions, and optionally generate post-hoc summary reports. We evaluate AgentPanel in terms of idea quality, exploration breadth, interaction effectiveness, candidate-selection efficiency, and practical utility. Offline experiments show that AgentPanel outperforms a centralized multi-agent debate baseline. A human study with 20 participants further shows that users value AgentPanel for perspective diversity and exploration support. In experience-based comparisons with commonly used LLM tools, 65% of participants favored AgentPanel for both breadth of research directions and overall suitability for early-stage exploration. The platform is publicly available at https://agentpanel.cc/.
  3. PerPilot: Personalizing VLM-based Mobile Agents via Memory and Exploration
    Wang, Xin and Cui, Zhiyao and Li, Hao and Zeng, Ya and Wang, Chenxu and Song, Ruiqi and Chen, Yihang and Shao, Kun and Zhang, Qiaosheng and Liu, Jinzhuo and Ren, Siyue and Hu, Shuyue and Wang, Zhen
    2025
    arXiv GitHub
    @unpublished{wang2025perpilotpersonalizingvlmbasedmobile, title = {PerPilot: Personalizing VLM-based Mobile Agents via Memory and Exploration}, author = {Wang, Xin and Cui, Zhiyao and Li, Hao and Zeng, Ya and Wang, Chenxu and Song, Ruiqi and Chen, Yihang and Shao, Kun and Zhang, Qiaosheng and Liu, Jinzhuo and Ren, Siyue and Hu, Shuyue and Wang, Zhen}, year = {2025}, eprint = {2508.18040}, archiveprefix = {arXiv}, primaryclass = {cs.AI}, arxiv = {https://arxiv.org/abs/2508.18040}, image = {perpilot.png}, keywords = {main}, github = {https://github.com/xinwang-nwpu/PerPilot} }
    Vision language model (VLM)-based mobile agents show great potential for assisting users in performing instruction-driven tasks. However, these agents typically struggle with personalized instructions – those containing ambiguous, user-specific context – a challenge that has been largely overlooked in previous research. In this paper, we define personalized instructions and introduce PerInstruct, a novel human-annotated dataset covering diverse personalized instructions across various mobile scenarios. Furthermore, given the limited personalization capabilities of existing mobile agents, we propose PerPilot, a plug-and-play framework powered by large language models (LLMs) that enables mobile agents to autonomously perceive, understand, and execute personalized user instructions. PerPilot identifies personalized elements and autonomously completes instructions via two complementary approaches: memory-based retrieval and reasoning-based exploration. Experimental results demonstrate that PerPilot effectively handles personalized tasks with minimal user intervention and progressively improves its performance with continued use, underscoring the importance of personalization-aware reasoning for next-generation mobile agents.
  4. Emergence of Social Norms in Generative Agent Societies: Principles and Architecture
    Ren, Siyue and Cui, Zhiyao and Song, Ruiqi and Wang, Zhen and Hu, Shuyue
    Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI-24 , 2024
    arXiv DOI GitHub Demo
    @inproceedings{ijcai2024p0874, title = {Emergence of Social Norms in Generative Agent Societies: Principles and Architecture}, author = {Ren, Siyue and Cui, Zhiyao and Song, Ruiqi and Wang, Zhen and Hu, Shuyue}, booktitle = {Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, {IJCAI-24}}, publisher = {International Joint Conferences on Artificial Intelligence Organization}, editor = {Larson, Kate}, pages = {7895--7903}, year = {2024}, month = aug, note = {Human-Centred AI}, doi = {10.24963/ijcai.2024/874}, url = {https://doi.org/10.24963/ijcai.2024/874}, arxiv = {https://arxiv.org/abs/2403.08251}, image = {crsec.png}, github = {https://github.com/sxswz213/CRSEC}, demo = {https://www.bilibili.com/video/BV1A142187EE}, keywords = {main} }
    Social norms play a crucial role in guiding agents towards understanding and adhering to standards of behavior, thus reducing social conflicts within multi-agent systems (MASs). However, current LLM-based (or generative) MASs lack the capability to be normative. In this paper, we propose a novel architecture, named CRSEC, to empower the emergence of social norms within generative MASs. Our architecture consists of four modules: Creation & Representation, Spreading, Evaluation, and Compliance. This addresses several important aspects of the emergent processes all in one: (i) where social norms come from, (ii) how they are formally represented, (iii) how they spread through agents’ communications and observations, (iv) how they are examined with a sanity check and synthesized in the long term, and (v) how they are incorporated into agents’ planning and actions. Our experiments deployed in the Smallville sandbox game environment demonstrate the capability of our architecture to establish social norms and reduce social conflicts within generative MASs. The positive outcomes of our human evaluation, conducted with 30 evaluators, further affirm the effectiveness of our approach. Our project can be accessed via the following link: https://github.com/sxswz213/CRSEC.

Other Works

  1. Organizing, Orchestrating, and Benchmarking Agent Skills at Ecosystem Scale
    Li, Hao and Mu, Chunjiang and Chen, Jianhao and Ren, Siyue and Cui, Zhiyao and Zhang, Yiqun and Bai, Lei and Hu, Shuyue
    2026
    arXiv
    @unpublished{li2026organizing, title = {Organizing, Orchestrating, and Benchmarking Agent Skills at Ecosystem Scale}, author = {Li, Hao and Mu, Chunjiang and Chen, Jianhao and Ren, Siyue and Cui, Zhiyao and Zhang, Yiqun and Bai, Lei and Hu, Shuyue}, year = {2026}, month = mar, eprint = {2603.02176}, archiveprefix = {arXiv}, primaryclass = {cs.AI}, arxiv = {https://arxiv.org/abs/2603.02176}, image = {agentskillos.png} }

Academic Services & Activities

Academic Services

  • Reviewer, DAI 2026.

Activities

  • Gave a talk at the WAIC Academic 2026 Multi-Modal Agents for Science (MMAS) Workshop.