Dr. Jiahao Wang is a Principal Research Engineer at the Central Research Institute of Huawei 2012 Laboratories. His current research focuses on multimodal foundation models and Computer-Using Agents (CUA), including GUI/CLI agents, OS/mobile navigation, visual coding, scalable agent data flywheels, and pre-training/post-training techniques for agentic multimodal models.
Before moving to multimodal agent research, he worked on AI safety and trustworthy foundation models, with a focus on LLM value alignment, hallucination evaluation, factuality alignment, and frontier AI governance. His earlier Ph.D. research at IRIP Lab, Beihang University, supervised by Prof. Yunhong Wang, focused on trustworthy multimodal learning and fine-grained video understanding.
I am interested in research collaborations and strong candidates in multimodal agents, GUI agents, CUA training, agentic RL, AI alignment, and trustworthy foundation models.
Research Agenda
- Computer-Using Agents: GUI grounding, GUI/CLI navigation, OS/mobile/web agents, visual coding, and unified action spaces across interfaces and tools.
- Agent Data Flywheels: scalable data generation, trajectory quality control, verifiable task expansion, evaluation-to-training loops, and data mixture design for pre-training and SFT.
- Agentic Training: post-training, reinforcement learning infrastructure, process reward modeling, on-policy distillation, and long-horizon task optimization.
- Trustworthy Foundation Models: value alignment, hallucination and factuality evaluation, pluralistic alignment, safety evaluation, and governance-oriented model assessment.
News
- 2026.04: We are recruiting outstanding interns and full-time researchers in multimodal agentic models, GUI agents, and CUA training. Master’s and Ph.D. candidates are welcome to connect.
- 2025.10: Our work Diverse Human Value Alignment for Large Language Models via Ethical Reasoning was accepted as an oral paper by AIES 2025.
- 2025.08: Our work SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs was accepted by ACM Multimedia 2025.
- 2025.08: Our work C-FAITH: A Chinese Fine-Grained Benchmark for Automated Hallucination Evaluation was accepted by CIKM 2025.
- 2024.12: We released a technical challenge on Huawei Chaspark about AI value alignment evaluation.
- 2024.05: Our work EvCap: Element-Aware Video Captioning was accepted by IEEE TCSVT.
- 2022.07: Our unsupervised video segmentation method PACE was accepted by IJCAI 2022.
Selected Publications

Diverse Human Value Alignment for Large Language Models via Ethical Reasoning
Jiahao Wang, Songkai Xue, Jinghui Li, Xiaozhen Wang, AIES 2025 Oral
- Built an ethical-reasoning paradigm for LLM value alignment from interdisciplinary ethical decision-making models.
- Integrated four complementary ethical theories for multi-lens ethical impact analysis and diverse human value alignment.
- Improved social norm inference and cultural sensitivity on SafeWorld-style value alignment evaluation.

SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs
Bei Yan, Zhiyuan Chen, Yuecong Min, Jie Zhang, Jiahao Wang, Xiaozhen Wang, Shiguang Shan, ACM Multimedia 2025
- Built a scalable hallucination benchmark with controllable image-instruction pairs and fine-grained ground-truth answers.
- Covered diverse hallucination types, tasks, and scenarios with over 30K image-instruction pairs.
- Evaluated over 20 representative LVLMs and revealed factuality hallucination and semantic perturbation sensitivity.

C-FAITH: A Chinese Fine-Grained Benchmark for Automated Hallucination Evaluation
Xu Zhang, Zhifei Liu, Jiahao Wang, Huixuan Zhang, Fan Xu, Junzhe Zhang, Xiaojun Wan, CIKM 2025
- Built a Chinese fine-grained benchmark for automated hallucination evaluation.
- Supports more precise assessment of factual consistency and hallucination failure modes in Chinese LLM outputs.
- Complements multimodal hallucination evaluation with a text-centric factuality benchmark.

PACE: Predictive and Contrastive Embedding for Unsupervised Action Segmentation
Jiahao Wang, Jie Qin, Yunhong Wang, Annan Li, IJCAI 2022
- Proposed a unified framework that exploits both predictability and similarity information for unsupervised action segmentation.
- Learned predictive and contrastive embeddings for accurate action boundary detection.
- Achieved up to 26.9% F1-score improvement over prior state-of-the-art methods.

Few-shot Fine-grained Action Recognition via Bidirectional Attention and Contrastive Meta-learning
Jiahao Wang, Yunhong Wang, Sheng Liu, Annan Li, ACM Multimedia 2021
- Proposed the few-shot fine-grained action recognition problem and a framework for recognizing unseen fine-grained actions with few support samples.
- Combined task-driven and saliency-supervised signals to capture subtle action details.
- Introduced contrastive meta-learning for discriminative representation learning in low inter-class variance scenarios.

EvCap: Element-Aware Video Captioning
Sheng Liu, Annan Li, Yuwei Zhao, Jiahao Wang, Yunhong Wang, IEEE TCSVT 2024
- Proposed element-aware usage of linguistic features to reduce video-captioning hallucinations.
- Designed a multimodal multi-branch encoder-decoder framework with flexible feature fusion.
- Achieved strong CIDEr improvements across MSVD, MSR-VTT, VATEX, and TVC.
Publication List
- Diverse Human Value Alignment for Large Language Models via Ethical Reasoning. AIES 2025 Oral.
Jiahao Wang, Songkai Xue, Jinghui Li, Xiaozhen Wang - SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs. ACM Multimedia 2025.
Bei Yan, Zhiyuan Chen, Yuecong Min, Jie Zhang, Jiahao Wang, Xiaozhen Wang, Shiguang Shan - C-FAITH: A Chinese Fine-Grained Benchmark for Automated Hallucination Evaluation. CIKM 2025.
Xu Zhang, Zhifei Liu, Jiahao Wang, Huixuan Zhang, Fan Xu, Junzhe Zhang, Xiaojun Wan - VAPO: ValueCoT-Enhanced Search-Based Prompt Optimization for Human Value Alignment. Preprint.
Xuening Feng, Jiahao Wang et al. - EvCap: Element-Aware Video Captioning. IEEE TCSVT 2024.
Sheng Liu, Annan Li, Yuwei Zhao, Jiahao Wang, Yunhong Wang - PACE: Predictive and Contrastive Embedding for Unsupervised Action Segmentation. IJCAI 2022.
Jiahao Wang, Jie Qin, Yunhong Wang, Annan Li - Bidirectional Maximum Entropy Training with Word Co-occurrence for Video Captioning. IEEE TMM 2022.
Sheng Liu, Annan Li, Jiahao Wang, Yunhong Wang - Few-Shot Fine-Grained Action Recognition via Bidirectional Attention and Contrastive Meta-Learning. ACM Multimedia 2021.
Jiahao Wang, Yunhong Wang, Sheng Liu, Annan Li - Will You Ever Become Popular? Learning to Predict Virality of Dance Clips. ACM TOMM 2021.
Jiahao Wang, Yunhong Wang, Nina Weng, Tianrui Chai, Faxi Zhang, Sansi Yu, Annan Li - Two-Stream Temporal Convolutional Network for Dynamic Facial Attractiveness Prediction. IEEE ICPR 2020.
Nina Weng, Jiahao Wang, Annan Li, Yunhong Wang - Assessing Action Quality via Attentive Spatio-Temporal Convolutional Networks. PRCV 2020.
Jiahao Wang, Zhengyin Du, Annan Li, Yunhong Wang - Atrous Temporal Convolutional Network for Video Action Segmentation. IEEE ICIP 2019.
Jiahao Wang, Zhengyin Du, Annan Li, Yunhong Wang
Honors and Awards
- 2024 Huawei Annual President’s Individual Award
- 2024 Huawei Battlefield Hero Award
- 2022, 2023 Huawei Rising Star Award
- 2022 Outstanding Doctoral Graduate Award of Beihang University
- 2017 Doctoral Scholarship of Beihang University
- 2017 Outstanding Undergraduate Graduate Award of Beijing
- 2016 Meritorious Winner of MCM/ICM
- 2016 Silver Medal of the 26th Feng Ru Competition of Beihang University
Education
- 2017.09 - 2022.07, Ph.D. in Computer Science and Technology, Beihang University, Beijing, China. Supervised by Prof. Yunhong Wang at IRIP Lab.
- 2013.09 - 2017.06, B.E. in Computer Science and Technology, Beihang University, Beijing, China.
Experience
- 2025 - Present, Central Research Institute / Foundation Model Lab, Huawei 2012 Laboratories, Shenzhen, China.
- 2022 - 2025, AI Safety and Trustworthiness Research, Huawei 2012 Laboratories, Shenzhen, China.
- 2021.05 - 2021.10, Visual Intelligence Center, Meituan, Beijing, China.