Portrait of Ziyue Lin

Ziyue Lin

Department of Data Science and Artificial Intelligence
The Hong Kong Polytechnic University

Year 1 PhD Student at PolyU DSAI

I work on reliable and data-efficient multimodal systems for visual understanding, generation, and interaction with the physical world.

Hi, my name is Ziyue Lin. Starting in September 2026, I will pursue my Ph.D. in the Department of Data Science and Artificial Intelligence (DSAI) at The Hong Kong Polytechnic University, supervised by Prof. Xingyi Yang.

I received my master’s degree in Artificial Intelligence from the Department of Mathematics at The University of Hong Kong and my B.S. degree from The Chinese University of Hong Kong, Shenzhen. I previously worked at SenseTime Research, focusing on multimodal low-level vision models.

My research spans Multimodal Large Language Models, Generative AI, and Computer Vision, with a particular interest in building reliable and data-efficient systems for understanding, generating, and interacting with the visual world.

01

Multimodal Intelligence

Building models that connect language with rich visual signals.

02

Generative AI

Creating controllable and coherent visual content from human intent.

03

Efficient Learning

Learning reliable representations with less data and computation.

04

Visual Understanding

Reasoning about complex scenes and the physical world.

DRDD, my first-author paper, was accepted to CVPR 2026.

Started working on Uni-Lens at SenseTime Research.

ARRA was accepted to AAAI 2026 as an oral presentation.

MLLM-Bench was accepted to NAACL 2025.

CVPR 2026DRDD project preview

Decoupled Residual Denoising Diffusion Models for Unified and Data-Efficient Image-to-Image Translation

Ziyue Lin, Jiahe Hou, Hongyu Xia, Xinrui Xie, Feifei Wang, Yuyin Zhou, Wei Wang, Jiawei Liu, Liangqiong Qu

Decouples diffusion into stochastic domain alignment and deterministic semantic mapping for unified, data-efficient image translation.

PreprintFedVLMBench project preview

FedVLMBench: Benchmarking Federated Fine-Tuning of Vision-Language Models

Weiying Zheng, Ziyue Lin, Pengxin Guo, Yuyin Zhou, Feifei Wang, Liangqiong Qu

A standardized benchmark and toolkit for privacy-preserving federated fine-tuning of vision-language models.

AAAI 2026 · OralARRA project preview

Unleashing the Potential of Large Language Models for Text-to-Image Generation through Autoregressive Representation Alignment

Xing Xie, Jiawei Liu, Ziyue Lin, Huijie Fan, Zhi Han, Yandong Tang, Liangqiong Qu

Aligns autoregressive representations to improve globally coherent text-to-image generation without architectural changes.

NAACL 2025MLLM-Bench project preview

MLLM-Bench: Evaluating Multi-modal LLMs using GPT-4V

Wentao Ge, Shunian Chen, Guiming Hardy Chen, et al., Ziyue Lin, et al.

A pairwise-comparison benchmark revealing the diverse capabilities of 21 widely used multimodal large language models.

The Hong Kong Polytechnic University

Ph.D. in Data Science and Artificial Intelligence

The University of Hong Kong

M.Sc. in Artificial Intelligence

CUHK-Shenzhen

B.Sc. in Data Science and Big Data Technology

SenseTime Research

Research intern · Multimodal low-level vision