
Ph.D. student in Computer Science, Yale University · advised by Rex Ying & David van Dijk
I study generative models for discrete and non-Euclidean data — sequences, permutations, and geometric structures — and how to steer them with reinforcement learning and distillation.
t = 0
I'm a Ph.D. student in Computer Science at Yale, co-advised by Prof. Rex Ying and Prof. David van Dijk. Before Yale I worked with Prof. Jian Tang at Mila on geometric generative modeling, and I hold a B.Eng. in Computer Science from the ACM Honors Class at Shanghai Jiao Tong University.
My research is on generative modeling, mostly diffusion and flow models for discrete and structured data: discrete diffusion for language that conditions on its whole trajectory (CaDDi), diffusion over permutations via reflected soft ranks (Soft-Rank Diffusion), and, in earlier work, geometric models on the torus and in SE(3) (DiffPack, E3Bind).
I'm equally interested in steering and accelerating these models. I've used policy gradients to learn a model's own generation order (ICML'26 Spotlight), RL to post-train LLMs to refine sequences through verifiable edits (STRIDE), and at Meta I'm working on distilling diffusion language models for efficient generation.
Yangtian Zhang*, Zhe Wang*, Arthur Gretton, Rex Ying, David van Dijk, Michalis Titsias, Jiaxin Shi (*equal contribution)
Treat the generation order as a latent variable. Tokens are inserted anywhere, the order is learned by variational inference, and left-to-right AR and discrete diffusion fall out as special cases.
@article{zhang2026insertion,
title={Variational Learning for Insertion-based Generation},
author={Zhang, Yangtian and Wang, Zhe and Gretton, Arthur and Ying, Rex and van Dijk, David and Titsias, Michalis and Shi, Jiaxin},
journal={International Conference on Machine Learning (ICML), Spotlight (top 2.2\%)},
year={2026}
}
Sizhuang He*, Yangtian Zhang*, Shiyang Zhang, David van Dijk (*equal contribution)
Soft-Rank Diffusion. Lift a permutation to continuous soft ranks, diffuse them with reflection at the boundaries, and denoise with contextual Plackett–Luce heads.
@article{he2026permutation,
title={Learning Permutation Distributions via Reflected Diffusion on Ranks},
author={He, Sizhuang and Zhang, Yangtian and Zhang, Shiyang and van Dijk, David},
journal={International Conference on Machine Learning (ICML)},
year={2026}
}
Daiheng Zhang, Shiyang Zhang, Sizhuang He, Yangtian Zhang, Syed Asad Rizvi, David van Dijk
An LLM that writes its reasoning as an executable chain of INSERT / DELETE / REPLACE edits, trained with SFT on shortest edit paths and then RL.
@article{zhang2026stride,
title={STRIDE: Post-Training LLMs to Reason and Refine Bio-Sequences via Edit Trajectories},
author={Zhang, Daiheng and Zhang, Shiyang and He, Sizhuang and Zhang, Yangtian and Rizvi, Syed Asad and van Dijk, David},
journal={International Conference on Machine Learning (ICML)},
year={2026}
}
Yangtian Zhang*, Leyao Wang*, Hiren Madhu, Ngoc Bui, Walter Roznyatovskiy, Rex Ying (*equal contribution)
Sparse user histories leave facets of a persona unobserved. Split each history into facets, link peers facet-by-facet in a multiplex graph, and borrow what is missing.
@article{zhang2026copersona,
title={CoPersona: Collaborative Persona Graphs for Robust LLM Personalization},
author={Zhang, Yangtian and Wang, Leyao and Madhu, Hiren and Bui, Ngoc and Roznyatovskiy, Walter and Ying, Rex},
journal={ACM SIGKDD Conference on Knowledge Discovery and Data Mining},
year={2026}
}
Yangtian Zhang*, Sizhuang He*, Daniel Levine, Lawrence Zhao, David Zhang, Syed A Rizvi, Emanuele Zappala, Rex Ying, David van Dijk (*equal contribution)
CaDDi. Each denoising step conditions on the entire trajectory, not just the last state, so a causal LM can run discrete diffusion and reuse pretrained weights unchanged.
@article{zhang2025caddi,
title={Non-Markovian Discrete Diffusion with Causal Language Models},
author={Zhang, Yangtian and He, Sizhuang and Levine, Daniel and Zhao, Lawrence and Zhang, David and Rizvi, Syed A and Zappala, Emanuele and Ying, Rex and van Dijk, David},
journal={Advances in Neural Information Processing Systems},
year={2025}
}
Chengkai Liu*, Yangtian Zhang*, Jianling Wang, Rex Ying, James Caverlee (*equal contribution)
FlowCF. Flow from a behavior-guided prior to binary implicit feedback with a discrete flow, for accurate and fast generative recommendation.
@article{liu2025flow,
title={Flow Matching for Collaborative Filtering},
author={Liu, Chengkai and Zhang, Yangtian and Wang, Jianling and Ying, Rex and Caverlee, James},
journal={ACM SIGKDD Conference on Knowledge Discovery and Data Mining},
year={2025}
}
Sizhuang He, Daniel Levine, Ivan Vrkic, Marco Francesco Bressana, David Zhang, Syed Asad Rizvi, Yangtian Zhang, Emanuele Zappala, David van Dijk
Flow matching recast as a Volterra integral equation and solved by a causal language model that attends over the whole discretized path.
@article{he2024calmflow,
title={CaLMFlow: Volterra Flow Matching using Causal Language Models},
author={He, Sizhuang and Levine, Daniel and Vrkic, Ivan and Bressana, Marco Francesco and Zhang, David and Rizvi, Syed Asad and Zhang, Yangtian and Zappala, Emanuele and van Dijk, David},
journal={arXiv preprint arXiv:2410.05292},
year={2024}
}
Jianghao Lin, Jiaqi Liu, Jiachen Zhu, Yunjia Xi, Chengkai Liu, Yangtian Zhang, Yong Yu, Weinan Zhang
Yangtian Zhang*, Zuobai Zhang*, Bozitao Zhong, Sanchit Misra, Jian Tang (*equal contribution)
Diffuse only what can move. Side-chain torsions are denoised on the torus, one angle at a time from χ₁ to χ₄, with a 60× smaller model than prior art.
@article{zhan2023diffpack,
title={DiffPack: A Torsional Diffusion Model for Autoregressive Protein Side-Chain Packing},
author={Zhang, Yangtian and Zhang, Zuobai and Zhong, Bozitao and Misra, Sanchit and Tang, Jian},
journal={Advances in Neural Information Processing Systems},
year={2023}
}
Yangtian Zhang*, Huiyu Cai*, Chence Shi, Bozitao Zhong, Jian Tang (*equal contribution)
Dock like AlphaFold folds. An equivariant network that iteratively refines the ligand pose inside the pocket, end to end.
@article{zhang2022e3bind,
title={E3Bind: An End-to-End Equivariant Network for Protein-Ligand Docking},
author={Zhang, Yangtian and Cai, Huiyu and Shi, Chence and Zhong, Bozitao and Tang, Jian},
journal={Proceedings of the International Conference on Learning Representations (ICLR)},
year={2023}
}

Minghao Xu, Zuobai Zhang, Jiarui Lu, Zhaocheng Zhu, Yangtian Zhang, Chang Ma, Runcheng Liu, Jian Tang
@article{xu2022peer,
title={PEER: A Comprehensive and Multi-task Benchmark for Protein Sequence Understanding},
author={Xu, Minghao and Zhang, Zuobai and Lu, Jiarui and Zhu, Zhaocheng and Zhang, Yangtian and Chang, Ma and Liu, Runcheng and Tang, Jian},
journal={Advances in Neural Information Processing Systems (Datasets and Benchmarks Track)},
volume={35},
pages={35156--35173},
year={2022}
}

Zhaocheng Zhu, Chence Shi, Zuobai Zhang, Shengchao Liu, Minghao Xu, Xinyu Yuan, Yangtian Zhang, Junkun Chen, Huiyu Cai, Jiarui Lu, Chang Ma, Runcheng Liu, Louis-Pascal Xhonneux, Meng Qu, Jian Tang
@article{zhu2022torchdrug,
title={TorchDrug: A Powerful and Flexible Machine Learning Platform for Drug Discovery},
author={Zhu, Zhaocheng and Shi, Chence and Zhang, Zuobai and Liu, Shengchao and Xu, Minghao and Yuan, Xinyu and Zhang, Yangtian and Chen, Junkun and Cai, Huiyu and Lu, Jiarui and others},
journal={Preprint},
year={2022}
}

Minghuan Liu, Hangyu Wang, Yangtian Zhang, Minkai Xu, Zhengbang Zhu, Weinan Zhang
@article{liuautogail,
title={Imitation Learning via Multi-Step Occupancy Measure Matching},
author={Liu, Minghuan and Wang, Hangyu and Zhang, Yangtian and Xu, Minkai and Zhu, Zhengbang and Zhang, Weinan},
journal={Preprint},
year={2022}
}
Reviewer for NeurIPS, ICML, ICLR, and AAAI.
Continuous normalizing flows, conditional flow matching, and how they relate to diffusion.