Our wedding by the Raritan River

Zhenyu (Zach) Wang 王振宇

Ph.D. Candidate in Statistics · Rutgers University

I am a Ph.D. candidate in Statistics at Rutgers University, advised by Zijian Guo and Yifan Hu.

My research spans statistical learning and large language models. My earlier work focuses on distributionally robust learning, causal invariance, and statistical inference. More recently, I have been drawing on ideas from statistics, optimization, and operations research to understand and improve how large language models learn and reason, with a focus on post-training and budget-aware LLM inference.

Before Rutgers, I received my master’s degree from Columbia University and my bachelor’s degree in Finance from Shanghai Jiao Tong University.

Portrait of Zhenyu (Zach) Wang

Recent news

Selected research

All research
Dr. OPD: do not just follow the teacher; learn which token-level teacher signals help make the student better.

LLM learning

Dr. OPD

Which teacher signals make a better student?

OPD provides dense token-level supervision. We formulate optimal OPD as bilevel optimization, learning how to weight teacher signals to improve the student’s performance.

Zhenyu Wang*, Tianze Wang*, Linjun Zhang, Yifan Hu

Preprint · 2026

The Pareto frontier of LLM test-time compute: the best attainable accuracy at each token budget; the figure is illustrative.

LLM reasoning

PRICE

NeurIPS 2026 Workshop · Oral

Also accepted at NeurIPS 2026 workshops: MLxOR · RAAAI · TTCL

How much accuracy can one more token buy?

We derive a closed-form asymptotic cost–accuracy Pareto frontier for LLM test-time compute, then use it to guide adaptive rollouts and voting under a token budget.

Zhenyu Wang, Xiaozhi Zhu, Yifan Hu

Preprint · 2026

Geometric illustration of DRoL: source models, their convex hull, and the robust solution with and without a model-class constraint.

Robust learning

DRoL

The Annals of Statistics · 2026

A leading journal in statistical theory and methodology.

How do we find the most robust model across data sources?

We identify the model that maximizes worst-case predictive performance over mixtures of source populations, with no target labels required.

Zhenyu Wang, Peter Bühlmann, Zijian Guo

The Annals of Statistics · 2026 · 54(2), 570–596