I work on pre-training at an AI startup. Previously, I co-founded Voyage AI and led its research to develop the best embeddings models and rerankers for semantic search and information retrieval in the industry.
I did my PhD at Stanford University, affiliated with Stanford AI Lab and the Stanford NLP Group.
My research interests broadly lie in synthetic data, multimodal learning, and reasoning.
News
- Jul 2026MoCa will appear at ACL 2026 as an Oral and SAC Highlight.
Language Models
- Synthetic Bootstrapped Pretraining ICLR 2026 x
- MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings ACL 2026OralSAC Highlight code x
- Chain of Thought Empowers Transformers to Solve Inherently Serial Problems ICLR 2024 x
- Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training ICLR 2024 code x
- Same Pre-training Loss, Better Downstream: Implicit Bias Matters for Language Models ICML 2023Oral code x
- Self-supervised Learning is More Robust to Dataset Imbalance ICLR 2022Spotlight code x
Domain Adaptation and Transfer Learning
- Cycle Self-Training for Domain Adaptation NeurIPS 2021 code
- Learning to Adapt to Evolving Domains NeurIPS 2020
- Meta-learning Transferable Representations with a Single Target Domain arXiv 2020
- Towards Understanding the Transferability of Deep Representations arXiv 2019
- Transferable Adversarial Training: A General Approach to Adapting Deep Classifiers ICML 2019Long Talk code
- Separate to Adapt: Open Set Domain Adaptation via Progressive Separation CVPR 2019 code