Joy Hsu

Joy Hsu

I am a Member of Technical Staff at OpenAI and an incoming Assistant Professor in the Department of Computing at Imperial College. Previously, I completed my PhD in Computer Science at Stanford University as a Knight-Hennessy Scholar and NSF Fellow, advised by Jiajun Wu. I also received my BS and MS from Stanford University.

You can reach me at joy.cj.hsu@gmail.com.

Research

My research focuses on multimodal reasoning models & agentic systems: machines that can perceive, interpret, and interact safely with their environments. I am particularly interested in building models that generalize across diverse domains and long-horizon tasks, and do so robustly and reliably in the real world.

Publications

PhD thesis thumbnail

Toward Generalizable and Intelligent Visual Reasoning Systems

Joy Hsu

PhD Thesis / Paper

APIVOT paper thumbnail

APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts

Emily Jin, Joy Hsu, Yiqing Xu, Weiyu Liu†, Nick Haber†, and Jiajun Wu†

NeurIPS 2026 / Paper / Project Page

Learning Situated Awareness paper thumbnail

Learning Situated Awareness in the Real World

Chuhan Li, Ruilin Han*, Joy Hsu*, Yongyuan Liang*, Rajiv Dhawan, Jiajun Wu, Ming-Hsuan Yang, and Xin Eric Wang

ICML 2026 / Paper / Project Page

Spotlight Presentation

Tool Bottleneck paper thumbnail

A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding

Christina Liu*, Alan Wang*, Joy Hsu, Jiajun Wu, and Ehsan Adeli

MIDL 2026 / Paper / Project Page

Neuro-Symbolic Decoding of Neural Activity paper thumbnail

Neuro-Symbolic Decoding of Neural Activity

Yanchen Wang*, Joy Hsu*, Ehsan Adeli†, and Jiajun Wu†

ICLR 2026 / Paper / Project Page

Hybrid world representations paper thumbnail

Discovering Hybrid World Representations with Co-Evolving Foundation Models

Jiajun Wu, Yunzhi Zhang, Hong-Xing Yu, Joy Hsu, and Jiayuan Mao

AAAI 2026 / Paper

Factored real-world scene generation paper thumbnail

From Programs to Poses: Factored Real-World Scene Generation via Learned Program Libraries

Joy Hsu, Emily Jin, Jiajun Wu, and Niloy J. Mitra

NeurIPS 2025 / Paper / Project Page

Maze abstractions paper thumbnail

What Makes a Maze Look Like a Maze?

Joy Hsu, Jiayuan Mao, Joshua B. Tenenbaum, Noah D. Goodman, and Jiajun Wu

ICLR 2025 / Paper / Project Page

Predicate Hierarchies paper thumbnail

Predicate Hierarchies Improve Few-Shot State Classification

Emily Jin*, Joy Hsu*, and Jiajun Wu

ICLR 2025 / Paper / Project Page

VDLM paper thumbnail

Visually Descriptive Language Model for Vector Graphics Reasoning

Zhenhailong Wang, Joy Hsu, Xingyao Wang, Kuan-Hao Huang, Manling Li, Jiajun Wu, and Heng Ji

TMLR / Paper / Project Page

CVPR Workshop On Multimodal Algorithmic Reasoning Spotlight Presentation

LARC paper thumbnail

Naturally Supervised 3D Visual Grounding with Language-Regularized Concept Learners

Chun Feng*, Joy Hsu*, Weiyu Liu, and Jiajun Wu

CVPR 2024 / Paper

Learning Planning Abstractions from Language paper thumbnail

Learning Planning Abstractions from Language

Weiyu Liu*, Geng Chen*, Joy Hsu, Jiayuan Mao†, and Jiajun Wu†

ICLR 2024 / Paper / Project Page

What's Left paper thumbnail

What’s Left? Concept Grounding with Logic-Enhanced Foundation Models

Joy Hsu*, Jiayuan Mao*, Joshua B. Tenenbaum, and Jiajun Wu

NeurIPS 2023 / Paper / Project Page

Visual Scratchpads paper thumbnail

Can Visual Scratchpads With Diagrammatic Abstractions Augment LLM Reasoning?

Joy Hsu, Gabriel Poesia, Jiajun Wu, and Noah D. Goodman

PMLR / Paper

NeurIPS ICBINB Workshop Best Poster Award

Composable Part-Based Manipulation paper thumbnail

Composable Part-Based Manipulation

Weiyu Liu, Jiayuan Mao, Joy Hsu, Tucker Hermans, Animesh Garg, and Jiajun Wu

CoRL 2023 / Paper / Project Page

Motion Question Answering paper thumbnail

Motion Question Answering via Modular Motion Programs

Mark Endo*, Joy Hsu*, Jiaman Li, and Jiajun Wu

ICML 2023 / Paper / Project Page

NS3D paper thumbnail

NS3D: Neuro-Symbolic Grounding of 3D Objects and Relations

Joy Hsu, Jiayuan Mao, and Jiajun Wu

CVPR 2023 / Paper / Project Page

CVPR Workshop On Compositional 3D Vision Oral Presentation

ProgramPort paper thumbnail

Programmatically Grounded, Compositionally Generalizable Robotic Manipulation

Renhao Wang*, Jiayuan Mao*, Joy Hsu, Hang Zhao, Jiajun Wu, and Yang Gao

ICLR 2023 / Paper / Project Page

Spotlight Presentation

DisCo paper thumbnail

DisCo: Improving Compositional Generalization in Visual Reasoning through Distribution Coverage

Joy Hsu, Jiayuan Mao, and Jiajun Wu

TMLR / Paper / Project Page

Geoclidean paper thumbnail

Geoclidean: Few-Shot Generalization in Euclidean Geometry

Joy Hsu, Jiajun Wu, and Noah D. Goodman

NeurIPS 2022 / Paper / Project Page

Hyperbolic representations paper thumbnail

Capturing Implicit Hierarchical Structure in 3D Biomedical Images with Self-Supervised Hyperbolic Representations

Joy Hsu*, Jeff Gu*, Gong-Her Wu, Wah Chiu, and Serena Yeung

NeurIPS 2021 / Paper / Project Page

NeurIPS Differential Geometry Workshop Contributed Talk

DARCNN paper thumbnail

DARCNN: Domain Adaptive Region-based Convolutional Neural Network for Unsupervised Instance Segmentation in Biomedical Images

Joy Hsu, Wah Chiu, and Serena Yeung

CVPR 2021 / Paper / Project Page

Stanford CS Honors Thesis thumbnail

Unsupervised Learning for Discovery in 2D & 3D Scenes: Towards Unbiased Understanding of Biomedical Images

Joy Hsu

Undergraduate Honors Thesis / Paper

Stanford Ben Wegbreit Prize for Best Thesis