Research

Multimodal systems grounded in vision and geometry.

I study how visual models can reason beyond labels by connecting images with language, pose structure, geometry, and domain knowledge.

01 / VLM

Vision-Language Models

My current work explores fine-grained visual-language reasoning: systems that can inspect visual evidence, describe it precisely, compare alternatives, and produce grounded natural-language feedback.

A major focus is human pose understanding, where language supervision becomes more useful when it is tied to geometric signals rather than image appearance alone.

multimodal learninggroundingVLMsinstruction tuning
02 / 3D

3D Human Understanding

I’m interested in representations that expose the geometry of a human pose — joint angles, symmetry, orientation, relative position, and body configuration — so models can reason about what a person is doing and how the pose differs from a target.

This direction connects 3D body estimation with multimodal reasoning and language generation.

3D posegeometryhuman understandingspatial reasoning
03 / CV

Computer Vision

My broader computer-vision work spans fine-grained classification, segmentation, representation learning, and applied visual intelligence. I’m especially interested in methods that stay interpretable enough to diagnose failure modes and useful enough to ship.

representation learningsegmentationclassificationvision systems
Research philosophy

Evidence before eloquence.

A multimodal model should not merely produce convincing text. The goal is to connect its language to measurable visual or geometric evidence, then evaluate where that connection succeeds or fails.