PHD CANDIDATE · MULTIMODAL INTELLIGENCE

Teaching machines to see structure beyond pixels.

I’m Manav Barot, a Ph.D. candidate at the State University of New York at Binghamton. I build vision-language systems that connect visual evidence with geometry, language, and human-centered reasoning.

FOCUSVLM · CV · 3D
BASEBINGHAMTON, NY
STATUSRESEARCH ACTIVE
00 / THESIS

Recognition is not enough. I’m interested in models that can observe, measure, compare, and then explain their reasoning in language.

RESEARCH VECTOR

Three layers of the same problem.

Visual intelligence becomes more useful when perception, geometry, and language are designed to reinforce one another.

CURRENT SYSTEM

PoseVLM

Geometry-aware multimodal reasoning for fine-grained human-pose understanding.

OPEN PROJECT INDEX ↗
01Visual Inputhuman pose imagery
023D Geometrybody structure + measurements
03Multimodal Modelvision + language reasoning
04Grounded Outputexplanation + correction
SELECTED SYSTEMS

Research that compiles.

RESEARCH NOTEBOOK

Ideas in motion.

OPEN WRITING ↗
END OF FORWARD PASS

See the work.
Interrogate the ideas.