Teaching machines to see structure beyond pixels.
I’m Manav Barot, a Ph.D. candidate at the State University of New York at Binghamton. I build vision-language systems that connect visual evidence with geometry, language, and human-centered reasoning.
Recognition is not enough. I’m interested in models that can observe, measure, compare, and then explain their reasoning in language.
Three layers of the same problem.
Visual intelligence becomes more useful when perception, geometry, and language are designed to reinforce one another.
Grounded VLM Reasoning
Fine-grained visual understanding that ties generated language back to observable evidence.
Human Geometry
Joint relationships, symmetry, angles, orientation, and spatial structure as reasoning signals.
Computer Vision
Representation learning, segmentation, fine-grained classification, and visual systems.
PoseVLM
Geometry-aware multimodal reasoning for fine-grained human-pose understanding.
OPEN PROJECT INDEX ↗