Portrait of MD Khalequzzaman Chowdhury Sayem

MD Khalequzzaman Chowdhury Sayem

Hi there, I'm Sayem! I work on 3D-aware vision and multimodal models that try to understand people: how their hands and bodies move, how they interact with objects, and how they behave over time.

Before this, I was fortunate to work with Prof. Seungryul Baek and Prof. Binod Bhattarai at UNIST, where I completed my M.S. I'm always happy to chat about research or possible collaborations.

News

Older news
  • Sep 2025Joined the UNIST Vision & Learning Lab as a Researcher.
  • Aug 2025Defended my master's thesis and graduated from UNIST.
  • Dec 2024QORT-Former was accepted to AAAI 2025.

Research

Geometry-Grounded Multimodal Reasoning

Using 3D geometry to find and fix spatial-reasoning failures in vision-language models.

HandVQA · CVPR 2026

Real-Time 3D Perception

Fast 3D understanding, from hand-object pose to stereo depth.

QORT-Former · AAAI 2025 / LIOPS

3D Humans & Avatars

Animatable 3D Gaussian avatars from monocular video, even with partial body visibility.

FlexiAvatar · ECCV 2026

Looking ahead: I'm interested in human behavior foundation models that learn how people act and interact over time.

Publications

* equal contribution. Full list on Google Scholar.

HandVQA teaser image

HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models

MD Khalequzzaman Chowdhury Sayem*, Mubarrat Tajoar Chowdhury*, Yihalem Yimolal Tiruneh, Muneeb A. Khan, Muhammad Salman Ali, Binod Bhattarai, Seungryul Baek

CVPR 2026Denver, CO, USADataset released

A benchmark grounded in 3D hand geometry. We found that VLMs struggle with fine-grained spatial reasoning about hands, and that explicit 3D supervision helps considerably.

Abstract

We introduce HandVQA, a large-scale benchmark with 1.6M+ geometry-derived VQA pairs spanning joint angles, distances, and relative spatial relations (X/Y/Z). We find that explicit 3D supervision improves spatial reasoning and also transfers to gesture recognition and hand-object interaction tasks.

BibTeX
@inproceedings{sayem2026handvqa,
  title     = {HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models},
  author    = {Sayem, MD Khalequzzaman Chowdhury and Chowdhury, Mubarrat Tajoar and Tiruneh, Yihalem Yimolal and Khan, Muneeb A. and Ali, Muhammad Salman and Bhattarai, Binod and Baek, Seungryul},
  booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  year      = {2026}
}
FlexiAvatar results across full-body, upper-body, and head-only inputs

FlexiAvatar: Unified 3D Gaussian Human Avatars Under Arbitrary Body Visibility

Yihalem Yimolal Tiruneh, Muhammad Salman Ali, Uyoung Jeong, Muneeb A. Khan, MD Khalequzzaman Chowdhury Sayem, Allanur Bayramgeldiyev, Binod Bhattarai, Seungryul Baek

ECCV 2026Malmö, Sweden

A single avatar framework that works with full-body, upper-body, or head-only monocular video.

Abstract

FlexiAvatar is a visibility-aware framework for reconstructing animatable 3D Gaussian human avatars. It optimizes only visible body regions, combines occlusion-robust SMPL-X tracking with part-specific refinement and diffusion-based completion, and improves reconstruction quality while reducing runtime and memory overhead in partial-visibility settings.

BibTeX
@inproceedings{tiruneh2026flexiavatar,
  title     = {FlexiAvatar: Unified 3D Gaussian Human Avatars Under Arbitrary Body Visibility},
  author    = {Tiruneh, Yihalem Yimolal and Ali, Muhammad Salman and Jeong, Uyoung and Khan, Muneeb A. and Sayem, MD Khalequzzaman Chowdhury and Bayramgeldiyev, Allanur and Bhattarai, Binod and Baek, Seungryul},
  booktitle = {Proceedings of the European Conference on Computer Vision (ECCV)},
  year      = {2026}
}
QORT-Former teaser image

QORT-Former: Query-Optimized Real-Time Transformer for Understanding Two Hands Manipulating Objects

Elkhan Ismayilzada*, MD Khalequzzaman Chowdhury Sayem*, Yihalem Yimolal Tiruneh, Mubarrat Tajoar Chowdhury, Muhammadjon Boboev, Seungryul Baek

AAAI 2025Philadelphia, PA, USACo-first author

A real-time Transformer for estimating the 3D pose of two hands and a manipulated object, running at 53.5 FPS.

Abstract

QORT-Former optimizes queries to balance efficiency and accuracy, leveraging hand-object contact information and a three-step feature update mechanism. It runs at 53.5 FPS on an RTX 3090 Ti and improves over prior methods on the H2O and FPHA datasets.

BibTeX
@inproceedings{ismayilzada2025qortformer,
  title     = {QORT-Former: Query-Optimized Real-Time Transformer for Understanding Two Hands Manipulating Objects},
  author    = {Ismayilzada, Elkhan and Sayem, MD Khalequzzaman Chowdhury and Tiruneh, Yihalem Yimolal and Chowdhury, Mubarrat Tajoar and Boboev, Muhammadjon and Baek, Seungryul},
  booktitle = {Proceedings of the AAAI Conference on Artificial Intelligence},
  year      = {2025}
}

Experience & Education