Research
My research broadly spans the field of computer vision, with specific interests including generative AI, diffusion models, action recognition, person recognition, and representation learning. You can check out my selected papers below, with important papers highlighted.
|
|
|
FlashSSM-3D: A Distilled State Space Model for Lightning-Fast Dense 3D Reconstruction
Nyle Siddiqui,
Mubarak Shah.
Conference on Neural Information Processing Systems (NeurIPS), 2026
We introduce a geometry-aware knowledge distillation framework for efficiently training a large-scale state space model for dense 3D reconstruction, achieving comparable reconstruction accuracy with substantially faster inference.
|
|
|
CaMBRAIN: Real-time, Continuous EEG Inference with Causal State Space Models
Abhilash Durgam,
Nyle Siddiqui,
Jeffrey A. Chan-Santiago,
Qiushi Fu,
Elakkat D. Gireesh,
Mubarak Shah.
Under Review
paper
We investigate the use of SSMs for continuous, real-time EEG brain signal decoding.
|
|
|
VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
Sirnam Swetha,
Rohit Gupta,
Parth Parag Kulkarni,
David G Shatwell,
Jeffrey A Chan Santiago,
Nyle Siddiqui,
Joseph Fioresi,
Mubarak Shah.
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
(Highlight Paper)
paper  /  code  /  project page
We introduce VRR-QA, a benchmark for visual relational reasoning in videos that tests whether models can infer relationships, actions, and contextual cues that are not explicitly shown. The benchmark contains 1K expert-annotated question-answer pairs across 1K creative video clips and reveals a substantial gap between current VideoQA models and human performance on implicit reasoning.
|
|
|
DVANet: Disentangling View and Action Features for Multi-View Action Recognition
Nyle Siddiqui,
Praveen Tirupattur,
Mubarak Shah.
AAAI Conference on Artificial Intelligence, Main Technical Track (AAAI), 2024
paper  /  code  /  project page
We propose a novel transformer decoder-based architecture in tandem with two supervised contrastive losses for multi-view action recognition. By disentangling the view-relevant features from action-relevant features, we enable our model to learn action features that are robust to change in viewpoints. We show that changes in viewpoint impart perturbations on learned action features, and thus, disentangling these perturbations improves overall action recognition performance.
|
|
|
DLCR: Leveraging Diffusion and Large Language Models for Effective Clothes-Changing Re-Identification
Nyle Siddiqui,
Alin Croitoru,
Gaurav Kumar Nayak,
Radu Tudor Ionescu,
Mubarak Shah.
IEEE/CVF Winter Conference on Applications of Computer Vision, Main Algorithms Track (WACV), 2025
(Selected for Oral Presentation!)
We propose a generative data expansion framework via diffusion for clothes-changing person Re-ID, which leverages pre-trained diffusion models and large language models to accurately generate images of individuals with different clothing attires. We address the challenges faced by CC-ReID models due to the limited clothing diversity in current CC-ReID datasets by genereating additional synthetic data that increases clothing diversity while preserving important personal features in the generated images. We also introduce two novel CC-ReID training strategies: progressive learning and test-time prediction refinement. Notably, training certain models with data generated by DLCR on the PRCC dataset resulted in improvements of up to 11.3% improvement in top-1 accuracy, with additonal enhanced performance on out-of-distribution test data.
|
|
|
Beyond Fixed Frames: Flexible SSM Training Unlocks Video Understanding Across Spatio-Temporal Scales
Nyle Siddiqui,
Rohit Gupta,
Swetha Sirnam,
Mubarak Shah.
Under Review
Currently under review. We instill VideoMamba with spatio-temporal flexibility and shows it performs better on a variety of action recognition tasks.
|
|
|
Machine and Deep Learning Applications to Mouse Dynamics for Continuous User Authentication
Nyle Siddiqui,
Rushit Dave,
Mounika Vanamala,
Naeem Seliya.
MDPI Journal of Machine Learning and Knowledge Extraction (MAKE), 2022
(4th most cited MAKE paper in 2022)
paper    
|
|
|
Applied Science PhD Intern - Amazon Ring AI
May 2026 - November 2026
Conducted research on video object detection and grounding using state space models for Ring cameras. Paper submitted to CVPR 2027.
|
|
|
Computer Vision PhD Intern - Metalenz
May 2025 - September 2025
Conducted novel research applying vision-language models to polarimetric data for facial anti-spoofing.
|
 |
ORCGS Doctoral Fellowship, 2022-2026
|
Professional Reviewing Experience
|
Feel free to steal this website's source code. Do not scrape the HTML from this page itself, as it includes analytics tags that you do not want on your own website — use the github code instead. Also, consider using Leonid Keselman's Jekyll fork of this page.
|
|