Ali Torabi

PhD Student in Computer Science

Spatial & Relational Reasoning in Vision-Language Models, 3D Scene Understanding, and Explainable AI

Ali Torabi

About Me

I am a PhD student in Computer Science at the University of Wyoming, working under the supervision of Dr. Yaqoob Majeed. My research focuses on developing novel techniques in Explainable AI (XAI) for Computer Vision, particularly in Weakly Supervised Semantic Segmentation (WSSS).

I'm passionate about bridging the gap between theoretical research and practical applications in deep learning. My work involves developing influence-guided Class Activation Mapping (CAM) techniques with instance-level refinement and GPU-optimized scaling for large-scale datasets such as PASCAL VOC and COCO.

My current work focuses on reasoning in LLMs and Vision-Language Models, especially relational reasoning and counting of 3D objects in complex scenes, across single images, multi-view 3D data, and video. I build controllable 3D benchmarks with exact ground truth, 3D scene-graph reasoners that lift depth and open-vocabulary perception into metric space, and R1-style reinforcement learning (GRPO with verifiable rewards) to improve the spatial reasoning of VLMs such as Qwen2.5-VL.

With a strong background in software engineering and industrial experience in iOS development with ML integration, I bring both academic rigor and practical expertise to my research. I have published in top venues including Springer Neural Processing Letters and have papers under review in leading journals.

Education

PhD in Computer Science
University of Wyoming

Research Focus

Spatial Reasoning in VLMs
Explainable AI

Recognition

NSF ART Grant
Multiple Publications

Research Interests

Spatial & Relational Reasoning in VLMs

Counting and relational reasoning about 3D objects in complex scenes, from single images to multi-view 3D data and video.

3D Scene Graphs & RL for Reasoning

Lifting depth and open-vocabulary perception into metric 3D scene graphs, and GRPO with verifiable rewards to teach VLMs to reason in 3D.

Weakly Supervised Semantic Segmentation

Developing novel CAM-based techniques with influence functions for improved segmentation with weak annotations.

Explainable AI for Computer Vision

Creating interpretable deep learning models with explanation methods to improve neural network predictions.

Vision Transformers & Deep Learning

Exploring transformer architectures for computer vision tasks and multi-modal learning applications.

GPU-Accelerated ML

Optimizing model training with CUDA, mixed-precision training, and multi-GPU parallelism for large-scale datasets.

Current Research

Counting and relational reasoning in 3D scenes with Vision-Language Models

Spatial Scene Bench renders, instance masks and depth

Spatial Scene Bench

Procedural 3D scenes with exact ground truth (depth, instance masks, camera poses, orbit videos) and occlusion-aware questions on counting, viewer-centric relations, metric distance and egocentric direction for evaluating VLMs.

3D Reasoning VLM Evaluation Video
3D scene graph reasoning

3D Scene-Graph Spatial Reasoner

Lifts open-vocabulary detection, SAM masks and metric depth into a 3D scene graph with amodal object completion and multi-view fusion, then answers counting and relational questions with a neuro-symbolic solver or graph-prompted LLM.

Grounding DINO Depth Anything V2 Scene Graphs
GRPO training for spatial reasoning

Spatial-R1: RL for 3D Reasoning in VLMs

R1-style GRPO with verifiable rewards (format, exact-count, MRA for metric answers) to improve counting and spatial reasoning of Qwen2.5-VL on images and multi-frame video, built on TRL with LoRA, Dr. GRPO and DAPO objectives.

GRPO Qwen2.5-VL TRL

Featured Projects

Computer Vision & Vision-Language Model Projects

Vision Transformer

Vision Transformer Classification

Complete PyTorch implementation of ViT with attention visualization and transfer learning capabilities.

PyTorch Transformers ViT
CLIP Search

CLIP Image Search Engine

Semantic image search using OpenAI CLIP for natural language queries with FAISS integration.

CLIP FAISS Gradio
Image Captioning

Image Captioning with VLMs

Automatic caption generation using BLIP, BLIP-2, and GIT vision-language models.

BLIP VLM Hugging Face
YOLO Detection

YOLO Object Detection

Real-time object detection with YOLOv8 supporting 80+ classes with video processing.

YOLOv8 Real-time OpenCV
Semantic Segmentation

Semantic Segmentation

Pixel-level classification using DeepLabV3 with 21-class PASCAL VOC segmentation.

DeepLab Segmentation PyTorch
Visual QA

Visual Question Answering

Answer questions about images using BLIP-2 and multi-modal reasoning.

VQA BLIP-2 Multi-modal

Publications

2025

Instance-Guided Class Activation Mapping for Weakly Supervised Semantic Segmentation

Ali Torabi, Yaqoob Majeed, Md. Mahbubur Rahman, Sanjog Gaihre

arXiv:2509.12496, 2025

2025

Integrating deep CNN models for multilingual Sign Language recognition: A SignLink-based approach for Bengali and English

Niamul Hassan Samin, Mustahidul Islam Ferdous, Renu Akter Suity, et al.

Research Square (Preprint), July 2025

2023

Using Cartesian Genetic Programming Approach with New Crossover Technique to Design Convolutional Neural Networks

Ali Torabi, Arash Sharifi, Mohammad Teshnehlab

Springer, Neural Processing Letters, 2023

Under Review

A New Crossover Technique for Cartesian Genetic Programming based on Population Alignment

Ali Torabi, Arash Sharifi, Mohammad Teshnehlab

International Journal of Bio-Inspired Computation

Under Review

A Web-Based Framework for Detecting Fake News in Monolingual Texts

Md. Abdur Rahman, Md. Hafizur Rahman Sumon, Shanta Islam, et al., Ali Torabi, Asif Alif

Online Social Networks and Media

Experience

Aug 2023 - Present

PhD Researcher

University of Wyoming

Conducting research on Explainable AI in Deep Learning and Image Recognition. Developing influence-guided CAM techniques for Weakly Supervised Semantic Segmentation with GPU-optimized scaling on PASCAL VOC and COCO datasets.

Summer 2025

Research Assistant (NSF ART Grant)

University of Wyoming

Designed and implemented a web-based 3D STL visualization platform with interactive cross-sectioning and advanced rendering using Three.js and WebGL for F3DT commercialization project.

Fall 2023 - Present

Teaching Assistant - Algorithm Design and Analysis

University of Wyoming

Served as TA for five consecutive semesters, providing mentorship in algorithmic problem solving and complexity analysis.

April 2020 - 2023

Senior iOS Developer (Freelancer)

Tehran, Iran

Developed iOS applications with ML integration. Built Deep CNN models for product image recognition and implemented ML functionality across iOS devices including watchOS.

Get In Touch

I'm always open to discussing research opportunities, collaborations, or potential positions in Computer Vision and Deep Learning.

Phone

+1 (949) 738-3001

Location

University of Wyoming
Laramie, WY, USA

GitHub

@selfishout