saurabh atreya
Illustrated portrait of Saurabh Atreya holding a cat

saurabh atreya

i'm an ml engineer at Sarvam AI, where i work on gpu infrastructure and the inference stack, a published computer vision researcher, working on biometrics, self-supervised learning and video understanding, a BITS Pilani, Hyderabad graduate, with a bachelor's in computer science, based in bengaluru, india, where outside of work, i am exploring football, manga, improv and aviation.

01 papers

5 peer-reviewed papers. select one to read the abstract, or find the same list on google scholar .

Beyond Real versus Fake: Towards Intent-Aware Video Analysis IEEE T-BIOM 2026

The rapid advancement of generative models has led to increasingly realistic deepfake videos, posing significant societal and security risks. While existing detection methods focus on distinguishing real from fake videos, such approaches fail to address a fundamental question: What is the intent behind a manipulated video? Towards addressing this question, we introduce IntentHQ: a new benchmark for human-centered intent analysis, shifting the paradigm from authenticity verification to contextual understanding of videos. IntentHQ consists of 5168 videos that have been meticulously collected and annotated with 23 fine-grained intent-categories, including "Financial fraud", "Indirect marketing", "Political propaganda", as well as "Fear mongering". We perform intent recognition with supervised and self-supervised multi-modality models that integrate spatio-temporal video features, audio processing, and text analysis to infer underlying motivations and goals behind videos. Our proposed model is streamlined to differentiate between a wide range of intent-categories.

ViM-Disparity: Bridging the Gap of Speed, Accuracy and Memory for Disparity Map Generation ICASSP 2025

In this work we propose a Visual Mamba (ViM) based architecture, to dissolve the existing trade-off for real-time and accurate model with low computation overhead for disparity map generation (DMG). Moreover, we proposed a performance measure that can jointly evaluate the inference speed, computation overhead and the accurateness of a DMG model.

KDC-MAE: Knowledge Distilled Contrastive Masked Auto-Encoder WACV 2025

In this work, we attempted to extend the thought and showcase a way forward for the Self-supervised Learning (SSL) learning paradigm by combining contrastive learning, self-distillation (knowledge distillation) and masked data modelling, the three major SSL frameworks, to learn a joint and coordinated representation. The proposed technique of SSL learns by the collaborative power of different learning objectives of SSL. Hence to jointly learn the different SSL objectives we proposed a new SSL architecture KDC-MAE, a complementary masking strategy to learn the modular correspondence, and a weighted way to combine them coordinately. Experimental results conclude that the contrastive masking correspondence along with the KD learning objective has lent a hand to performing better learning for multiple modalities over multiple tasks.

Enhancing 3D-Air Signature by Pen Tip Tail Trajectory Awareness: Dataset and Featuring by Novel Spatio-temporal CNN IJCB 2023

This work proposes a novel process of using pen tip and tail 3D trajectory for air signature. To acquire the trajectories we developed a new pen tool and a stereo camera was used. We proposed SliT-CNN, a novel 2D spatial-temporal convolutional neural network (CNN) for better featuring of the air signature. In addition, we also collected an air signature dataset from 45 signers. Skilled forgery signatures per user are also collected. A detailed benchmarking of the proposed dataset using existing techniques and proposed CNN on existing and proposed dataset exhibit the effectiveness of our methodology.

Sclera Segmentation and Joint Recognition Benchmarking Competition: SSRBC 2023 IJCB 2023

This paper presents the summary of the Sclera Segmentation and Joint Recognition Benchmarking Competition (SSRBC 2023) held in conjunction with IEEE International Joint Conference on Biometrics (IJCB 2023). Different from the previous editions of the competition, SSRBC 2023 not only explored the performance of the latest and most advanced sclera segmentation models, but also studied the impact of segmentation quality on recognition performance. Five groups took part in SSRBC 2023 and submitted a total of six segmentation models and one recognition technique for scoring. The submitted solutions included a wide variety of conceptually diverse deep-learning models and were rigorously tested on three publicly available datasets, i.e., MASD, SBVPI and MOBIUS. Most of the segmentation models achieved encouraging segmentation and recognition performance. Most importantly, we observed that better segmentation results always translate into better verification performance.

02 patents

Three-Dimensional Air Signature System

A biometric system that captures a signature drawn freehand in mid-air, tracking both the tip and tail of a purpose-built pen in three dimensions to produce a signature that is far harder to forge than its two-dimensional equivalent.

filed
july 2024
published
january 2026
office
Indian Patent Office 

03 history

  1. ml engineer Sarvam AI · bengaluru, india
  2. ml research intern Inria · sophia antipolis, france
  3. b.e. computer science BITS Pilani · hyderabad, india

04 news

  1. my bachelor's thesis paper “Beyond Real versus Fake…” was accepted at IEEE T-BIOM
  2. our patent “Three-Dimensional Air Signature System” was published at the Indian Patent Office
  3. started working at Sarvam AI as an ML engineer
  4. graduated with a B.E. in computer science from BITS Hyderabad
  5. started working at Inria France as a research intern
  6. our paper “ViM-Disparity…” was accepted at ICASSP 2025
  7. our paper “KDC-MAE…” was accepted at WACV 2025
  8. our paper “Sclera Segmentation…” was accepted at IJCB 2023
  9. our paper “Enhancing 3D-Air Signature…” was accepted at IJCB 2023
  10. started a 4-year bachelor's degree in CS at BITS Hyderabad

05 contact

email is the best way to reach me. i'm constantly involved in projects across computer vision, ai infra and security, so if you're working on something exciting, i'd love to hear about it.