Skip to content
Saurabh Gupta

Fast track

30-second brief

Role fit

Applied AI roles that combine modelling depth, evaluation discipline, and production ownership.

Open to Senior Data Scientist, Machine Learning Engineer, AI / GenAI Engineer, or MLOps Engineer.

What I deliver

Problem framing, modelling and retrieval, evaluation, deployment, and operational ownership across the ML lifecycle.

Three reasons to continue

  1. ~5 years’ experience spans generative AI, computer vision, predictive modelling, and the data and cloud infrastructure around them.

  2. Production reliability is treated as part of the model: validation, failure paths, infrastructure parity, and operational ownership are designed in from the start.

  3. Evaluation evidence stays tied to intended use, while project evidence makes failure behavior, provenance, and deployment constraints explicit.

Proof paths

Data Science Engineer at Dodge Construction Network

Saurabh Gupta

I build AI systems that get measured, not just demoed.

shipping ML to production
~5 yearsshipping ML to production
in 5+ ML hackathons
Top 5in 5+ ML hackathons
companies across AI & data
3companies across AI & data
professional awards
2professional awards
Portrait of Saurabh Gupta

About

I am a Senior Data Scientist and Applied AI Engineer with ~5 years building and deploying production-grade machine learning and generative AI systems across NLP, computer vision, predictive analytics, and the cloud infrastructure that carries them.

Most of my work is end-to-end ownership of the ML lifecycle: designing custom RAG architectures and deep learning models, then getting them past the notebook and into scalable services on AWS. I build with rigorous data validation, deterministic dependency management, and containerized deployment, which is usually what separates a model that works from a model that keeps working.

The through-line is measurement. I use evaluation evidence to assess whether an AI system is meeting its intended use.

Domains I have worked in

Construction data
LLM-based company trade classification, residential-versus-commercial permit modelling, and real-time scoring services over commercial construction and building-permit records.
Supply chain
Data quality automation and partner-data ingestion pipelines for a global beverage manufacturer, including schema validation, reconciliation and anomaly detection across 20+ datasets under delivery SLAs agreed with external stakeholders.
Mobility & parking
Vehicle re-identification from entry and exit imagery, enforcement routing and illegal-parking revenue prediction across urban parking networks in Norway, Denmark, Germany, and Sweden.
Master data management
Entity resolution, crosswalk-safe remediation and large-scale data quality automation in Reltio, recovering missing organization records and purging invalid contacts without losing cross-source lineage.

Open to

Senior Data Scientist, Machine Learning Engineer, AI / GenAI Engineer, MLOps Engineer, currently based in Bengaluru, Karnataka; open to relocating worldwide.

Skills & tooling

PRODUCTION AIframe → model → evaluate → ship
5 branches · 43 capabilities
Core competencies
  • Generative & Agentic AI
  • LLM Orchestration
  • Retrieval-Augmented Generation
  • MLOps
  • Natural Language Processing
  • Computer Vision
Languages
  • Python
  • SQL
  • Shell scripting
Libraries & frameworks
  • LangChain
  • LangGraph
  • Ollama
  • FastAPI
  • Pydantic
  • PyTorch
  • TensorFlow
  • Scikit-Learn
  • Hugging Face
Cloud & infrastructure
  • AWS ECS
  • AWS Fargate
  • SageMaker
  • Bedrock
  • AWS Batch
  • Glue
  • ECR
  • RDS
  • Docker
  • Terraform
  • GitHub Actions
  • MLflow
  • DVC
  • PostgreSQL
  • Alembic
Machine & deep learning
  • Linear & Logistic Regression
  • Decision Trees
  • Random Forest
  • XGBoost
  • K-Means
  • CNNs
  • Siamese Networks
  • Transfer Learning
  • TF-IDF
  • BM25

Experience

Dodge Construction Network

Bengaluru, Karnataka

full time

Data Science Engineer

Jul 2025 - Present

1 yr 3 mos

Role scope

Owned applied AI delivery across construction-data products, from problem framing and model development to production deployment, data enrichment, and operational reliability.

Selected work

2 of 6

  1. Trade classification engine

    Generative AI

    Engineered an LLM-powered company trade classification pipeline on LangChain, LangGraph and AWS Bedrock.

    classification accuracy
    85%classification accuracy
    less manual processing
    70%less manual processing
  2. Production AI services

    AI Platform

    Built and deployed real-time AI backend services on FastAPI, Docker and AWS ECS/Fargate, with rate limiting, elastic auto-scaling, strict Pydantic validation and Alembic-managed database migrations.

Systems & tools

  • LangChain
  • LangGraph
  • AWS Bedrock
  • FastAPI
  • Pydantic
  • Docker
  • AWS ECS
  • Fargate
  • AWS Batch
  • Terraform
  • MLflow
  • DVC
  • PostgreSQL
  • Alembic
  • Reltio MDM

Sigmoid Analytics

Bengaluru, Karnataka

full time

Senior Data Scientist

May 2025 - Jul 2025

2 mos

Role scope

Delivered data-quality and ingestion solutions for partner datasets, translating reliability requirements into repeatable checks, workflows, and stakeholder-ready insights.

Selected work

2 entries

  1. Data-quality automation

    Data Quality

    Automated data quality workflows including schema validation, null checks, reconciliation, and anomaly detection, and established delivery SLAs with external stakeholders.

    datasets covered
    20+datasets covered
  2. Partner ingestion validation

    Data Engineering

    Developed scalable validation pipelines for partner data ingestion, improving reliability and refresh consistency.

    less manual effort
    40%less manual effort

Systems & tools

  • Python
  • SQL
  • Data Validation
  • Anomaly Detection

Get My Parking

Bengaluru, Karnataka

full time

Associate Data Scientist

Jan 2022 - May 2025

3 yrs 4 mos

Role scope

Built data products for parking operations, combining vehicle intelligence, geospatial analysis, and predictive modelling to improve field decisions and platform reliability.

Selected work

2 of 10

  1. Vehicle image matching

    Computer Vision

    Built an AI-powered vehicle image similarity system to match vehicles at entry and exit, mitigating errors from License Plate Recognition cameras. Used a Siamese network with contrastive learning and transfer learning on a ResNet backbone.

    top-3 match accuracy
    87%top-3 match accuracy
  2. Low-light enhancement

    Computer Vision

    Built a low-light image enhancement pipeline with Zero-DCE to recover vehicle details in challenging illumination and improve the quality of images passed to the similarity model.

Data Science Intern

Jul 2021 - Jan 2022

6 mos

Role scope

Supported digital route-planning research by preparing external and internal data, testing assumptions, and turning exploratory analysis into actionable model inputs.

Selected work

3 entries

  1. Route-planning research

    Applied Research

    Worked with the team testing Digital Route Planning (DRP), covering the full project pipeline from collection to analysis.

  2. External-data acquisition

    Data Engineering

    Wrote web scraping scripts to collect weather and event data using requests, Scrapy and BeautifulSoup, and extracted raw data from MySQL servers.

  3. Parking-pattern validation

    Statistical Analysis

    Performed hypothesis testing to validate whether weather and event conditions affect city parking patterns, helping the team assess model performance.

Systems & tools

  • PyTorch
  • Siamese Networks
  • ResNet
  • YOLOv8
  • Zero-DCE
  • XGBoost
  • Scikit-Learn
  • KMeans
  • Python
  • MySQL
  • Scrapy
  • Power BI

Projects

Things I built outside of work, with evaluation details where available. Source links are included where available.

  • Generative AI

    Vernacular RAG Voicebot for Government Schemes

    A public-prototype assistant for government-scheme questions in Hindi and Bengali voice or English, Hindi, and Bengali text, with a grounded RAG workflow and local-by-default inference.

    English context recall · offline
    0.97English context recall · offline
    English context precision · offline
    0.91English context precision · offline
    English faithfulness · answered rows
    0.996English faithfulness · answered rows
    • LangGraph
    • Ollama
    • Gemma
    • AI4Bharat Conformer
    • Parler-TTS
    • BM25
    • Cross-Encoder Reranking
    • RAGAS
    • Docker
    • FastAPI
    Local multilingual RAG voicebotText input in English, Hindi, or Bengali and voice input in Hindi or Bengali converge on hybrid retrieval. After reranking, a retrieval-confidence gate runs before Gemma. Gemma generates from accepted context or refuses; a groundedness verifier after generation routes failed verification to fallback before the localized response.LOCAL / MULTILINGUAL RAGLOCAL STACKVOICE HI / BNTEXT EN / HI / BNHYBRID RETRIEVALBM25DENSEFSCHEME INDEX + FUSIONX-ENCODER123RERANKGATEGEMMAGROUND / REFUSEVERIFYOR FALLBACKANSWER OR FALLBACKlocalized text · audio only for voice turnsTEXT/ AUDIOINPUT → RETRIEVE → GATE → GENERATE / REFUSE → VERIFY / FALLBACK
  • Machine learning

    MSME Financial Health Card

    An explainable credit-workflow prototype scored on synthetic snapshots of alternative-data fields, keeping missing repayment history explicit while scoring the remaining evidence.

    eligibility ROC-AUC · synthetic holdout
    0.978eligibility ROC-AUC · synthetic holdout
    generated-PD adj. R² · synthetic holdout
    0.905generated-PD adj. R² · synthetic holdout
    approval spread · synthetic portfolio
    2.4 ptsapproval spread · synthetic portfolio
    • Python
    • Scikit-Learn
    • XGBoost
    • FastAPI
    • Pandas
    • Credit Risk Modeling
    MSME financial health card credit modelSynthetic snapshots of GST, UPI, banking, and payroll fields flow through health scoring, risk models, anomaly checks, and prototype policy rules to produce a neutral decision card. The workflow is a prototype and its policy thresholds are unvalidated for real underwriting.CREDIT HEALTH / PROTOTYPEUNVALIDATEDGSTUPIBANKPAYSCORE / RISK / RULESDECISION CARDOUTPUT FIELDSHEALTH SCORERISK BANDREASON CODESVERSIONSPROTOTYPE BOUNDARYSYNTHETIC SNAPSHOT · POLICY THRESHOLDS UNVALIDATEDSYNTHETIC INPUT → SCORE / RISK / RULES → DECISION CARD

Credentials

Honors & awards

  • 2024

    Value Evangelist - Grow Together Award

    Get My Parking

    For the deployment and enhancement of Digital Route Planning applications across Denmark, Norway and Sweden.

  • 2023

    Go Getter Award

    Get My Parking

    For automating reporting workflows and generating operational insights for the business.

Certifications

  • DeepLearning.AI

    Neural Networks and Deep Learning

  • Coding Blocks

    Machine Learning Masters

  • CampusX

    Docker for Machine Learning

  • CampusX

    GenAI using Ollama

  • Hackerrank

    HackerRank Advanced SQL

Competitions

  • 2025

    Rank 1 of 1600+ participants

    HackerEarth Machine Learning Challenge

    First place across the full participant field in a competitive machine learning challenge.

  • 2024

    1st Runner Up amongst 12+ teams

    Get My Parking

    For creating a car and number plate matching image similarity model by combining SOTA computer vision techniques like CLIP (Contrastive Language-Image Pre-training) by OpenAI and fuzzy matching methods like Levenshtein distance to accelerate the review case resolution for ExpressAI product.

  • 2021

    Ranked 4/342 participants

    Dockship

    Fourth place across the full participant field in a competitive machine learning challenge.

  • 2020

    Ranked 20/2300+ participants

    HackerEarth

    Achieved 20th place across the full participant field in a competitive machine learning challenge.

Education

  • 2017-2021

    Lovely Professional University

    Bachelor of Technology (Hons.) · Computer Science Engineering

    CGPA 8.36/10

    Punjab, India

Get in touch

Open to thoughtful collaborations

Have a hard AI problem worth measuring?

If you are taking an ML system beyond the notebook, tightening evaluation around a GenAI product, or exploring a technical collaboration, I would be glad to hear what you are working on.

Best place to start

Put the problem on the table.

A short note about the context, the constraint, and what a useful outcome looks like is plenty. Email is the most reliable way to reach me.

Direct email

Start a conversationBengaluru, Karnataka · open to working across time zones

Good reasons to reach out

  1. 01

    Build something together

    Production AI, retrieval systems, evaluation pipelines, or an applied ML idea with a real user and a clear measure of success.

  2. 02

    Pressure-test an idea

    Compare architecture choices, experiment design, or how to tell whether a model is actually improving the product around it.

  3. 03

    Share the work

    Open-source experiments, hackathons, technical writing, or a thoughtful conversation with people working on similar problems.