Skip to content
Saurabh Gupta

Fast track

30-second brief

Role fit

Applied AI roles that combine modelling depth, evaluation discipline, and production ownership.

Open to Senior Data Scientist, Machine Learning Engineer, AI / GenAI Engineer, or MLOps Engineer.

What I deliver

Problem framing, modelling and retrieval, evaluation, deployment, and operational ownership across the ML lifecycle.

Three reasons to continue

  1. ~5 years’ experience spans generative AI, computer vision, predictive modelling, and the data and cloud infrastructure around them.

  2. Production reliability is treated as part of the model: validation, failure paths, infrastructure parity, and operational ownership are designed in from the start.

  3. Evaluation evidence stays tied to intended use, while project evidence makes failure behavior, provenance, and deployment constraints explicit.

Proof paths

Vernacular RAG Voicebot for Government Schemes

A multilingual, stateful RAG voice-assistant prototype for government-scheme discovery, with explicit evidence gating and grounded generation.

Offline evaluation replays one fixed synthetic set of 100 semantic questions in English, Hindi, and Bengali: 100 questions per language and 300 language-runs in total.

Government-scheme discovery across text and speech

The prototype treats scheme discovery as a question-answering problem: a person can describe a need or eligibility situation without first knowing the scheme name or source document. English, Hindi, and Bengali text access keeps that search from depending on an English-only query.

Hindi and Bengali voice input serves the same path when typing is not the preferred mode, and vernacular speech is returned only for voice turns. Language choice is explicit because the IndicConformer speech service requires the caller to declare the spoken language rather than relying on automatic detection.

Product surface

Multilingual Voicebot
Starting API
voicebot-api127.0.0.1:8200

$ uvicorn src.api.app:app

INFO Loading multilingual-e5-large and bge-reranker-v2-m3

INFO FAISS + BM25 indexes loaded

200 GET /healthcheck  761 chunks ready

Recorded multi-turn API replay from the project evaluation artifacts. The follow-up resolves “it” from session memory, retrieves APY evidence again, and returns a second grounded response.

The public portfolio replays the run locally in the browser. It does not expose the private model server or API credentials.

Watch the RAG system work

Scrub the timeline to see source documents become chunks, ranks become evidence, and evidence become a response.

01 / 07Paused

Detect changes

Parse PDFs and text without flattening their structure

A SHA-256 manifest skips unchanged PDF and TXT files during incremental ingestion.

Source files become page-aware content with tables and reading order intact.

Retrieval and uncertainty are separate decisions

After normalization, multilingual-e5 semantic search covers paraphrase and intent while BM25 protects exact scheme names, eligibility terms, and distinctive wording. Reciprocal-rank fusion combines their ranked candidates without pretending their raw scores share one scale; bge-reranker-v2-m3 then evaluates the fused query-passage pairs more closely.

The reranked set passes only when its strongest score reaches 0.55 and leads the second result by at least 0.03. Weak evidence routes to a fallback; accepted evidence enters a grounded prompt and then a verifier that also fails closed. General conversation deliberately bypasses retrieval and rejoins the path at vernacular translation.

Conversation state and evaluation stay decoupled

LangGraph checkpoints retain conversation content and graph state under the session thread ID. A separate Alembic-managed chat_sessions table keeps the title, timestamps, and turn count needed to list conversations. The browser retains the active ID, loads the session list, and asks the replay endpoint for the selected history, so resume behavior does not make product metadata part of graph internals.

Evaluation stays outside the request path because the local judge and metric calls are a repeatable regression harness, not user-facing work. Retrieval precision and recall use all 100 runs per language because retrieval happens before gating; answer-quality metrics use only the 84 English, 79 Hindi, and 74 Bengali answered rows; answered rate uses all runs so refusal coverage remains visible.

Evidence lab

Evidence figure

Multilingual retrieval and answer quality

How consistent is the final RAG pipeline across the three languages?

Retrieval · all language-runs100 rows per language; retrieval runs before the evidence gate.

Context precision

English0.908
Hindi0.858
Bengali0.853

Context recall

English0.970
Hindi0.943
Bengali0.923
Answer quality · answered rows onlyEnglish n=84 · Hindi n=79 · Bengali n=74. Evidence-gated fallbacks are coverage outcomes, not answer-quality rows.

Faithfulness

English0.996
Hindi0.994
Bengali0.989

Answer relevancy

English0.892
Hindi0.889
Bengali0.887

Answer correctness

English0.738
Hindi0.747
Bengali0.738

Answer similarity

English0.951
Hindi0.947
Bengali0.952

Shared 0.0-1.0 scale; higher is better.

Interpretation note Retrieval facts use all 100 rows per language; answer-quality facts use only answered rows. Do not average them into one overall score.

Values and denominators
Retrieval metrics · all language-runs
MetricEnglish
n=100
Hindi
n=100
Bengali
n=100
Context precision0.9080.8580.853
Context recall0.9700.9430.923
Answer quality · answered rows only
MetricEnglish
n=84
Hindi
n=79
Bengali
n=74
Faithfulness0.9960.9940.989
Answer relevancy0.8920.8890.887
Answer correctness0.7380.7470.738
Answer similarity0.9510.9470.952
Evidence figure

Evidence-gated answer coverage

How often did the configured evidence gate permit a grounded answer?

All 100 runs per language · configured gate 0.70

English0.840
Hindi0.790
Bengali0.740

Read Answered rate moves from 0.84 in English to 0.79 in Hindi and 0.74 in Bengali while answered-row quality remains comparatively stable.

Values and denominators
Answered rate · all language-runs
MetricEnglish
n=100
Hindi
n=100
Bengali
n=100
Answered rate0.8400.7900.740

What this establishes, and what to test next

The fixed three-language replay shows that the final configured pipeline retrieves useful context and keeps answered responses grounded across English, Hindi, and Bengali within this synthetic corpus-derived testset. It also makes the coverage difference visible instead of hiding refusals inside answer-quality averages. It does not establish production usage, open-domain generalization, or quality outside these 300 language-runs.

The most informative next experiment is a human-reviewed parallel Hindi and Bengali set written natively rather than machine-translated, scored alongside the current English-normalized path. That comparison would separate translation effects from retrieval behavior and test whether the same confidence thresholds preserve the quality-versus-coverage balance for native vernacular questions.

Evaluation setup

Embedding model
intfloat/multilingual-e5-large
Judge / generator
gemma4:12b
Retrieval threshold
0.55
Retrieval-gap threshold
0.03