Vernacular RAG Voicebot for Government Schemes
A multilingual, stateful RAG voice-assistant prototype for government-scheme discovery, with explicit evidence gating and grounded generation.
Offline evaluation replays one fixed synthetic set of 100 semantic questions in English, Hindi, and Bengali: 100 questions per language and 300 language-runs in total.
Government-scheme discovery across text and speech
The prototype treats scheme discovery as a question-answering problem: a person can describe a need or eligibility situation without first knowing the scheme name or source document. English, Hindi, and Bengali text access keeps that search from depending on an English-only query.
Hindi and Bengali voice input serves the same path when typing is not the preferred mode, and vernacular speech is returned only for voice turns. Language choice is explicit because the IndicConformer speech service requires the caller to declare the spoken language rather than relying on automatic detection.
Product surface
$ uvicorn src.api.app:app
INFO Loading multilingual-e5-large and bge-reranker-v2-m3
INFO FAISS + BM25 indexes loaded
200 GET /healthcheck 761 chunks ready
Recorded multi-turn API replay from the project evaluation artifacts. The follow-up resolves “it” from session memory, retrieves APY evidence again, and returns a second grounded response.
The public portfolio replays the run locally in the browser. It does not expose the private model server or API credentials.
Watch the RAG system work
Scrub the timeline to see source documents become chunks, ranks become evidence, and evidence become a response.
Detect changes
Parse PDFs and text without flattening their structure
A SHA-256 manifest skips unchanged PDF and TXT files during incremental ingestion.
Source files become page-aware content with tables and reading order intact.
Retrieval and uncertainty are separate decisions
After normalization, multilingual-e5 semantic search covers paraphrase and intent while BM25 protects exact scheme names, eligibility terms, and distinctive wording. Reciprocal-rank fusion combines their ranked candidates without pretending their raw scores share one scale; bge-reranker-v2-m3 then evaluates the fused query-passage pairs more closely.
The reranked set passes only when its strongest score reaches 0.55 and leads the second result by at least 0.03. Weak evidence routes to a fallback; accepted evidence enters a grounded prompt and then a verifier that also fails closed. General conversation deliberately bypasses retrieval and rejoins the path at vernacular translation.
Conversation state and evaluation stay decoupled
LangGraph checkpoints retain conversation content and graph state under the session thread ID. A separate Alembic-managed chat_sessions table keeps the title, timestamps, and turn count needed to list conversations. The browser retains the active ID, loads the session list, and asks the replay endpoint for the selected history, so resume behavior does not make product metadata part of graph internals.
Evaluation stays outside the request path because the local judge and metric calls are a repeatable regression harness, not user-facing work. Retrieval precision and recall use all 100 runs per language because retrieval happens before gating; answer-quality metrics use only the 84 English, 79 Hindi, and 74 Bengali answered rows; answered rate uses all runs so refusal coverage remains visible.
Evidence lab
Multilingual retrieval and answer quality
How consistent is the final RAG pipeline across the three languages?
Context precision
Context recall
Faithfulness
Answer relevancy
Answer correctness
Answer similarity
Shared 0.0-1.0 scale; higher is better.
Interpretation note Retrieval facts use all 100 rows per language; answer-quality facts use only answered rows. Do not average them into one overall score.
Values and denominators
| Metric | English n=100 | Hindi n=100 | Bengali n=100 |
|---|---|---|---|
| Context precision | 0.908 | 0.858 | 0.853 |
| Context recall | 0.970 | 0.943 | 0.923 |
| Metric | English n=84 | Hindi n=79 | Bengali n=74 |
|---|---|---|---|
| Faithfulness | 0.996 | 0.994 | 0.989 |
| Answer relevancy | 0.892 | 0.889 | 0.887 |
| Answer correctness | 0.738 | 0.747 | 0.738 |
| Answer similarity | 0.951 | 0.947 | 0.952 |
Evidence-gated answer coverage
How often did the configured evidence gate permit a grounded answer?
All 100 runs per language · configured gate 0.70
Read Answered rate moves from 0.84 in English to 0.79 in Hindi and 0.74 in Bengali while answered-row quality remains comparatively stable.
Values and denominators
| Metric | English n=100 | Hindi n=100 | Bengali n=100 |
|---|---|---|---|
| Answered rate | 0.840 | 0.790 | 0.740 |
What this establishes, and what to test next
The fixed three-language replay shows that the final configured pipeline retrieves useful context and keeps answered responses grounded across English, Hindi, and Bengali within this synthetic corpus-derived testset. It also makes the coverage difference visible instead of hiding refusals inside answer-quality averages. It does not establish production usage, open-domain generalization, or quality outside these 300 language-runs.
The most informative next experiment is a human-reviewed parallel Hindi and Bengali set written natively rather than machine-translated, scored alongside the current English-normalized path. That comparison would separate translation effects from retrieval behavior and test whether the same confidence thresholds preserve the quality-versus-coverage balance for native vernacular questions.
Evaluation setup
- Embedding model
- intfloat/multilingual-e5-large
- Judge / generator
- gemma4:12b
- Retrieval threshold
- 0.55
- Retrieval-gap threshold
- 0.03