Hybrid Retrieval for Indic Language Question Answering using an Improved RAG Pipeline
DOI:
https://doi.org/10.63503/acset.102Keywords:
Retrieval-Augmented Generation (RAG), Hinglish, Indic natural language processing, Reducing hallucinations, Combined search methods, Languages with few resourcesAbstract
Although Large Language Models (LLMs) have progressed significantly, Retrieval-Augmented Generation (RAG) still falls short for under-resourced Indic languages. Existing embedding systems, built primarily for English, face issues with script mismatches and limited word coverage, leading to near-complete retrieval failures in mixed Hinglish contexts. This paper introduces Hybrid Indic-RAG, a simple framework to reduce meaning shifts. It uses a targeted word-matching layer to connect Romanised Hinglish with standard Devanagari Hindi, paired with a two-part scoring system that favours exact word matches over fuzzy semantic signals. Tests in fields like healthcare and books show it cuts retrieval errors by 45% and lifts fact accuracy to 75%, well ahead of standard RAG methods. Our approach offers a low-cost way to build stable, error-free question-answering tools for India's language ecosystem.
Downloads
Published
Conference Proceedings Volume
Section
License
Copyright (c) 2026 The Author(s)

This work is licensed under a Creative Commons Attribution 4.0 International License.