Hybrid Retrieval for Indic Language Question Answering using an Improved RAG Pipeline

Authors

  • Anshika Saini Department of Computer Science and Engineering, Chandigarh University, Punjab, India Author
  • Gorisha Department of Computer Science and Engineering, Chandigarh University, Punjab, India Author
  • Bearda Savita Rajpal Department of Computer Science and Engineering, Chandigarh University, Punjab, India Author
  • Munish Kumar Department of Computer Science and Engineering, Chandigarh University, Punjab, India Author

DOI:

https://doi.org/10.63503/acset.102

Keywords:

Retrieval-Augmented Generation (RAG), Hinglish, Indic natural language processing, Reducing hallucinations, Combined search methods, Languages with few resources

Abstract

Although Large Language Models (LLMs) have progressed significantly, Retrieval-Augmented Generation (RAG) still falls short for under-resourced Indic languages. Existing embedding systems, built primarily for English, face issues with script mismatches and limited word coverage, leading to near-complete retrieval failures in mixed Hinglish contexts. This paper introduces Hybrid Indic-RAG, a simple framework to reduce meaning shifts. It uses a targeted word-matching layer to connect Romanised Hinglish with standard Devanagari Hindi, paired with a two-part scoring system that favours exact word matches over fuzzy semantic signals. Tests in fields like healthcare and books show it cuts retrieval errors by 45% and lifts fact accuracy to 75%, well ahead of standard RAG methods. Our approach offers a low-cost way to build stable, error-free question-answering tools for India's language ecosystem.

Downloads

Published

2026-09-15

Conference Proceedings Volume

Section

Articles

How to Cite

Anshika Saini, Gorisha, Bearda Savita Rajpal, & Munish Kumar. (2026). Hybrid Retrieval for Indic Language Question Answering using an Improved RAG Pipeline. Adroid Conference Series: Engineering and Technology, 2(3), 9-18. https://doi.org/10.63503/acset.102