Open Access DOI Assigned

Retrieval Augmented Generation Using Multimodal Large Language Models for Real-Time Knowledge-Grounded Question Answering

Volume 3, Issue 4

  • Author(s)Dr. K. Sujatha
  • AffiliationIndependent Reseracher
  • Page No.138-145
  • Volume, Issue & YearVolume 3, Issue 4, April 2026
  • Published On2026/04/30
  • JournalInternational Journal of Advanced Multidisciplinary Application (IJAMA)
  • ISSN No.3048-9350
  • DOIhttps://doi.org/10.5281/zenodo.20068396

Abstract

The exponential growth of heterogeneous digital information across structured and unstructured repositories presents a critical challenge for large language models (LLMs): the inability to access and reason over dynamically evolving knowledge without costly model retraining. This paper introduces a comprehensive Retrieval Augmented Generation (RAG) framework that integrates multimodal large language models (MLLMs) with real-time, knowledge-grounded question answering systems. The proposed architecture — MultiRAG — combines a dense bi-encoder retrieval backbone with a cross-modal fusion module capable of jointly indexing and retrieving text, images, tables, and structured data. Retrieved multimodal evidence is processed by a vision-language model (VLM) serving as the generative backbone, conditioned on retrieved context through a novel cross-attention grounding mechanism that attenuates hallucination by enforcing faithfulness constraints at the token level. Experiments conducted on four benchmark datasets — Natural Questions, WebQA, MultiModalQA, and a custom real-time knowledge update benchmark (RKUB-2024) — demonstrate that MultiRAG achieves 87.3% Exact Match on open-domain QA, 91.4% answer faithfulness score, and 6.7× reduction in hallucination rate compared to vanilla LLM baselines. Real-time knowledge ingestion pipeline latency averages 340 ms per document, supporting continuous knowledge grounding without model fine-tuning. The system reduces hallucination by 82% over standard LLM deployment and outperforms all retrieval-augmented baselines by 4.2–9.8 percentage points across evaluation metrics

Keywords: Retrieval Augmented Generation, Multimodal LLM, Knowledge-Grounded QA, Dense Retrieval, Cross-Modal Fusion, Vision-Language Models, Hallucination Mitigation, Real-Time Knowledge Updating, Open-Domain QA

References

  1. [1] Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., & Zisserman, A. (2022). Flamingo: A visual language model for few-shot learning. NeurIPS 2022.
  2. [2] Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., & Sifre, L. (2022). Improving language models by retrieving from trillions of tokens. ICML 2022.
  3. [3] Chang, Y., Wang, X., Wang, J., Wu, Y., Yang, L., Zhu, K., & Xie, X. (2022). WebQA: Multihop and multimodal QA. CVPR 2022.
  4. [4] Guu, K., Lee, K., Tung, Z., Pasupat, P., & Chang, M. W. (2020). REALM: Retrieval augmented language model pre-training. ICML 2020.
  5. [5] Izacard, G., & Grave, E. (2021). Leveraging passage retrieval with generative models for open domain question answering. EACL 2021.
  6. [6] Izacard, G., Lewis, P., Lomeli, M., Hosseini, L., Petroni, F., Schick, T., & Riedel, S. (2023). Atlas: Few-shot learning with retrieval augmented language models. JMLR, 24(1).
  7. [7] Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12).
  8. [8] Jia, C., Yang, Y., Xia, Y., Chen, Y. T., Parekh, Z., Pham, H., & Duerig, T. (2021). Scaling up visual and vision-language representation learning with noisy text supervision. ICML 2021.
  9. [9] Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., & Yih, W. T. (2020). Dense passage retrieval for open-domain question answering. EMNLP 2020.
  10. [10] Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., & Petrov, S. (2019). Natural Questions: A benchmark for question answering research. TACL, 7, 452–466.
  11. [11] Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. NeurIPS 2020.
  12. [12] Li, J., Li, D., Savarese, S., & Hoi, S. (2023). BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. ICML 2023.
  13. [13] Liu, H., Li, C., Wu, Q., & Lee, Y. J. (2023). Visual instruction tuning (LLaVA). NeurIPS 2023.
  14. [14] Mallen, A., Asai, A., Zhong, V., Das, R., Khashabi, D., & Hajishirzi, H. (2023). When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. ACL 2023.
  15. [15] Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., & Sutskever, I. (2021). Learning transferable visual models from natural language supervision (CLIP). ICML 2021.
  16. [16] Shi, F., Chen, X., Misra, K., Scales, N., Dohan, D., Chi, E. H., & Zhou, D. (2023). Large language models can be easily distracted by irrelevant context. ICML 2023.
  17. [17] Talmor, A., Yoran, O., Catav, A., Lahav, D., Wang, Y., Asai, A., & Berant, J. (2021). MultiModalQA: Complex question answering over text, tables and images. ICLR 2021.
  18. [18] Yuan, L., Chen, D., Chen, Y. L., Codella, N., Dai, X., Gao, J., & Zhang, L. (2021). Florence: A new foundation model for computer vision. arXiv:2111.11432.

Explore Our Related Journals

Looking for the right journal for your next manuscript? Explore our international peer-reviewed journals covering engineering, management, computer science, artificial intelligence and multidisciplinary research.