SC-HyDE: Mitigating Hallucinations in Chinese Legal Question Answering via Self-Corrected Hypothetical Document Embeddings

Authors

  • Haoran Li School of mathematics and statistics, Northeastern University at Qinhuangdao, Qinhuangdao, China

DOI:

https://doi.org/10.54097/f24wnv40

Keywords:

Retrieval-Augmented Generation, Legal Question Answering, Hallucination Mitigation, Hypothetical Document Embeddings, Self-Correction.

Abstract

Retrieval-Augmented Generation (RAG) has been cited as a potential technique for reducing hallucinations in Large Language Models (LLM). However, applying RAG in the Chinese legal landscape remains a challenge due to the large semantic gap between informal user queries and professional legal terminology, and the risk of retrieval noise and false information generation. To address these concerns, this paper proposes SC-HyDE, a novel, training-free system that blends generative expansion with self-reflective filtering. First, the domain based Hypothetical Document Embeddings (HyDE) system is applied to generate an imaginative legal analysis, effectively filling out the semantic gap and restoring retrieval accuracy. Second, since HyDE can produce hallucinating legal reference points within its latent space, there is a simple module incorporated for Self-Correction Critic. This module compares the relevance of retrieved documents to the original query by using a zero-shot prompting strategy while excluding noise before the second generation stage. The proposed framework is evaluated against a sample of 588 cases from official legal datasets. The experimental results show that SC-HyDE performs well above standard RAG and vanilla HyDE levels. In particular, the method drops the Noise Rate from 88.33% to 25.06% and the Hallucination Rate from 25.68% to 13.61%, offering an effective approach for those high stakes legal consultancies where factual accuracy is a major concern.

References

[1] P. Lewis et al., “Retrieval-augmented generation for knowledge-intensive NLP tasks,” in Advances in Neural Information Processing Systems 33 (NeurIPS 2020), 2020, pp. 9459–9474J.

[2] C. Xiao et al., “CAIL2018: A large-scale dataset for legal judgment prediction,” in Proc. 27th Int. Joint Conf. Artificial Intelligence (IJCAI-18), 2018, pp. 4487–4493

[3] L. Gao, X. Ma, J. Lin, and J. Callan, “Precise zero-shot dense retrieval without relevance labels,” in Proc. 61st Annu. Meeting Assoc. Computational Linguistics (Vol. 1: Long Papers), 2023, pp. 1762–1777

[4] A. Asai, Z. Wu, Y. Wang, A. Sil, and H. Hajishirzi, “Self-RAG: Learning to retrieve, generate, and critique through self-reflection,” in Proc. 12th Int. Conf. Learning Representations (ICLR 2024), 2024.

[5] S. Yan, J.-C. Gu, Y.-Z. Jiang, and Z.-H. Ling, “Corrective retrieval augmented generation,” in Proc. 2024 Conf. North American Chapter Assoc. Computational Linguistics (NAACL 2024),

[6] Z. Fei et al., “LawBench: Benchmarking legal knowledge of large language models,” in Proc. 2024 Conf. Empirical Methods in Natural Language Processing (EMNLP 2024), 2024, pp. 7933–7962.

[7] S. Yue et al., “DISC-LawLLM: Fine-tuning large language models for intelligent legal services,” arXiv:2309.11325. 2023.

[8] A. B. Hou, W. Jurayj, N. Holzenberger, A. Blair-Stanek, and B. Van Durme, “Gaps or hallucinations? Gazing into machine-generated legal analysis for fine-grained text evaluations,” arXiv:2409.09947. 2024.

[9] V. Noël, E. Y. Seidou, C. K. Capo-Chichi, and G. Amari, “HalluGraph: Auditable hallucination detection for legal RAG systems via knowledge graph alignment,” arXiv:2512.01659. 2025.

[10] P. Trivedi et al., “Self-rationalization improves LLM as a fine-grained judge,” arXiv:2410.05495. 2024.

Downloads

Published

13-08-2026

Issue

Section

Articles

How to Cite

Li, H. (2026). SC-HyDE: Mitigating Hallucinations in Chinese Legal Question Answering via Self-Corrected Hypothetical Document Embeddings. Mathematical Modeling and Algorithm Application, 9(3), 25-31. https://doi.org/10.54097/f24wnv40