Curriculum Learning Query Rewriting Optimization Based on Reinforcement Learning

Authors

  • Wenbo Tong International College, Hebei University, Baoding, Hebei, 071000, China

DOI:

https://doi.org/10.54097/wqtsra13

Keywords:

Query rewriting, reinforcement learning, curriculum learning, HotpotQA.

Abstract

Query rewriting is the key to improving the effectiveness of information retrieval, but existing reinforcement learning (RL) based methods often face the challenges of unstable training and slow convergence when the difficulty of the query is uneven. Inspired by the human learning process of "from easy to difficult", curriculum learning (CL), as a general training strategy that can improve the speed of model generalization and convergence, provides new ideas for solving this problem. This study uses CL to provide a sample scheduling strategy for RL training, with the core being a three-level adaptive curriculum learning scheduler. The scheduler dynamically adjusts the mixing ratio of extracting samples of different difficulty levels from the dataset based on the real-time performance (reward) of the model during training, thereby achieving a progressive training strategy for RL models from easy to difficult. Experiments on the HotpotQA dataset show that compared with the random sampling baseline, this method improves the reward value, average performance and training stability by 17.4%, 24.9% and 51.9% respectively, which preliminarily verifies the effectiveness of CL in natural language processing tasks.

References

[1] Ma, X., Gong, Y., He, P., et al. (2023) Query rewriting in retrieval-augmented large language models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 5303-5315.

[2] Peng, W., Li, G., Jiang, Y., et al. (2024) Large language model based long-tail query rewriting in taobao search. Companion Proceedings of the ACM Web Conference 2024, 20-28.

[3] Chaudhari, S., Aggarwal, P., Murahari, V., et al. (2024) RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs. arXiv preprint arXiv:2404.08555. Retrieved from https://arxiv.org/abs/2404.08555

[4] Wang, X., Zhou, Y., Chen, H., et al. (2024) Curriculum learning: Theories, approaches, applications, tools, and future directions in the era of large language models. Companion Proceedings of the ACM Web Conference 2024, 1306-1310.

[5] Yang, Z., Qi, P., Zhang, S., et al. (2018) HotpotQA: A dataset for diverse, explainable multi-hop question answering. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 2369-2380.

[6] Wang, S., Zhang, S., Zhang, J., et al. (2024) Reinforcement learning enhanced llms: A survey. arXiv preprint arXiv:2412.10400. Retrieved from https://arxiv.org/abs/2412.10400

[7] Xi, Z., Chen, W., Hong, B., et al. (2024) Training large language models for reasoning through reverse curriculum reinforcement learning. arXiv preprint arXiv:2402.05808. Retrieved from https://arxiv.org/abs/2402.05808

[8] Parashar, S., Gui, S., Li, X., et al. (2025) Curriculum reinforcement learning from easy to hard tasks improves LLM reasoning. arXiv preprint arXiv:2506.06632. Retrieved from https://arxiv.org/abs/2506.06632

[9] Shi, T., Wu, Y., Song, L., et al. (2025) Efficient reinforcement finetuning via adaptive curriculum learning. arXiv preprint arXiv:2504.05520. Retrieved from https://arxiv.org/abs/2504.05520

[10] Zhou, H., Huang, H., Zhao, Z., et al. (2026) Lost in benchmarks? rethinking large language model benchmarking with item response theory. Proceedings of the AAAI Conference on Artificial Intelligence, 40(41), 35085-35093.

[11] Song, M., Zheng, M., Li, Z., et al. (2025) Fastcurl: Curriculum reinforcement learning with stage-wise context scaling for efficient training r1-like reasoning models. arXiv preprint arXiv:2503.17287. Retrieved from https://arxiv.org/abs/2503.17287

Downloads

Published

13-08-2026

Issue

Section

Articles

How to Cite

Tong, W. (2026). Curriculum Learning Query Rewriting Optimization Based on Reinforcement Learning. Mathematical Modeling and Algorithm Application, 9(3), 77-83. https://doi.org/10.54097/wqtsra13