IntentRouter-QAC: Distilling LLM-Derived Intent Signals into Small Language Models for Context-Aware Autosuggest and Search Entry Routing
DOI:
https://doi.org/10.51903/jtie.v5i1.563Keywords:
query autocomplete, autosuggest, search entry routing, intent-aware search, small language modelsAbstract
Search boxes increasingly function as routing interfaces: a partial query may require a completion, immediate submission, a context-sensitive suggestion, or a fallback action. This paper evaluates IntentRouter-QAC as a lightweight lexical routing framework and tests product retrieval independently. Five mutually exclusive operational outcomes are derived within the 20,000-row AmazonQAC test file and evaluated with a strict session-disjoint temporal split; product-query provenance is not used as a route class. Across five seeds, the combined word/character TF-IDF router reached 0.473 mean accuracy (SD 0.031) and 0.330 mean macro-F1 (SD 0.018), indicating uneven performance across routes. On the 951-row temporal test set, context-nearest completion achieved 0.033 Success@10 and 0.025 MRR@10, compared with 0.029 and 0.023 for confidence-gated routing. The paired MRR@10 difference was 0.002 (95% CI -0.007 to 0.010; Holm-adjusted p = 1.000). Product retrieval on 480 WANDS queries and 42,994 products produced a different pattern: a word/character lexical hybrid reached 0.690 nDCG@10 and 0.504 Exact-MRR@10, significantly exceeding the title-only baseline on both measures. Route-aware decision-making is therefore useful as an architectural separation of entry actions, but it does not automatically improve exact autocomplete ranking over a strong session-context baseline.
References
amazon/AmazonQAC · Datasets at Hugging Face. (n.d.). Retrieved July 31, 2026, from https://huggingface.co/datasets/amazon/AmazonQAC
Bai, J., Wang, H., Wu, Q., & Zhang, B. (2026). Privacy-Robust Incrementality Estimation in Cookieless Settings via Uplift Modeling: Reproducible Evidence from the Hillstrom E-Mail Experiment. Journal of Technology Informatics and Engineering, 5(1), 17–38. https://doi.org/10.51903/jtie.v5i1.468
Bar-Yossef, Z., & Kraus, N. (2011). Context-Sensitive Query Auto-Completion. Proceedings of the 20th International Conference on World Wide Web, 107–116. https://doi.org/10.1145/1963405.1963424
Broder, A. (2002). A Taxonomy of Web Search. ACM SIGIR Forum, 36(2), 3–10. https://doi.org/10.1145/792550.792552
Cai, F., & de Rijke, M. (2016). A Survey of Query Auto Completion in Information Retrieval. Foundations and Trends® in Information Retrieval, 10(4), 273–363. https://doi.org/10.1561/1500000055
Chang, X., Lu, Y., & Zhong, Z. S. (2026). Review-Grounded Explainable Recommendation with Faithfulness Evaluation on Amazon Reviews. JEECS (Journal of Electrical Engineering and Computer Sciences), 11(1), 9–22. https://doi.org/10.54732/jeecs.v11i1.2
Chen, Y., Liu, S., Liu, Z., Sun, W., Baltrunas, L., & Schroeder, B. (2022). WANDS: Dataset for Product Search Relevance Assessment (pp. 128–141). https://doi.org/10.1007/978-3-030-99736-6_9
Chen, Y., & Xu, H. (2026). Trust-Calibrated Multilingual RAG for Humanitarian Information Platforms: Empirical Evaluation on OMoS-QA for Migration Information Access. International Journal of Graphic Design, 4(1), 141–164. https://doi.org/10.51903/ijgd.v4i1.3552
Chenyu Li, Jingwen Bai, & Samuel Wang. (2024). Evidence-Chain Reliable RAG: Word-Level Hallucination Detection, Source Attribution, and Provenance Explanation for LLM Applications. Journal of Advanced Computing Systems, 4(2), 76–92. https://doi.org/10.69987/JACS.2024.40207
Chenyu Li, Wenhao Su, & Eric Zhang. (2023). Lightweight Hallucination Firewall for Enterprise LLM Applications: Evidence Consistency, Self-Checking, and Small-Model Detection on TruthfulQA. Journal of Advanced Computing Systems, 3(1), 49–65. https://doi.org/10.69987/JACS.2023.30104
Devlin, J., Chang, M.-W., Lee, K., Google, K. T., & Language, A. I. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North, 4171–4186. https://doi.org/10.18653/V1/N19-1423
Dror, R., Baumer, G., Shlomov, S., & Reichart, R. (2018). The Hitchhiker’s Guide to Testing Statistical Significance in Natural Language Processing. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1383–1392. https://doi.org/10.18653/v1/P18-1128
Everaert, D., Patki, R., Zheng, T., & Potts, C. (2024). AmazonQAC: A Large-Scale, Naturalistic Query Autocomplete Dataset. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track, 1046–1055. https://doi.org/10.18653/v1/2024.emnlp-industry.78
Hinton, G., Vinyals, O., & Dean, J. (2015). Distilling the Knowledge in a Neural Network. https://arxiv.org/pdf/1503.02531
Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., & Liu, Q. (2020). TinyBERT: Distilling BERT for Natural Language Understanding. Findings of the Association for Computational Linguistics: EMNLP 2020, 4163–4174. https://doi.org/10.18653/v1/2020.findings-emnlp.372
Jiayi Nie, & David Zheng. (2023). Ambiguity-Aware HDFS Log Anomaly Detection with Retrieval-Augmented Failure Narratives and Selective Refusal. Journal of Advanced Computing Systems, 3(1), 66–80. https://doi.org/10.69987/JACS.2023.30105
Jiaying Jin, Tina Huang, & Sam Lu. (2024a). A Model-Risk-Friendly Probability of Default Workflow: Calibration, Distribution-Free Uncertainty Quantification, and SHAP Explanations on the UCI Credit Card Default Dataset. Journal of Advanced Computing Systems, 4(6), 74–85. https://doi.org/10.69987/JACS.2024.40606
Jiaying Jin, Tina Huang, & Sam Lu. (2024b). Cost-Sensitive Learning, Simulated PU Learning, and One-Class Autoencoding for Extreme-Imbalance Credit Card Fraud Detection. Journal of Advanced Computing Systems, 4(6), 64–73. https://doi.org/10.69987/JACS.2024.40605
Jin, J. (2025a). Evidence-Chain Reliable RAG: Hallucination Detection, Source Attribution, and Deterministic Provenance Explanations. Journal of Technology Informatics and Engineering, 4(2), 520–533. https://doi.org/10.51903/jtie.v4i2.535
Jin, J. (2025b). LLM-Style Evidence Cards for Scientific Search Interfaces: A UI/UX Design Framework for Retrieval Transparency, Ranking Trust, and Visual Evidence Hierarchy. International Journal of Graphic Design, 3(2), 397–414. https://doi.org/10.51903/IJGD.V3I2.3698
Jinyi Mu, Yifei Lu, & Michelle Smith. (2023). LLM-Assisted Incrementality (Uplift) Modeling for Email Advertising: From Feature Interactions to Interpretable Audience–Creative–Channel Policies. Journal of Advanced Computing Systems, 3(1), 31–48. https://doi.org/10.69987/JACS.2023.30103
Joachims, T. (2002). Optimizing Search Engines Using Clickthrough Data. Proceedings of the Eighth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 133–142. https://doi.org/10.1145/775047.775067
Joulin, A., Grave, É., Bojanowski, P., & Mikolov, T. (2017). Bag of Tricks for Efficient Text Classification. In the Association for Computational Linguistics (Vol. 2, pp. 427–431). https://aclanthology.org/E17-2068/
Kuo, M.-J., Zheng, D., & Hires, J. (2025). Federated Topic-Preference Learning for Knowledge-Grounded Chat with Differential Privacy. Journal of Technology Informatics and Engineering, 4(2), 385–401. https://doi.org/10.51903/jtie.v4i2.502
Li, C., Liu, G., & Zhao, Z. (2026). Cost-Aware LLM-Style Routing for AIOps Log Analysis: Log Parsing, Anomaly Detection, Fault Diagnosis, and Incident Summarization on LogEval Task Files. Journal of Technology Informatics and Engineering, 5(2), 91–103. https://doi.org/10.51903/jtie.v5i2.538
Pàmies-Estrems, D., & Garcia-Alfaro, J. (2023). On the Self-Adjustment of Privacy Safeguards for Query Log Streams. Computers & Security, 134, 103450. https://doi.org/10.1016/j.cose.2023.103450
Ponte, J. M., & Croft, W. B. (1998). A Language Modeling Approach to Information Retrieval. Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, 275–281. https://doi.org/10.1145/290941.291008
Radlinski, F., & Joachims, T. (2005). Query Chains. Proceedings of the Eleventh ACM SIGKDD International Conference on Knowledge Discovery in Data Mining, 239–248. https://doi.org/10.1145/1081870.1081899
Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 3980–3990. https://doi.org/10.18653/v1/D19-1410
Robertson, S., & Zaragoza, H. (2009). The Probabilistic Relevance Framework: BM25 and Beyond. Foundations and Trends® in Information Retrieval, 4(1–2), 1–174. https://doi.org/10.1561/1500000019
Rose, D. E., & Levinson, D. (2004). Understanding User Goals in Web Search. Proceedings of the 13th International Conference on World Wide Web, 13–19. https://doi.org/10.1145/988672.988675
Shenghan Lu, & David Zhou. (2023). LLM-Augmented Customer Representation Learning for Next-Purchase Prediction in Online Retail. Journal of Advanced Computing Systems, 3(3), 50–64. https://doi.org/10.69987/JACS.2023.30305
Srinivasan, K., Raman, K., Samanta, A., Liao, L., Bertelli, L., & Bendersky, M. (2022). QUILL: Query Intent with Large Language Models using Retrieval Augmentation and Multi-stage Distillation. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: Industry Track, 492–501. https://doi.org/10.18653/v1/2022.emnlp-industry.50
Su, W., Chen, S., & Zhao, C. (2025). Budgeted Multi-Hop Retrieval Agent for Compositional Question Answering: A Retrieval-Policy Evaluation on the Official MultiHop-RAG Benchmark. Journal of Technology Informatics and Engineering, 4(3), 649–662. https://doi.org/10.51903/jtie.v4i3.543
Su, W., Rao, H., & Ma, E. (2026). Privacy and Data-Integrity Risk Cards for LLM Agents: A UI/UX Design Framework for Secure Human Oversight under Prompt-Injection Attacks. International Journal of Graphic Design, 4(1), 186–191. https://doi.org/10.51903/ijgd.v4i1.3699
Wang, Z., & Culotta, A. (2020). Identifying Spurious Correlations for Robust Text Classification. Findings of the Association for Computational Linguistics: EMNLP 2020, 3431–3440. https://doi.org/10.18653/v1/2020.findings-emnlp.308
Xiaohan Chang, Tong Ye, & Sophia Luo. (2023). LLM-as-Reranker for Personalized Recommendation: Popularity Bias Mitigation and Faithful Natural-Language Explanations on MovieLens 100K. Journal of Advanced Computing Systems, 3(8), 61–78. https://doi.org/10.69987/JACS.2023.30806
Xin, Q. (2026). Self-Supervised Customer Representation Learning for Segmentation and Next-Purchase Prediction on UCI Online Retail. J-INTECH, 14(01), 20–37. https://doi.org/10.32664/j-intech.v14i01.2229
Ye, T., Mu, J., & Hunter, J. (2026). Off-Policy Evaluation and Conservative Policy Selection for Slot-Level Dynamic Bidding and Ranking on the Open Bandit Dataset (Small). Journal of Technology Informatics and Engineering, 5(1), 178–199. https://doi.org/10.51903/jtie.v5i1.503
Yifei Lu, Jinyi Mu, & Thao Tran. (2024). Uncertainty-Aware Uplift Modeling for Safer Marketing Targeting: Conformal Prediction and Bayesian Calibration with LCB Policies. Journal of Advanced Computing Systems, 4(5), 84–101. https://doi.org/10.69987/JACS.2024.40507
Yuanzheng Chen, Yitian Zhang, David Chau, & Matt Sherman. (2023). Credit Card Default Risk Tiering with Probability Calibration and Uncertainty-Driven Rejection: A Reproducible Study on the UCI Credit Card Clients Dataset. Journal of Advanced Computing Systems, 3(4), 31–47. https://doi.org/10.69987/JACS.2023.30403
Yunhe Li. (2024). Findable then Explainable: Retrieval–Summary Integration for Code Intelligence on a Lightweight CodeSearchNet Subset. Journal of Advanced Computing Systems, 4(7), 65–82. https://doi.org/10.69987/JACS.2024.40706
Zhong, Z. S., Chen, J., Zhong, E., & Sun, X. (2025). Evidence-Calibrated RAG for Unanswerable Question Answering: Retrieval Coverage, Abstention Calibration, and Hallucination-Proxy Analysis on SQuAD 2.0. Journal of Technology Informatics and Engineering, 4(2), 502–520. https://doi.org/10.51903/jtie.v4i2.536
Zhong, Z. S., Chenyu Li, & Rao, H. (2026). Trajectory Reliability Prediction for Generalist AI Agents: Tool-Use Failure Analysis and Success Forecasting on ZClawBench. Journal of Technology Informatics and Engineering, 5(1), 341–360. https://doi.org/10.51903/jtie.v5i1.539
Ziliang Samuel Zhong, Ruiyan Ma, & Hailey Zhao. (2023). Human-Uncertainty Distillation for Calibrated Vision Models on CIFAR-10H. Journal of Advanced Computing Systems, 3(2), 77–89. https://doi.org/10.69987/JACS.2023.30206
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Haowei Tu, Yinchen Shi, Sophia Chen

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

