A Systematic Literature Review on Software Testing Prediction Models

Authors

DOI:

https://doi.org/10.51903/jtie.v5i2.526

Keywords:

Software testing prediction, Software defect prediction, Machine learning, deep learning, Explainable AI

Abstract

Software testing prediction models play a critical role in improving software quality, reducing faults, and optimizing testing. Although past studies show that these models have improved over time, little focus has been given to them. Many studies rely on limited datasets that do not fully capture the complexity of real software systems, which limits how well the models can generalize. There is also insufficient evidence on how these models use existing datasets in practical, real-life settings, since most evaluations are conducted under controlled or experimental conditions. In addition, key aspects such as interpretability, generalizability, and practical usability are still not adequately addressed, which reduces trust and slows down adoption in practice. This study presents a systematic literature review of recent empirical research on predictive approaches in software testing, focusing on machine learning, deep learning, and hybrid techniques. A structured methodology was used, including clearly defined inclusion and exclusion criteria, systematic searches across major academic databases, quality assessment, and data extraction from 22 selected studies. The analysis considered model types, datasets, feature methods, evaluation metrics, and methodological approaches. The findings show a shift from traditional statistical models to advanced artificial intelligence techniques, including graph neural networks, contrastive learning, and deep fuzzy clustering. In addition, optimization and data balancing techniques improve predictive performance, while explainable artificial intelligence enhances model interpretability. However, challenges such as limited cross-project generalization and insufficient industrial validation still exist. 

References

Abdulwareth, A. J., & Al-Shargabi, A. A. (2021). Toward a Multi-Criteria Framework for Selecting Software Testing Tools. IEEE Access, 9, 158872–158891. https://doi.org/10.1109/ACCESS.2021.3128071

Aggarwal, A., Kumar, S., & Gupta, R. (2024). Testing coverage based NHPP software reliability growth modeling with testing effort and change-point. International Journal of System Assurance Engineering and Management, 15(11), 5157–5166. https://doi.org/10.1007/s13198-024-02504-7

Ali, A., & Gravino, C. (2019). A systematic literature review of software effort prediction using machine-learning methods. Journal of Software: Evolution and Process, 31(10), e2211. https://doi.org/10.1002/smr.2211

Arani, A. K., Le, T. H. M., Zahedi, M., & Babar, M. A. (2024). Systematic Literature Review on Application of Learning-Based Approaches in Continuous Integration. IEEE Access, 12, 135419–135450. https://doi.org/10.1109/ACCESS.2024.3424276

Arshad, A., Riaz, S., Jiao, L., & Murthy, A. (2018). The Empirical Study of Semi-Supervised Deep Fuzzy C-Mean Clustering for Software Fault Prediction. IEEE Access, 6, 47047–47061. https://doi.org/10.1109/ACCESS.2018.2866082

Aziz, S. R., Khan, T. A., & Nadeem, A. (2019). Experimental Validation of Inheritance Metrics’ Impact on Software Fault Prediction. IEEE Access, 7, 85262–85275. https://doi.org/10.1109/ACCESS.2019.2924040

Behera, A. K., Chaudhury, P., & Dash, Ch. S. K. (2025). A comprehensive survey on intelligent software reliability prediction. Discover Computing, 28(1), 90. https://doi.org/10.1007/s10791-025-09597-z

Boloori, A., Zamanifar, A., & Farhadi, A. (2024). Enhancing software defect prediction models using metaheuristics with a learning to rank approach. Discover Data, 2(1), 11. https://doi.org/10.1007/s44248-024-00016-0

Boukhlif, M., Hanine, M., Kharmoum, N., Ruigómez Noriega, A., García Obeso, D., & Ashraf, I. (2024). Natural Language Processing-Based Software Testing: A Systematic Literature Review. IEEE Access, 12, 79383–79400. https://doi.org/10.1109/ACCESS.2024.3407753

Chan, P. Y. P., & Keung, J. (2024). Validating Unsupervised Machine Learning Techniques for Software Defect Prediction with Generic Metamorphic Testing. IEEE Access, 12, 165155–165172. https://doi.org/10.1109/ACCESS.2024.3494044

Chengxiao, Y., Owais, J., Ahmed, A., Khan, M. O., Xiaoyang, Z., & Tunio, M. H. (2023). Software Defect Prediction Using Machine Learning—A Systematic Literature Review. 2023 20th International Computer Conference on Wavelet Active Media Technology and Information Processing (ICCWAMTIP), 1–9. https://doi.org/10.1109/ICCWAMTIP60502.2023.10387043

Choras, M., Springer, T., Kozik, R., Lopez, L., Martinez-Fernandez, S., Ram, P., Rodriguez, P., & Franch, X. (2020). Measuring and Improving Agile Processes in a Small-Size Software Development Company. IEEE Access, 8, 78452–78466. https://doi.org/10.1109/ACCESS.2020.2990117

Dai, P., Zhu, H., Wu, J., & He, H. (2025). An integrated graph neural network model for joint software defect prediction and code quality assessment. Scientific Reports, 16(1), 1677. https://doi.org/10.1038/s41598-025-31209-5

De Silva, D., & Hewawasam, L. (2024). The Impact of Software Testing on Serverless Applications. IEEE Access, 12, 51086–51099. https://doi.org/10.1109/ACCESS.2024.3384459

Durrani, U. K., Akpinar, M., Fatih Adak, M., Talha Kabakus, A., Maruf Öztürk, M., & Saleh, M. (2024). A Decade of Progress: A Systematic Literature Review on the Integration of AI in Software Engineering Phases and Activities (2013-2023). IEEE Access, 12, 171185–171204. https://doi.org/10.1109/ACCESS.2024.3488904

EskandariNasab, M., Hamdi, S. M., & Filali Boubrahimi, S. (2026). FlaPLeT: A full-stack web platform for end-to-end time series data processing and machine learning in solar flare prediction. SoftwareX, 33, 102540. https://doi.org/10.1016/j.softx.2026.102540

Faraji, A., & Pombo, N. (2025). AI-Driven Software Test Automation: An AI4SE-Oriented Survey of Techniques, Tools, and Challenges. IEEE Access, 13, 183296–183313. https://doi.org/10.1109/ACCESS.2025.3623944

Fatima, S., Hemmati, H., & C. Briand, L. (2024). FlakyFix: Using Large Language Models for Predicting Flaky Test Fix Categories and Test Code Repair. IEEE Transactions on Software Engineering, 50(12), 3146–3171. https://doi.org/10.1109/TSE.2024.3472476

Gong, L., Jiang, S., & Jiang, L. (2020). Conditional Domain Adversarial Adaptation for Heterogeneous Defect Prediction. IEEE Access, 8, 150738–150749. https://doi.org/10.1109/ACCESS.2020.3017101

Gordieiev, O., Rainer, A., Kharchenko, V., Pishchukhina, O., & Gordieieva, D. (2024). A Unified Approach to the Development of Technology-Based Software Quality Models on the Example of Blockchain Systems. IEEE Access, 12, 118875–118889. https://doi.org/10.1109/ACCESS.2024.344827

Gurcan, F., Dalveren, G. G. M., Cagiltay, N. E., Roman, D., & Soylu, A. (2022). Evolution of Software Testing Strategies and Trends: Semantic Content Analysis of Software Research Corpus of the Last 40 Years. IEEE Access, 10, 106093–106109. https://doi.org/10.1109/ACCESS.2022.3211949

Haddaway, N. R., Page, M. J., Pritchard, C. C., & McGuinness, L. A. (2022). PRISMA2020: An R package and Shiny app for producing PRISMA 2020-compliant flow diagrams, with interactivity for optimised digital transparency and Open Synthesis Campbell Systematic Reviews, 18, e1230. https://doi.org/10.1002/cl2.1230

Hemandez-Molinos, Ma. J., Sanchez-Garcia, A. J., & Barrientos-Martinez, R. E. (2021). Classification Algorithms for Software Defect Prediction: A Systematic Literature Review. 2021 9th International Conference in Software Engineering Research and Innovation (CONISOFT), 189–196. https://doi.org/10.1109/CONISOFT52520.2021.00034

Khatibsyarbini, M., Isa, M. A., Jawawi, D. N. A., Hamed, H. N. A., & Mohamed Suffian, M. D. (2019). Test Case Prioritization Using Firefly Algorithm for Software Testing. IEEE Access, 7, 132360–132373. https://doi.org/10.1109/ACCESS.2019.2940620

Kitchenham, B., & Charters, S. (2007). Guidelines for performing systematic literature reviews in software engineering (EBSE Technical Report EBSE-2007-01). School of Computer Science and Mathematics, Keele University and Department of Computer Science, University of Durham

Kim, Y. S., Song, K. Y., & Chang, I. H. (2023). Prediction and Comparative Analysis of Software Reliability Model Based on NHPP and Deep Learning. Applied Sciences, 13(11), 6730. https://doi.org/10.3390/app13116730

Komee, & Pachauri, B. (2025). Cost analysis of enhanced software reliability model with Weibull testing effort function using simulated annealing. Scientific Reports, 15(1), 41927. https://doi.org/10.1038/s41598-025-25841-4

Li, N., Shepperd, M., & Guo, Y. (2020). A Systematic Review of Unsupervised Learning Techniques for Software Defect Prediction (arXiv:1907.12027). arXiv. https://doi.org/10.48550/arXiv.1907.12027

Li, W., & Fang, C.-C. (2025). Enhanced Software Testing Model Under Software Warranty Policy Considering Debugger Learning and Time Postponement Factors. IEEE Access, 13, 85116–85130. https://doi.org/10.1109/ACCESS.2025.3569201

Malhotra, R., & Patidar, J. (2025). Enhanced software change proneness prediction in android applications using balanced data techniques and advanced language models. Discover Computing, 28(1), 51. https://doi.org/10.1007/s10791-025-09552-y

Mehmood, A., Ilyas, Q. M., Ahmad, M., & Shi, Z. (2024). Test Suite Optimization Using Machine Learning Techniques: A Comprehensive Study. IEEE Access, 12, 168645–168671. https://doi.org/10.1109/ACCESS.2024.3490453

Navarro Cedeño, G. O., Moya, K. C., Dormond, A. S., González-Torres, A., & Rojas-Hernández, Y. (2023). Systematic Literature Review: Machine Learning for Software Fault Prediction. 2023 IEEE 41st Central America and Panama Convention (CONCAPAN XLI), 1–6. https://doi.org/10.1109/CONCAPANXLI59599.2023.10517566

Olivares-Galindo, J. A., Sánchez-García, Á. J., Barrientos-Martínez, R. E., & Ocharán-Hernández, J. O. (2023). Ensemble Classifiers in Software Defect Prediction: A Systematic Literature Review. 2023 11th International Conference in Software Engineering Research and Innovation (CONISOFT), 1–8. https://doi.org/10.1109/CONISOFT58849.2023.00011

Petersen, K., Vakkalanka, S., & Kuzniarz, L. (2021). Guidelines for conducting systematic mapping studies in software engineering: An update. Information and Software Technology, 133, 106455. https://doi.org/10.1016/j.infsof.2020.106455

Phung, K., Ogunshile, E., & Aydin, M. E. (2025). Domain-specific implications of error-type metrics in risk-based software fault prediction. Software Quality Journal, 33(1), 7. https://doi.org/10.1007/s11219-024-09704-1

Sahar, S., Younas, M., Khan, M. M., & Sarwar, M. U. (2024). DP-CCL: A Supervised Contrastive Learning Approach Using CodeBERT Model in Software Defect Prediction. IEEE Access, 12, 22582–22594. https://doi.org/10.1109/ACCESS.2024.3362896

Tao, C., Gao, J., & Wang, T. (2019). Testing and Quality Validation for AI Software–Perspectives, Issues, and Practices. IEEE Access, 7, 120164–120175. https://doi.org/10.1109/ACCESS.2019.2937107

Tosun, A., Bener, A. B., & Akbarinasaji, S. (2017). A systematic literature review on the applications of Bayesian networks to predict software quality. Software Quality Journal, 25(1), 273–305. https://doi.org/10.1007/s11219-015-9297-z

Tumar, I., Hassouneh, Y., Turabieh, H., & Thaher, T. (2020). Enhanced Binary Moth Flame Optimization as a Feature Selection Algorithm to Predict Software Fault Prediction. IEEE Access, 8, 8041–8055. https://doi.org/10.1109/ACCESS.2020.2964321

Urbikain-Pelayo, G., Olvera-Trejo, D., López De Lacalle, L. N., & Elías-Zúñiga, A. (2025). Turn+: A MATLAB-based software for dynamic turning, chatter analysis, and surface roughness prediction. SoftwareX, 32, 102401. https://doi.org/10.1016/j.softx.2025.102401

Wahono, R. S. (2015). A systematic literature review of software defect prediction: Research trends, datasets, methods and frameworks. Journal of Software Engineering, 1(1).

Wang, X., Ali, S., & Taibi, D. (2025). The Landscape of Quantum Software Testing Tools. IEEE Software, 42(5), 136–140. https://doi.org/10.1109/MS.2025.3578154

Weng, S., Feng, Y., Yin, Y., Zhang, Z., & Xu, B. (2026). Data preparation and quality for code-centric generative software engineering tasks: A systematic literature review. Frontiers of Computer Science, 20(9), 2009203. https://doi.org/10.1007/s11704-025-41376-3

Wohlin, C., Runeson, P., Höst, M., Ohlsson, M. C., Regnell, B., & Wesslén, A. (2022). Experimentation in software engineering (2nd ed.). Springer

Zou, Q.-Y., Yang, Z.-Y., Qiu, X.-R., Yu, J.-H., Shi, Y.-Y., & Chen, N. (2025). Counterfactual Contrastive Explanations for Software Defect Prediction: Toward Better Model Understanding and Accuracy. IEEE Transactions on Reliability, 74(4), 5000–5014. https://doi.org/10.1109/TR.2025.3611079

Downloads

Published

2026-08-21

Data Availability Statement

No new datasets were generated or analyzed in this study. This research is based on a systematic literature review of publicly available, peer-reviewed publications retrieved from academic databases, including IEEE Xplore, ScienceDirect, SpringerLink, ACM Digital Library, and Google Scholar. All sources used in the analysis are cited and listed in the reference section of the article.

How to Cite

A Systematic Literature Review on Software Testing Prediction Models. (2026). Journal of Technology Informatics and Engineering, 5(2), 167-187. https://doi.org/10.51903/jtie.v5i2.526