LLM-Inspired Ontology-Based Semantic Enrichment for FinTech M&A Intelligence in SEC Structured Disclosures

Authors

  • Guanzheng Zhao Financial Engineering, Stevens Institute of Technology, NJ, USA
  • Dingyuan Zhang Business Analytics, University of Rochester, NY, USA
  • Sisi Meng Accounting, University of Rochester, NY, USA

DOI:

https://doi.org/10.51903/jtie.v5i2.566

Keywords:

event intelligence, FinTech, mergers and acquisitions, ontology-based semantic enrichment, SEC structured disclosures

Abstract

This study develops an ontology-based event-intelligence framework for FinTech merger-and-acquisition evidence in U.S. Securities and Exchange Commission (SEC) structured disclosures. Five quarterly SEC Financial Statement Data Sets from 2025Q1 through 2026Q1 contain 32,254 filings, 7,335 registrants, and 18,312,494 numerical XBRL facts; validation adds a 1,800-filing full-text Form 8-K sample and 192 FDIC events. A deterministic ontology maps XBRL tag names, labels, and documentation to acquisition, disposition, valuation, integration, risk, and payment concepts. The method is therefore LLM-inspired semantic enrichment rather than direct LLM extraction. Logistic regression, decision tree, and random forest classifiers are evaluated in three expanding forward-quarter tests with training-only preprocessing and threshold selection. A proximal protocol retains semantically close predictors, whereas a strict protocol excludes label-generating variables and deterministic descendants. Across the rolling tests, proximal logistic-regression M&A detection attains mean F1 = 0.981941, ROC-AUC = 0.999517, and average precision = 0.998315; strict performance falls to F1 = 0.735640, ROC-AUC = 0.930870, and average precision = 0.824107. FinTech M&A F1 declines from 0.921198 to 0.449503. Strict random forests yield F1 = 0.798129 for integration risk and 0.753987 for valuation signals. In independent full text, strict main-text-plus-exhibit F1 is 0.297482; among 24 automatically linked FDIC events in rolling test quarters, 9 are detected. Near-perfect scores therefore describe ontology reconstruction, not transaction-level accuracy.

References

Ali, S., Abuhmed, T., El-Sappagh, S., Muhammad, K., Alonso-Moral, J. M., Confalonieri, R., Guidotti, R., Del Ser, J., Díaz-Rodríguez, N., & Herrera, F. (2023). Explainable Artificial Intelligence (XAI): What we know and what is left to attain trustworthy artificial intelligence. Information Fusion, 99, 101805. https://doi.org/10.1016/j.inffus.2023.101805

Andrade, G., Mitchell, M., & Stafford, E. (2001). New evidence and perspectives on mergers. Journal of Economic Perspectives, 15(2), 103–120. https://doi.org/10.1257/jep.15.2.103

Araci, D. (2019). FinBERT: Financial sentiment analysis with pre-trained language models. arXiv. https://arxiv.org/abs/1908.10063

Arner, D. W., Barberis, J., & Buckley, R. P. (2016). The evolution of FinTech: A new post-crisis paradigm? Georgetown Journal of International Law, 47(4), 1271–1319. https://doi.org/10.2139/ssrn.2676553

Asai, A., Wu, Z., Wang, Y., Sil, A., & Hajishirzi, H. (2024). Self-RAG: Learning to retrieve, generate, and critique through self-reflection. International Conference on Learning Representations. https://openreview.net/forum?id=hSyW5go0v8

Autio, C., Schwartz, R., Dunietz, J., Jain, S., Stanley, M., Tabassi, E., Hall, P., & Roberts, K. (2024). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.600-1

Babina, T., Bahaj, S., Buchak, G., De Marco, F., Foulis, A., Gornall, W., Mazzola, F., & Yu, T. (2025). Customer data access and fintech entry: Early evidence from open banking. Journal of Financial Economics, 169, 103950. https://doi.org/10.1016/j.jfineco.2024.103950

Bartlett, R. P., Morse, A., Stanton, R., & Wallace, N. (2022). Consumer-lending discrimination in the FinTech era. Journal of Financial Economics, 143(1), 30–56. https://doi.org/10.1016/j.jfineco.2021.05.047

Basel Committee on Banking Supervision. (2024). Digitalisation of finance. Bank for International Settlements. https://www.bis.org/bcbs/publ/d575.htm

Benidis, K., Rangapuram, S. S., Flunkert, V., Wang, Y., Maddix, D., Türkmen, A. C., Gasthaus, J., Bohlke-Schneider, M., Salinas, D., Stella, L., Aubet, F.-X., Callot, L., & Januschowski, T. (2023). Deep learning for time series forecasting: Tutorial and literature survey. ACM Computing Surveys, 55(6), Article 121. https://doi.org/10.1145/3533382

Bettencourt, N., Ding, X., & Giesecke, K. (2026). The Stanford EDGAR Filings Dataset: Reconstructing U.S. corporate and financial disclosures into layout-faithful and token-efficient pretraining data. arXiv. https://arxiv.org/abs/2606.18192

Betton, S., Eckbo, B. E., & Thorburn, K. S. (2008). Corporate takeovers. In B. E. Eckbo (Ed.), Handbook of empirical corporate finance (Vol. 2, pp. 291–430). Elsevier.

Bhatia, G., Nagoudi, E. M. B., Cavusoglu, H., & Abdul-Mageed, M. (2024). FinTral: A family of GPT-4 level multimodal financial large language models. arXiv. https://arxiv.org/abs/2402.10986

Bischl, B., Binder, M., Lang, M., Pielok, T., Richter, J., Coors, S., Thomas, J., Ullmann, T., Becker, M., Boulesteix, A.-L., Deng, D., & Lindauer, M. (2023). Hyperparameter optimization: Foundations, algorithms, best practices, and open challenges. WIREs Data Mining and Knowledge Discovery, 13(2), e1484. https://doi.org/10.1002/widm.1484

Bochkay, K., Brown, S. V., Leone, A. J., & Tucker, J. W. (2023). Textual analysis in accounting: What's next? Contemporary Accounting Research, 40(2), 765–805. https://doi.org/10.1111/1911-3846.12825

Breiman, L. (2001). Random forests. Machine Learning, 45, 5–32. https://doi.org/10.1023/A:1010933404324

Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., ... Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901.

Bücker, M., Szepannek, G., Gosiewska, A., & Biecek, P. (2022). Transparency, auditability, and explainability of machine learning models in credit scoring. Journal of the Operational Research Society, 73(1), 70–90. https://doi.org/10.1080/01605682.2021.1922098

Chalkidis, I., Jana, A., Hartung, D., Bommarito, M., Androutsopoulos, I., Katz, D., & Aletras, N. (2022). LexGLUE: A benchmark dataset for legal language understanding in English. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4310–4330. https://doi.org/10.18653/v1/2022.acl-long.297

Chen, M. A., Wu, Q., & Yang, B. (2019). How valuable is FinTech innovation? The Review of Financial Studies, 32(5), 2062–2106. https://doi.org/10.1093/rfs/hhy130

Chen, Y., Zhou, S., & Lin, E. (2025, December). Accounting-aware evidence retrieval for institutional due diligence of tokenized trade receivable RWA. Journal of Technology Informatics and Engineering, 4(3), 649–663. https://doi.org/10.51903/jtie.v4i3.542

Chen, Z., Li, S., Smiley, C., Ma, Z., Shah, S., & Wang, W. Y. (2022). ConvFinQA: Exploring the chain of numerical reasoning in conversational finance question answering. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 6279–6292. https://doi.org/10.18653/v1/2022.emnlp-main.421

Cornelli, G., Frost, J., Gambacorta, L., Rau, P. R., Wardrop, R., & Ziegler, T. (2023). Fintech and big tech credit: Drivers of the growth of digital lending. Journal of Banking & Finance, 148, 106742. https://doi.org/10.1016/j.jbankfin.2022.106742

Daud, S. N. M., Ahmad, A. H., Khalid, A., & Azman-Saini, W. N. W. (2022). FinTech and financial stability: Threat or opportunity? Finance Research Letters, 47, 102667. https://doi.org/10.1016/j.frl.2021.102667

Demirgüç-Kunt, A., Klapper, L., Singer, D., & Ansar, S. (2022). The Global Findex Database 2021: Financial inclusion, digital payments, and resilience in the age of COVID-19. World Bank. https://doi.org/10.1596/978-1-4648-1897-4

Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL-HLT 2019, 4171–4186. https://doi.org/10.18653/v1/N19-1423

Dwivedi, R., Dave, D., Naik, H., Singhal, S., Omer, R., Patel, P., Qian, B., Wen, Z., Shah, T., Morgan, G., & Ranjan, R. (2023). Explainable AI (XAI): Core ideas, techniques, and solutions. ACM Computing Surveys, 55(9), Article 194. https://doi.org/10.1145/3561048

Edge, D., Trinh, H., Cheng, N., Bradley, J., Chao, A., Mody, A., Truitt, S., & Larson, J. (2024). From local to global: A graph RAG approach to query-focused summarization. arXiv. https://arxiv.org/abs/2404.16130

Es, S., James, J., Espinosa Anke, L., & Schockaert, S. (2024). RAGAs: Automated evaluation of retrieval augmented generation. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, 150–158. https://doi.org/10.18653/v1/2024.eacl-demo.16

Federal Deposit Insurance Corporation. (2026a). BankFind Suite: Bank Structure Changes. https://banks.data.fdic.gov/bankfind-suite/oscr

Federal Deposit Insurance Corporation. (2026b). BankFind Suite: Bulk Data and API. https://banks.data.fdic.gov/bankfind-suite/bulkData

Federal Financial Institutions Examination Council. (2026). National Information Center data download: Relationships and transformations. https://www.ffiec.gov/npw/FinancialReport/DataDownload

Feyen, E., Natarajan, H., & Saal, M. (2022). Fintech and the future of finance: Market and policy implications. World Bank. https://www.worldbank.org/en/publication/fintech-and-the-future-of-finance

Financial Stability Board. (2024). The financial stability implications of artificial intelligence. https://www.fsb.org/2024/11/the-financial-stability-implications-of-artificial-intelligence/

Financial Stability Board. (2025). Monitoring adoption of artificial intelligence and related vulnerabilities in the financial sector. https://www.fsb.org/2025/10/monitoring-adoption-of-artificial-intelligence-and-related-vulnerabilities-in-the-financial-sector/

Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., Wang, M., & Wang, H. (2023). Retrieval-augmented generation for large language models: A survey. arXiv. https://arxiv.org/abs/2312.10997

Gomber, P., Koch, J. A., & Siering, M. (2017). Digital finance and FinTech: Current research and future research directions. Journal of Business Economics, 87, 537–580. https://doi.org/10.1007/s11573-017-0852-x

Guo, X., Xia, H., Liu, Z., Cao, H., Yang, Z., Liu, Z., Wang, S., Niu, J., Wang, C., Wang, Y., Liang, X., Huang, X., Zhu, B., Wei, Z., Chen, Y., Shen, W., & Zhang, L. (2023). FinEval: A Chinese financial domain knowledge evaluation benchmark for large language models. arXiv. https://arxiv.org/abs/2308.09975

Harford, J. (2005). What drives merger waves? Journal of Financial Economics, 77(3), 529–560. https://doi.org/10.1016/j.jfineco.2004.05.004

Hewamalage, H., Ackermann, K., & Bergmeir, C. (2023). Forecast evaluation for data scientists: Common pitfalls and best practices. Data Mining and Knowledge Discovery, 37, 788–832. https://doi.org/10.1007/s10618-022-00894-5

Hicks, S. A., Strümke, I., Thambawita, V., Hammou, M., Riegler, M. A., Halvorsen, P., & Parasa, S. (2022). On evaluation metrics for medical applications of artificial intelligence. Scientific Reports, 12, 5979. https://doi.org/10.1038/s41598-022-09954-8

Huang, A. H., Wang, H., & Yang, Y. (2023). FinBERT: A large language model for extracting information from financial text. Contemporary Accounting Research, 40(2), 806–841. https://doi.org/10.1111/1911-3846.12832

Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., & Liu, T. (2025). A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 43(2), Article 42. https://doi.org/10.1145/3703155

International Monetary Fund, Monetary and Capital Markets Department. (2024). Global Financial Stability Report, October 2024: Steadying the course—Uncertainty, artificial intelligence, and financial stability. International Monetary Fund. https://doi.org/10.5089/9798400277573.082

International Organization for Standardization. (2023). ISO/IEC 42001:2023: Information technology—Artificial intelligence—Management system. https://www.iso.org/standard/81230.html

Januschowski, T., Wang, Y., Torkkola, K., Erkkilä, T., Hasson, H., & Gasthaus, J. (2022). Forecasting with trees. International Journal of Forecasting, 38(4), 1473–1481. https://doi.org/10.1016/j.ijforecast.2021.10.004

Jeong, S., Baek, J., Cho, S., Hwang, S. J., & Park, J. (2024). Adaptive-RAG: Learning to adapt retrieval-augmented large language models through question complexity. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 7036–7050. https://doi.org/10.18653/v1/2024.naacl-long.389

Ji, S., Pan, S., Cambria, E., Marttinen, P., & Yu, P. S. (2022). A survey on knowledge graphs: Representation, acquisition, and applications. IEEE Transactions on Neural Networks and Learning Systems, 33(2), 494–514. https://doi.org/10.1109/TNNLS.2021.3070843

Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), Article 248. https://doi.org/10.1145/3571730

Jiang, Y., Chen, J., Makri, E., Chen, J., Li, P., Maatouk, A., Tassiulas, L., Brenner, E., Xiang, B., & Ying, R. (2026). Fin-RATE: A real-world financial analytics and tracking evaluation benchmark for LLMs on SEC filings. arXiv. https://arxiv.org/abs/2602.07294

Jin, J., Huang, T., & Lu, S. (2024, June). A model-risk-friendly probability of default workflow: Calibration, distribution-free uncertainty quantification, and SHAP explanations on the UCI credit card default dataset. Journal of Advanced Computing Systems, 4(6), 74–85. https://doi.org/10.69987/JACS.2024.40606

Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9), 100804. https://doi.org/10.1016/j.patter.2023.100804

Kapoor, S., Cantrell, E. M., Peng, K., Pham, T. H., Bail, C. A., Gundersen, O. E., Hofman, J. M., Hullman, J., Lones, M. A., Malik, M. M., Nanayakkara, P., Poldrack, R. A., Raji, I. D., Roberts, M., Salganik, M. J., Serra-Garcia, M., Stewart, B. M., Vandewiele, G., & Narayanan, A. (2024). REFORMS: Consensus-based recommendations for machine-learning-based science. Science Advances, 10(18), eadk3452. https://doi.org/10.1126/sciadv.adk3452

Lee, J., Stevens, N., Han, S. C., & Song, M. (2024). A survey of large language models in finance (FinLLMs). arXiv. https://arxiv.org/abs/2402.02315

Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474.

Li, C., Bai, J., & Wang, S. (2024, February). Evidence-chain reliable RAG: Word-level hallucination detection, source attribution, and provenance explanation for LLM applications. Journal of Advanced Computing Systems, 4(2), 76–92. https://doi.org/10.69987/JACS.2024.40207

Li, C., Su, W., & Zhang, E. (2023, January). Lightweight hallucination firewall for enterprise LLM applications: Evidence consistency, self-checking, and small-model detection on TruthfulQA. Journal of Advanced Computing Systems, 3(1), 49–65. https://doi.org/10.69987/JACS.2023.30104

Li, F. (2010). The information content of forward-looking statements in corporate filings—A naïve Bayesian machine learning approach. Journal of Accounting Research, 48(5), 1049–1102. https://doi.org/10.1111/j.1475-679X.2010.00382.x

Li, Z., Zhang, K., & Wong, A. (2026, June). Numerical-reasoning guardrails for a quant research assistant: A compact reproducible benchmark using SEC and FRED data. Journal of Technology Informatics and Engineering, 5(2), 75–90. https://doi.org/10.51903/jtie.v5i2.541

Li, Z., Zhou, S., & Zhou, Z. (2025, May). Financial risk dashboard design for institutional RWA investors: Visual hierarchy, chart comprehension, and explainability in FinChart-Bench. International Journal of Graphic Design, 3(1), 196–210. https://doi.org/10.51903/ijgd.v3i1.3715

Liu, N. F., Zhang, T., & Liang, P. (2023). Evaluating verifiability in generative search engines. Findings of the Association for Computational Linguistics: EMNLP 2023, 7001–7025. https://doi.org/10.18653/v1/2023.findings-emnlp.467

Liu, X.-Y., Wang, G., Yang, H., & Zha, D. (2023). FinGPT: Democratizing Internet-scale data for financial large language models. arXiv. https://arxiv.org/abs/2307.10485

Longo, L., Brcic, M., Cabitza, F., Choi, J., Confalonieri, R., Del Ser, J., Guidotti, R., Hayashi, Y., Herrera, F., Holzinger, A., Jiang, R., Khosravi, H., Lecue, F., Malgieri, G., Paez, A., Samek, W., Schneider, J., Speith, T., & Stumpf, S. (2024). Explainable artificial intelligence (XAI) 2.0: A manifesto of open challenges and interdisciplinary research directions. Information Fusion, 106, 102301. https://doi.org/10.1016/j.inffus.2024.102301

Lopez-Lira, A., & Tang, Y. (2023). Can ChatGPT forecast stock price movements? Return predictability and large language models. arXiv. https://arxiv.org/abs/2304.07619

Loughran, T., & McDonald, B. (2011). When is a liability not a liability? Textual analysis, dictionaries, and 10-Ks. The Journal of Finance, 66(1), 35–65. https://doi.org/10.1111/j.1540-6261.2010.01625.x

Loughran, T., & McDonald, B. (2016). Textual analysis in accounting and finance: A survey. Journal of Accounting Research, 54(4), 1187–1230. https://doi.org/10.1111/1475-679X.12123

Loukas, L., Fergadiotis, M., Chalkidis, I., Spyropoulou, E., Malakasiotis, P., Androutsopoulos, I., & Paliouras, G. (2022). FiNER: Financial numeric entity recognition for XBRL tagging. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 4419–4431. https://doi.org/10.18653/v1/2022.acl-long.303

Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems, 30.

Mallen, A., Asai, A., Zhong, V., Das, R., Khashabi, D., & Hajishirzi, H. (2023). When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9802–9822. https://doi.org/10.18653/v1/2023.acl-long.546

Malmendier, U., & Tate, G. (2008). Who makes acquisitions? CEO overconfidence and the market's reaction. Journal of Financial Economics, 89(1), 20–43. https://doi.org/10.1016/j.jfineco.2007.07.002

Manakul, P., Liusie, A., & Gales, M. J. F. (2023). SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 9004–9017. https://doi.org/10.18653/v1/2023.emnlp-main.557

Meng, S., Chen, J., & Zheng, I. (2026, April). LLM-inspired offline reranking for financial search: Query rewriting, hybrid retrieval, and listwise relevance ranking on FiQA. Journal of Technology Informatics and Engineering, 5(1), 361–378. https://doi.org/10.51903/jtie.v5i1.537

Min, S., Krishna, K., Lyu, X., Lewis, M., Yih, W.-t., Koh, P. W., Iyyer, M., Zettlemoyer, L., & Hajishirzi, H. (2023). FActScore: Fine-grained atomic evaluation of factual precision in long-form text generation. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 12076–12100. https://doi.org/10.18653/v1/2023.emnlp-main.741

Molnar, C. (2022). Interpretable machine learning: A guide for making black box models explainable (2nd ed.). https://christophm.github.io/interpretable-ml-book/

Morgan Stanley. (2026). 5 forces driving M&A in 2026. https://www.morganstanley.com/insights/articles/mergers-and-acquisitions-outlook-2026-activity

Nauta, M., Trienes, J., Pathak, S., Nguyen, E., Peters, M., Schmitt, Y., Schlötterer, J., van Keulen, M., & Seifert, C. (2023). From anecdotal evidence to quantitative evaluation methods: A systematic review on evaluating explainable AI. ACM Computing Surveys, 55(13s), Article 295. https://doi.org/10.1145/3583558

Nie, Y., Kong, Y., Dong, X., Mulvey, J. M., Poor, H. V., Wen, Q., & Zohren, S. (2024). A survey of large language models for financial applications: Progress, prospects and challenges. arXiv. https://arxiv.org/abs/2406.11903

Pan, S., Luo, L., Wang, Y., Chen, C., Wang, J., & Wu, X. (2024). Unifying large language models and knowledge graphs: A roadmap. IEEE Transactions on Knowledge and Data Engineering, 36(7), 3580–3599. https://doi.org/10.1109/TKDE.2024.3352100

Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830.

Peng, C., Xia, F., Naseriparsa, M., & Osborne, F. (2023). Knowledge graphs: Opportunities and challenges. Artificial Intelligence Review, 56(11), 13071–13102. https://doi.org/10.1007/s10462-023-10465-9

Petropoulos, F., Apiletti, D., Assimakopoulos, V., Babai, M. Z., Barrow, D. K., Ben Taieb, S., Bergmeir, C., Bessa, R. J., Bijak, J., Boylan, J. E., Browell, J., Carnevale, C., Castle, J. L., Cirillo, P., Clements, M. P., Cordeiro, C., Cyrino Oliveira, F. L., De Baets, S., Dokumentov, A., ... Ziel, F. (2022). Forecasting: Theory and practice. International Journal of Forecasting, 38(3), 705–871. https://doi.org/10.1016/j.ijforecast.2021.11.001

PricewaterhouseCoopers. (2026). Global M&A industry trends: 2026 outlook. https://www.pwc.com/gx/en/services/deals/trends.html

Pushkarna, M., Zaldivar, A., & Kjartansson, O. (2022). Data cards: Purposeful and transparent dataset documentation for responsible AI. Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, 1776–1826. https://doi.org/10.1145/3531146.3533231

Rainio, O., Teuho, J., & Klén, R. (2024). Evaluation metrics and statistical tests for machine learning. Scientific Reports, 14, 6086. https://doi.org/10.1038/s41598-024-56706-x

Rhodes-Kropf, M., Robinson, D. T., & Viswanathan, S. (2005). Valuation waves and merger activity: The empirical evidence. Journal of Financial Economics, 77(3), 561–603. https://doi.org/10.1016/j.jfineco.2004.06.015

Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). “Why should I trust you?” Explaining the predictions of any classifier. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1135–1144. https://doi.org/10.1145/2939672.2939778

Rudin, C., Chen, C., Chen, Z., Huang, H., Semenova, L., & Zhong, C. (2022). Interpretable machine learning: Fundamental principles and 10 grand challenges. Statistics Surveys, 16, 1–85. https://doi.org/10.1214/21-SS133

Saad-Falcon, J., Khattab, O., Potts, C., & Zaharia, M. (2024). ARES: An automated evaluation framework for retrieval-augmented generation systems. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 338–354. https://doi.org/10.18653/v1/2024.naacl-long.20

Saeed, W., & Omlin, C. W. (2023). Explainable AI (XAI): A systematic meta-survey of current challenges and future opportunities. Knowledge-Based Systems, 263, 110273. https://doi.org/10.1016/j.knosys.2023.110273

Securities and Exchange Commission Investor.gov. (2021). How to read an 8-K. https://www.investor.gov/introduction-investing/general-resources/news-alerts/alerts-bulletins/how-read-8

Securities and Exchange Commission. (2004). Additional Form 8-K disclosure requirements and acceleration of filing date. https://www.sec.gov/rules-regulations/2004/03/additional-form-8-k-disclosure-requirements-acceleration-filing-date

Securities and Exchange Commission. (2023). Form 8-K: Current report pursuant to Section 13 or 15(d) of the Securities Exchange Act of 1934. https://www.sec.gov/files/form8-k.pdf

Securities and Exchange Commission. (2025a). 2025 XBRL taxonomies update. https://www.sec.gov/newsroom/whats-new/2503-2025-xbrl-taxonomies-update

Securities and Exchange Commission. (2025b). Developer resources. https://www.sec.gov/about/developer-resources

Securities and Exchange Commission. (2025c). EDGAR application programming interfaces (APIs). https://www.sec.gov/search-filings/edgar-application-programming-interfaces

Securities and Exchange Commission. (2025d). EDGAR Release 25.1. https://www.sec.gov/submit-filings/edgar-news-announcements/edgar-release-25-1

Securities and Exchange Commission. (2026a). Financial Statement Data Sets. https://www.sec.gov/data-research/sec-markets-data/financial-statement-data-sets

Securities and Exchange Commission. (2026b). Exchange Act Form 8-K: Compliance and disclosure interpretations. https://www.sec.gov/rules-regulations/staff-guidance/compliance-disclosure-interpretations/exchange-act-form-8-k

Song, H., Kim, M., Park, D., Shin, Y., & Lee, J.-G. (2023). Learning from noisy labels with deep neural networks: A survey. IEEE Transactions on Neural Networks and Learning Systems, 34(11), 8135–8153. https://doi.org/10.1109/TNNLS.2022.3152527

Su, W., Chen, S., & Zhao, C. (2025, December). Budgeted multi-hop retrieval agent for compositional question answering: A retrieval-policy evaluation on the official MultiHop-RAG benchmark. Journal of Technology Informatics and Engineering, 4(3), 649–662. https://doi.org/10.51903/jtie.v4i3.543

Tabassi, E. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0) (NIST AI 100-1). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.100-1

Tashman, L. J. (2000). Out-of-sample tests of forecasting accuracy: An analysis and review. International Journal of Forecasting, 16(4), 437–450. https://doi.org/10.1016/S0169-2070(00)00065-0

Thakor, A. V. (2020). Fintech and banking: What do we know? Journal of Financial Intermediation, 41, 100833. https://doi.org/10.1016/j.jfi.2019.100833

Tonmoy, S. M. T. I., Zaman, S. M. M., Jain, V., Rani, A., Rawte, V., Chadha, A., & Das, A. (2024). A comprehensive survey of hallucination mitigation techniques in large language models. arXiv. https://arxiv.org/abs/2401.01313

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.

Wiles, O., Gowal, S., Stimberg, F., Rebuffi, S.-A., Ktena, I., Dvijotham, K., & Cemgil, A. T. (2022). A fine-grained analysis on distribution shift. International Conference on Learning Representations. https://openreview.net/forum?id=Dl4LetuLdyK

Wu, Q., Bai, J., & Zhou, X. (2023, March). Evidence-grounded financial RAG: Reducing numerical hallucination in LLM-generated corporate risk memos. Journal of Advanced Computing Systems, 3(3), 65–84. https://doi.org/10.69987/JACS.2023.30306

Wu, S., Irsoy, O., Lu, S., Dabravolski, V., Dredze, M., Gehrmann, S., Kambadur, P., Rosenberg, D., & Mann, G. (2023). BloombergGPT: A large language model for finance. arXiv. https://arxiv.org/abs/2303.17564

Xie, Q., Han, W., Chen, Z., Xiang, R., Zhang, X., He, Y., Xiao, M., Li, D., Dai, Y., Feng, D., Xu, Y., Kang, H., Kuang, Z., Yuan, C., Yang, K., Luo, Z., Zhang, T., Liu, Z., Xiong, G., ... Huang, J. (2024). FinBen: A holistic financial benchmark for large language models. arXiv. https://arxiv.org/abs/2402.12659

Xie, Q., Han, W., Zhang, X., Lai, Y., Peng, M., Lopez-Lira, A., & Huang, J. (2023). PIXIU: A large language model, instruction data and evaluation benchmark for finance. arXiv. https://arxiv.org/abs/2306.05443

Yan, S.-Q., Gu, J.-C., Zhu, Y., & Ling, Z.-H. (2024). Corrective retrieval augmented generation. arXiv. https://arxiv.org/abs/2401.15884

Yang, H., Liu, X.-Y., & Wang, C. D. (2023). FinGPT: Open-source financial large language models. arXiv. https://arxiv.org/abs/2306.06031

Zhang, K., Chen, Y., & Qian, A. (2025, October). Evidence-grounded accounting disclosure review cards: A visual communication framework for LLM-style explanations over SEC financial statements and notes. International Journal of Graphic Design, 3(2), 395. https://doi.org/10.51903/ijgd.v3i2.3710

Zhang, K., Meng, S., & Zhou, E. (2023, February). Evidence-grounded trading desk risk memos over SEC filings: Retrieval-augmented generation with XBRL numeric verification. Journal of Advanced Computing Systems, 3(2), 60–76. https://doi.org/10.69987/JACS.2023.30205

Zhao, G., & Zhang, K. (2024, March). From earnings calls to earnings events: An LLM-guided separation of discretionary and systematic volatility trading signals. Journal of Advanced Computing Systems, 4(3), 126–143. https://doi.org/10.69987/JACS.2024.40309

Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., Du, Y., Yang, C., Chen, Y., Chen, Z., Jiang, J., Ren, R., Li, Y., Tang, X., Liu, Z., ... Wen, J.-R. (2023). A survey of large language models. arXiv. https://arxiv.org/abs/2303.18223

Zhong, Z. S., Chen, J., Zhong, E., & Sun, X. (2025, August). Evidence-calibrated RAG for unanswerable question answering: Retrieval coverage, abstention calibration, and hallucination-proxy analysis on SQuAD 2.0. Journal of Technology Informatics and Engineering, 4(2), 502–520. https://doi.org/10.51903/jtie.v4i2.536

Zhou, S., Chen, Y., & Lee, K. (2026, June). Accounting-aware evidence-constrained agents for disclosure, settlement, and secondary-market risk monitoring in tokenized RWA infrastructure. Journal of Technology Informatics and Engineering, 5(2), 60–74. https://doi.org/10.51903/jtie.v5i2.544

Zhou, S., Li, Z., & Wang, E. (2023, July). Evidence-grounded RAG for tokenized trade receivable disclosure QA under U.S. capital market standards. Journal of Advanced Computing Systems, 3(7), 41–57. https://doi.org/10.69987/JACS.2023.30704

Zhou, S., Li, Z., & Wang, E. (2024, August). Long-document RAG for contractual and insurance clause analysis in receivables RWA structures. Journal of Advanced Computing Systems, 4(8), 88–104. https://doi.org/10.69987/JACS.2024.40810

Downloads

Published

2026-07-29

How to Cite

LLM-Inspired Ontology-Based Semantic Enrichment for FinTech M&A Intelligence in SEC Structured Disclosures. (2026). Journal of Technology Informatics and Engineering, 5(2), 119-141. https://doi.org/10.51903/jtie.v5i2.566