Predictive pattern analysis in administrative and tax litigation through generative AI and regular expressions
DOI:
https://doi.org/10.47975/ijdl.v7.1305Keywords:
Large language models; document classification; regular expressions; predictive pattern analysis; generative AI.Abstract
Text classification is the task of assigning a piece of text to an appropriate category, with the set of possible categories varying by domain. In this study, we evaluate ChatGPT’s accuracy in classifying judicial rulings issued by an Argentine court, using prompts based on regular expression matches. We compare its performance to two alternative approaches: (1) classification based on natural language descriptions of the categories, and (2) a traditional regex-based algorithm. The regex patterns were developed by domain experts employed at the partner organization, and all experiments were conducted in Spanish. Our results show that ChatGPT achieves high accuracy when prompted with both natural language descriptions and regex-based category identifications. However, generating outputs with ChatGPT is significantly more time-consuming. Despite this limitation, the model’s accessible interface and minimal technical requirements make it a promising tool for automating document classification. Furthermore, the use of regex-based prompts enables human oversight and active involvement in the classification process—an advantage in domains where interpretability and accountability are critical. All code, prompts, and evaluation materials are publicly available at <https://github.com/FAIR-IALAB-UBA/ai-regex>.
Downloads
References
ABUKAR AHMED, Abdinor; KHALIF ALI, Mohamed. User-centric adoption of democratized generative AI: Focus on human-machine interaction and overcoming challenges. International Journal of Engineering Trends and Technology, v. 72, issue 9, p. 78-95, 2024.
AGARWAL, Basant; MITTAL, Namita. Text classification using machine learning methods-a survey. In: Proceedings of the Second International Conference on Soft Computing for Problem Solving (SocProS 2012). New Dehli: Springer, 2014. p. 701-709.
BHAMBRI, Gaurav. Information overload in business organizations and entrepreneurship: An analytical review of the literature. Business Information Review, v. 38, issue 4, 2021, p. 193–200.
BIDERMAN, Stella Biderman, et al. Lessons from the trenches on reproducible evaluation of language models. 2024. Available at: <https://arxiv.org/abs/2405.14782>.; REUEL, Anka; et al. Open problems in technical AI governance. 2024. Available at: <https://arxiv.org/abs/2407.14981>.
BURDEN, John. Evaluating AI evaluation: perils and prospects. 2024. Available at: <https://arxiv.org/abs/2407.09221>.
CABELLERO, William N.; JENKINS, Phillip R. On large language models in national security applications. Stat, v. 14, issue 2, e70057, 2025.
FRONTIER MODEL FORUM. Issue brief: Early best practices for frontier ai safety evaluations. 2024. Available at: <https://www.frontiermodelforum.org/updates/early-best-practices-for-frontier-ai-safety-evaluations/>.
JIANG, Shuo; HU, Jie; MAGEE, Christopher L.; LUO, Jianxi. Deep learning for technical document classification. IEEE Transactions on Engineering Management, v. 71, p. 1163-1179, 2022.
KLEIN, L. K; EARL, E.; CUNDICK, D. Reducing information overload in your organization. Harvard Business Review, 2023.
KOSTINA, Arina; DIKAIAKOS, Marios D.; STEFANIDIS, Dimosthenis; PALLIS, George Pallis. Large language models for text classification: case study and comprehensive review. 2025. Available at: <https://arxiv.org/abs/2501.08457>.
KOWSARI, Kamram; JAFARI MEIMANDI, Kiana; HEIDARYSAFA, Mojtaba; MENDU, Sanjana; BARNES, Laura; BROWN, Donald. Text classification algorithms: A survey. Information, v. 10, issue 4, p. 150, 2019.
LI, Yinghao; RAMPRASAD, Rampi; ZHANG, Chao. A simple but effective approach to improve structured language model output for information extraction. 2024. Available at: <https://arxiv.org/abs/2402.13364>.
MAHONEY, Christian; GRONVALL, Peter; HUBER-FLIFLET; ZHANG, Jianping. Explainable text classification techniques in legal document review: locating rationales without using human annotated training text snippets. In: 2022 IEEE International Conference on Big Data (Big Data). IEEE, p. 2044-2051, 2022.
MCINTOSH, Timothy R., et al. Inadequacies of large language model benchmarks in the era of generative artificial intelligence. IEEE Transactions on Artificial Intelligence, 2025. Available at: <https://arxiv.org/abs/2402.09880>.
NATIONAL JUDICIAL STATISTICS SYSTEM (SNEJ). Annual Report: Technical Report. National Judicial Statistics System (SNEJ): Argentina, 2022.
PALANIVINAYAGAM, Ashokkumar; ZIAD EL-BAYEH, Claude; DAMAŠEVICIUS, Robertas. Twenty years of machine-learning-based text classification: A systematic review. Algorithms, v. 16, issue 5, p. 236, 2023.
PASKOV, Patricia; BERGLUND, Lukas; SMITH, Everett; SODER, Lisa. GPAI evaluations standards taskforce: towards effective AI governance. 2024. Available at: <https://arxiv.org/abs/2411.13808>.
PEÑA ALMANSA, Alejandro; et al. Leveraging large language models for topic classification in the domain of public affairs. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 2023.
PEREZ, Ethan; et al. Discovering language model behaviors with model-written evaluations. In: Findings of the Association for Computational Linguistics. Toronto: ACL, 2023. p. 13387–13434.
ROHERA, Pritika; GINIMAV, Chaitrali; SAWANT, Gayatri; JOSHI, Raviraj. Better to ask in english? Evaluating factual accuracy of multilingual LLMms in english and low-resource languages. Available at: <https://arxiv.org/abs/2504.20022>.
SAHOO, Pranab; SINGH, Ayush Kumar; SAHA, Sriparna; JAIN, Vinija; MONDAL, Samrat; CHADHA, Aman. A systematic survey of prompt engineering in large language models: techniques and applications. 2024. Available at: <https://arxiv.org/abs 2402.07927>.
SCHUT, Lisa; GAL, Yarin; FARQUHAR, Sebastian. Do multilingual LLMs think in english? 2025. Available at: <https://arxiv.org/abs/2502.15603>.
SCLAR, Melanie; CHOI, Yejin; TSVETKOV, Yulia; SUHR, Alane. Quantifying language models’ sensitivity to spurious features in prompt design or: how I learned to start worrying about prompt formatting. 2023. Available at: <https://arxiv.org/abs/2310.11324>.
SEGER, Elizabeth. What do we mean when we talk about “AI democratisation”. 2023. Available at: <https://www.governance.ai/analysis/what-do-we-mean-when-we-talk-about-ai-democratisation>.
SUN, Xiaofei; LI, Xiaoya; LI, Jiwei; WU, Fei; GUO, Shangwei; ZHANG, Tianwei; WANG, Guoyin. Text classification via large language models. 2023. Available at: <https://arxiv.org/abs/2305.08377>.
SUPREME COURT OF ARGENTINA. Statistics For The Year 2024: Technical report. Supreme Court of Argentina, 2024.
WAGH, Vedangi; KHANDVE, Snehal; JOSHI, Isha; WANI, Apurva; KALE, Geetanjali; JOSHI, Raviraj. Comparative study of long document classification. In: TENCON 2021-2021 IEEE Region 10 Conference (TENCON). 2021. p. 732–737.
WEIDINGER, Laura; et al. Sociotechnical safety evaluation of generative AI systems. 2023. Available at: <https://arxiv.org/abs/2310.11986>.
YE, Junjie; XU, Nuo; WANG, Yikun; ZHOU, Jie; ZHANG, Qi; GUI, Tao; HUANG, Xuanjing. LLM-DA: Data augmentation via large language models for few-shot named entity recognition. 2024. Available at: <https://arxiv.org/abs/2402.14568>.
YOON, Sungwook Yoon; et. al. Digital innovation in public administration through intelligent public sector automation (IPSA): Strategies and challenges. Journal of Multimedia Information System, v. 11, issue 4, p. 249-260, 2024.
ZHANG, Shuai; GU, Xiaodong; CHEN, Yuting; SHEN, Beijun. Infere: step-by-step regex generation via chain of inference. In: 38th IEEE/ACM International Conference on Automated Software Engineering (ASE). IEEE, 2023. p. 1505–1515.
Downloads
Published
How to Cite
License
Copyright (c) 2026 Juan Gustavo Corvalán, Sofía Tammaro, Gisel Alvarado, Luca Nicolás Forziati Gangi, Carina Mariel Papini, Mariana Sánchez Caparrós, Agostina Celeste Jara Rey, Florencia Paola Croci, Lola Ramos Pereyra, Melania Gadea, Luciano Dalla Via (Autor)

This work is licensed under a Creative Commons Attribution 4.0 International License.
This journal is licensed by
Creative Commons Attribution-NonCommercial-4.0 International.
Attribution-ShareAlike 4.0 International (CC BY-NC 4.0)
Submission and publication of paper are free; Works evaluated by blind double review; the Journal uses S_cites and CrossCheck (anti-plagiarism); and complies with the COPE Editors Guide - Committee on Publication Ethics, in addition to the Elsevier and SciELO recommendations.













