Predictive pattern analysis in administrative and tax litigation through generative AI and regular expressions
DOI:
https://doi.org/10.47975/ijdl.v7.1305Palavras-chave:
Large language models; document classification; regular expressions; predictive pattern analysis; generative AI.Resumo
Text classification is the task of assigning a piece of text to an appropriate category, with the set of possible categories varying by domain. In this study, we evaluate ChatGPT’s accuracy in classifying judicial rulings issued by an Argentine court, using prompts based on regular expression matches. We compare its performance to two alternative approaches: (1) classification based on natural language descriptions of the categories, and (2) a traditional regex-based algorithm. The regex patterns were developed by domain experts employed at the partner organization, and all experiments were conducted in Spanish. Our results show that ChatGPT achieves high accuracy when prompted with both natural language descriptions and regex-based category identifications. However, generating outputs with ChatGPT is significantly more time-consuming. Despite this limitation, the model’s accessible interface and minimal technical requirements make it a promising tool for automating document classification. Furthermore, the use of regex-based prompts enables human oversight and active involvement in the classification process—an advantage in domains where interpretability and accountability are critical. All code, prompts, and evaluation materials are publicly available at <https://github.com/FAIR-IALAB-UBA/ai-regex>.
Downloads
Referências
ABUKAR AHMED, Abdinor; KHALIF ALI, Mohamed. User-centric adoption of democratized generative AI: Focus on human-machine interaction and overcoming challenges. International Journal of Engineering Trends and Technology, v. 72, issue 9, p. 78-95, 2024.
AGARWAL, Basant; MITTAL, Namita. Text classification using machine learning methods-a survey. In: Proceedings of the Second International Conference on Soft Computing for Problem Solving (SocProS 2012). New Dehli: Springer, 2014. p. 701-709.
BHAMBRI, Gaurav. Information overload in business organizations and entrepreneurship: An analytical review of the literature. Business Information Review, v. 38, issue 4, 2021, p. 193–200.
BIDERMAN, Stella Biderman, et al. Lessons from the trenches on reproducible evaluation of language models. 2024. Available at: <https://arxiv.org/abs/2405.14782>.; REUEL, Anka; et al. Open problems in technical AI governance. 2024. Available at: <https://arxiv.org/abs/2407.14981>.
BURDEN, John. Evaluating AI evaluation: perils and prospects. 2024. Available at: <https://arxiv.org/abs/2407.09221>.
CABELLERO, William N.; JENKINS, Phillip R. On large language models in national security applications. Stat, v. 14, issue 2, e70057, 2025.
FRONTIER MODEL FORUM. Issue brief: Early best practices for frontier ai safety evaluations. 2024. Available at: <https://www.frontiermodelforum.org/updates/early-best-practices-for-frontier-ai-safety-evaluations/>.
JIANG, Shuo; HU, Jie; MAGEE, Christopher L.; LUO, Jianxi. Deep learning for technical document classification. IEEE Transactions on Engineering Management, v. 71, p. 1163-1179, 2022.
KLEIN, L. K; EARL, E.; CUNDICK, D. Reducing information overload in your organization. Harvard Business Review, 2023.
KOSTINA, Arina; DIKAIAKOS, Marios D.; STEFANIDIS, Dimosthenis; PALLIS, George Pallis. Large language models for text classification: case study and comprehensive review. 2025. Available at: <https://arxiv.org/abs/2501.08457>.
KOWSARI, Kamram; JAFARI MEIMANDI, Kiana; HEIDARYSAFA, Mojtaba; MENDU, Sanjana; BARNES, Laura; BROWN, Donald. Text classification algorithms: A survey. Information, v. 10, issue 4, p. 150, 2019.
LI, Yinghao; RAMPRASAD, Rampi; ZHANG, Chao. A simple but effective approach to improve structured language model output for information extraction. 2024. Available at: <https://arxiv.org/abs/2402.13364>.
MAHONEY, Christian; GRONVALL, Peter; HUBER-FLIFLET; ZHANG, Jianping. Explainable text classification techniques in legal document review: locating rationales without using human annotated training text snippets. In: 2022 IEEE International Conference on Big Data (Big Data). IEEE, p. 2044-2051, 2022.
MCINTOSH, Timothy R., et al. Inadequacies of large language model benchmarks in the era of generative artificial intelligence. IEEE Transactions on Artificial Intelligence, 2025. Available at: <https://arxiv.org/abs/2402.09880>.
NATIONAL JUDICIAL STATISTICS SYSTEM (SNEJ). Annual Report: Technical Report. National Judicial Statistics System (SNEJ): Argentina, 2022.
PALANIVINAYAGAM, Ashokkumar; ZIAD EL-BAYEH, Claude; DAMAŠEVICIUS, Robertas. Twenty years of machine-learning-based text classification: A systematic review. Algorithms, v. 16, issue 5, p. 236, 2023.
PASKOV, Patricia; BERGLUND, Lukas; SMITH, Everett; SODER, Lisa. GPAI evaluations standards taskforce: towards effective AI governance. 2024. Available at: <https://arxiv.org/abs/2411.13808>.
PEÑA ALMANSA, Alejandro; et al. Leveraging large language models for topic classification in the domain of public affairs. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 2023.
PEREZ, Ethan; et al. Discovering language model behaviors with model-written evaluations. In: Findings of the Association for Computational Linguistics. Toronto: ACL, 2023. p. 13387–13434.
ROHERA, Pritika; GINIMAV, Chaitrali; SAWANT, Gayatri; JOSHI, Raviraj. Better to ask in english? Evaluating factual accuracy of multilingual LLMms in english and low-resource languages. Available at: <https://arxiv.org/abs/2504.20022>.
SAHOO, Pranab; SINGH, Ayush Kumar; SAHA, Sriparna; JAIN, Vinija; MONDAL, Samrat; CHADHA, Aman. A systematic survey of prompt engineering in large language models: techniques and applications. 2024. Available at: <https://arxiv.org/abs 2402.07927>.
SCHUT, Lisa; GAL, Yarin; FARQUHAR, Sebastian. Do multilingual LLMs think in english? 2025. Available at: <https://arxiv.org/abs/2502.15603>.
SCLAR, Melanie; CHOI, Yejin; TSVETKOV, Yulia; SUHR, Alane. Quantifying language models’ sensitivity to spurious features in prompt design or: how I learned to start worrying about prompt formatting. 2023. Available at: <https://arxiv.org/abs/2310.11324>.
SEGER, Elizabeth. What do we mean when we talk about “AI democratisation”. 2023. Available at: <https://www.governance.ai/analysis/what-do-we-mean-when-we-talk-about-ai-democratisation>.
SUN, Xiaofei; LI, Xiaoya; LI, Jiwei; WU, Fei; GUO, Shangwei; ZHANG, Tianwei; WANG, Guoyin. Text classification via large language models. 2023. Available at: <https://arxiv.org/abs/2305.08377>.
SUPREME COURT OF ARGENTINA. Statistics For The Year 2024: Technical report. Supreme Court of Argentina, 2024.
WAGH, Vedangi; KHANDVE, Snehal; JOSHI, Isha; WANI, Apurva; KALE, Geetanjali; JOSHI, Raviraj. Comparative study of long document classification. In: TENCON 2021-2021 IEEE Region 10 Conference (TENCON). 2021. p. 732–737.
WEIDINGER, Laura; et al. Sociotechnical safety evaluation of generative AI systems. 2023. Available at: <https://arxiv.org/abs/2310.11986>.
YE, Junjie; XU, Nuo; WANG, Yikun; ZHOU, Jie; ZHANG, Qi; GUI, Tao; HUANG, Xuanjing. LLM-DA: Data augmentation via large language models for few-shot named entity recognition. 2024. Available at: <https://arxiv.org/abs/2402.14568>.
YOON, Sungwook Yoon; et. al. Digital innovation in public administration through intelligent public sector automation (IPSA): Strategies and challenges. Journal of Multimedia Information System, v. 11, issue 4, p. 249-260, 2024.
ZHANG, Shuai; GU, Xiaodong; CHEN, Yuting; SHEN, Beijun. Infere: step-by-step regex generation via chain of inference. In: 38th IEEE/ACM International Conference on Automated Software Engineering (ASE). IEEE, 2023. p. 1505–1515.
Downloads
Publicado
Como Citar
Licença
Copyright (c) 2026 Juan Gustavo Corvalán, Sofía Tammaro, Gisel Alvarado, Luca Nicolás Forziati Gangi, Carina Mariel Papini, Mariana Sánchez Caparrós, Agostina Celeste Jara Rey, Florencia Paola Croci, Lola Ramos Pereyra, Melania Gadea, Luciano Dalla Via (Autor)

Este trabalho está licenciado sob uma licença Creative Commons Attribution 4.0 International License.
Autores que publicam nesta revista concordam com os seguintes termos:
- Autores mantém os direitos autorais e concedem à revista o direito de primeira publicação, com o trabalho simultaneamente licenciado sob a Creative Commons - Atribuição 4.0 Internacional que permite o compartilhamento do trabalho com reconhecimento da autoria e publicação inicial nesta revista.
- Autores têm autorização para assumir contratos adicionais separadamente, para distribuição não-exclusiva da versão do trabalho publicada nesta revista (ex.: publicar em repositório institucional ou como capítulo de livro), com reconhecimento de autoria e publicação inicial nesta revista.
- Autores têm permissão e são estimulados a publicar e distribuir seu trabalho online (ex.: em repositórios institucionais ou na sua página pessoal) a qualquer ponto antes ou durante o processo editorial, já que isso pode gerar alterações produtivas, bem como aumentar o impacto e a citação do trabalho publicado (Veja O Efeito do Acesso Livre).













