Predictive pattern analysis in administrative and tax litigation through generative AI and regular expressions

Autores

DOI:

https://doi.org/10.47975/ijdl.v7.1305

Palavras-chave:

Large language models; document classification; regular expressions; predictive pattern analysis; generative AI.

Resumo

Text classification is the task of assigning a piece of text to an appropriate category, with the set of possible categories varying by domain. In this study, we evaluate ChatGPT’s accuracy in classifying judicial rulings issued by an Argentine court, using prompts based on regular expression matches. We compare its performance to two alternative approaches: (1) classification based on natural language descriptions of the categories, and (2) a traditional regex-based algorithm. The regex patterns were developed by domain experts employed at the partner organization, and all experiments were conducted in Spanish. Our results show that ChatGPT achieves high accuracy when prompted with both natural language descriptions and regex-based category identifications. However, generating outputs with ChatGPT is significantly more time-consuming. Despite this limitation, the model’s accessible interface and minimal technical requirements make it a promising tool for automating document classification. Furthermore, the use of regex-based prompts enables human oversight and active involvement in the classification process—an advantage in domains where interpretability and accountability are critical. All code, prompts, and evaluation materials are publicly available at <https://github.com/FAIR-IALAB-UBA/ai-regex>.

Downloads

Não há dados estatísticos.

Biografia do Autor

Juan Gustavo Corvalán, Universidad de Buenos Aires (Buenos Aires, Argentina)

Director of the Innovation and Artificial Intelligence Laboratory at the University of Buenos Aires Law School (Buenos Aires, Argentina). Holds a Doctorate in Legal Sciences from the Universidad del Salvador. Director of the Postgraduate Program in Artificial Intelligence and Law at the University of Buenos Aires (UBA). Director of the Diploma Program in Law 4.0 at Austral University. Co-creator of Prometea, the first predictive artificial intelligence system serving the justice system. Co-creator of PretorIA and Academic Director for the implementation of that system at the Constitutional Court of Colombia. General Project Director under the agreement to implement Prometea at the Inter-American Court of Human Rights. Currently serves as Deputy Prosecutor General for Contentious-Administrative and Tax Matters before the Superior Court of Justice of the Autonomous City of Buenos Aires. E-mail: corvalanjuang@gmail.com.     

Sofía Tammaro, Universidad de Buenos Aires (Buenos Aires, Argentina)

Criminal Justice Agenda Lead at the Innovation and Artificial Intelligence Laboratory (IALAB) at tbe Faculty of Law, University of Buenos Aires (Buenos Aires, Argentina). Lecturer in the Postgraduate Program on Artificial Intelligence and Law, University of Buenos Aires. E-mail: sofiatammaro97@gmail.com.

Gisel Alvarado, Universidad de Buenos Aires (Buenos Aires, Argentina)

Member of the Innovation and Artificial Intelligence Laboratory (IALAB) at the University of Buenos Aires Law School (Buenos Aires, Argentina). Law student at the Catholic University of Cuyo. E-mail: alvaradofgise@gmail.com.

Luca Nicolás Forziati Gangi, Universidad de Buenos Aires (Buenos Aires, Argentina)

Member of the Innovation and Artificial Intelligence Laboratory (IALAB) at the University of Buenos Aires Law School (Buenos Aires, Argentina). Bachelor of Laws from the Universidad Argentina de la Empresa. Lawyer. E-mail: lforziatigangi@uade.edu.ar.

Carina Mariel Papini, Universidad de Buenos Aires (Buenos Aires, Argentina)

Member of the Innovation and Artificial Intelligence Laboratory (IALAB) at the University of Buenos Aires Law School (Buenos Aires, Argentina). Academic Coordinator of the Continuing Education Program on Artificial Intelligence and Law at the University of Buenos Aires Law School. Currently pursuing a Master’s degree in Constitutional Procedural Law at the University of Buenos Aires and a Master’s degree in Artificial Intelligence at the Centro Europeo de Posgrado. E-mail: carinapapini79@gmail.com.

Mariana Sánchez Caparrós, Universidad de Buenos Aires (Buenos Aires, Argentina)

Member of the Innovation and Artificial Intelligence Laboratory (IALAB) at the University of Buenos Aires Law School (Buenos Aires, Argentina). Holds a Doctorate in Legal Sciences from UCA and a Master’s degree in Administrative Law from Austral University. E-mail: mariana.sanchezcaparros@ialab.com.ar.

Agostina Celeste Jara Rey, Universidad de Buenos Aires (Buenos Aires, Argentina)

Abogada con Posgrado en Cuestiones Complejas en Materia Societaria, por la Universidad Nacional del Nordeste. Maestrando en Derecho Empresario por la Universidad Nacional del Nordeste. Investigadora en Inteligencia Artificial en el Laboratorio de Inteligencia Artificial de la Universidad de Buenos Aires. Legal Engineer Jr en Legal Hub.

Florencia Paola Croci, Universidad de Buenos Aires (Buenos Aires, Argentina)

Member of the Innovation and Artificial Intelligence Laboratory (IALAB) at the University of Buenos Aires Law School (Buenos Aires, Argentina). Law graduate of the University of Buenos Aires. E-mail: florenciacroci@ialab.com.ar.

Lola Ramos Pereyra, Universidad de Buenos Aires (Buenos Aires, Argentina)

Member of the Innovation and Artificial Intelligence Laboratory (IALAB) at the University of Buenos Aires Law School (Buenos Aires, Argentina). Academic Coordination Assistant for the Postgraduate Program in Artificial Intelligence and Law at the University of Buenos Aires. E-mail: loliramos99gmail.com.

Melania Gadea, Universidad de Buenos Aires (Buenos Aires, Argentina)

Member of the Innovation and Artificial Intelligence Laboratory (IALAB) at the University of Buenos Aires Law School (Buenos Aires, Argentina). Law graduate of the University of Buenos Aires. E-mail: melaniagadea00@gmail.com.

Luciano Dalla Via, Universidad de Buenos Aires (Buenos Aires, Argentina)

Member of the Innovation and Artificial Intelligence Laboratory (IALAB) at the University of Buenos Aires Law School (Buenos Aires, Argentina). Law graduate of the Pontifical Catholic University of Argentina. Holder of a Master’s degree in Law and New Technologies from the School of Legal Practice at the Complutense University of Madrid. E-mail: dallavia.ldv@gmail.com.

Referências

ABUKAR AHMED, Abdinor; KHALIF ALI, Mohamed. User-centric adoption of democratized generative AI: Focus on human-machine interaction and overcoming challenges. International Journal of Engineering Trends and Technology, v. 72, issue 9, p. 78-95, 2024.

AGARWAL, Basant; MITTAL, Namita. Text classification using machine learning methods-a survey. In: Proceedings of the Second International Conference on Soft Computing for Problem Solving (SocProS 2012). New Dehli: Springer, 2014. p. 701-709.

BHAMBRI, Gaurav. Information overload in business organizations and entrepreneurship: An analytical review of the literature. Business Information Review, v. 38, issue 4, 2021, p. 193–200.

BIDERMAN, Stella Biderman, et al. Lessons from the trenches on reproducible evaluation of language models. 2024. Available at: <https://arxiv.org/abs/2405.14782>.; REUEL, Anka; et al. Open problems in technical AI governance. 2024. Available at: <https://arxiv.org/abs/2407.14981>.

BURDEN, John. Evaluating AI evaluation: perils and prospects. 2024. Available at: <https://arxiv.org/abs/2407.09221>.

CABELLERO, William N.; JENKINS, Phillip R. On large language models in national security applications. Stat, v. 14, issue 2, e70057, 2025.

FRONTIER MODEL FORUM. Issue brief: Early best practices for frontier ai safety evaluations. 2024. Available at: <https://www.frontiermodelforum.org/updates/early-best-practices-for-frontier-ai-safety-evaluations/>.

JIANG, Shuo; HU, Jie; MAGEE, Christopher L.; LUO, Jianxi. Deep learning for technical document classification. IEEE Transactions on Engineering Management, v. 71, p. 1163-1179, 2022.

KLEIN, L. K; EARL, E.; CUNDICK, D. Reducing information overload in your organization. Harvard Business Review, 2023.

KOSTINA, Arina; DIKAIAKOS, Marios D.; STEFANIDIS, Dimosthenis; PALLIS, George Pallis. Large language models for text classification: case study and comprehensive review. 2025. Available at: <https://arxiv.org/abs/2501.08457>.

KOWSARI, Kamram; JAFARI MEIMANDI, Kiana; HEIDARYSAFA, Mojtaba; MENDU, Sanjana; BARNES, Laura; BROWN, Donald. Text classification algorithms: A survey. Information, v. 10, issue 4, p. 150, 2019.

LI, Yinghao; RAMPRASAD, Rampi; ZHANG, Chao. A simple but effective approach to improve structured language model output for information extraction. 2024. Available at: <https://arxiv.org/abs/2402.13364>.

MAHONEY, Christian; GRONVALL, Peter; HUBER-FLIFLET; ZHANG, Jianping. Explainable text classification techniques in legal document review: locating rationales without using human annotated training text snippets. In: 2022 IEEE International Conference on Big Data (Big Data). IEEE, p. 2044-2051, 2022.

MCINTOSH, Timothy R., et al. Inadequacies of large language model benchmarks in the era of generative artificial intelligence. IEEE Transactions on Artificial Intelligence, 2025. Available at: <https://arxiv.org/abs/2402.09880>.

NATIONAL JUDICIAL STATISTICS SYSTEM (SNEJ). Annual Report: Technical Report. National Judicial Statistics System (SNEJ): Argentina, 2022.

PALANIVINAYAGAM, Ashokkumar; ZIAD EL-BAYEH, Claude; DAMAŠEVICIUS, Robertas. Twenty years of machine-learning-based text classification: A systematic review. Algorithms, v. 16, issue 5, p. 236, 2023.

PASKOV, Patricia; BERGLUND, Lukas; SMITH, Everett; SODER, Lisa. GPAI evaluations standards taskforce: towards effective AI governance. 2024. Available at: <https://arxiv.org/abs/2411.13808>.

PEÑA ALMANSA, Alejandro; et al. Leveraging large language models for topic classification in the domain of public affairs. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 2023.

PEREZ, Ethan; et al. Discovering language model behaviors with model-written evaluations. In: Findings of the Association for Computational Linguistics. Toronto: ACL, 2023. p. 13387–13434.

ROHERA, Pritika; GINIMAV, Chaitrali; SAWANT, Gayatri; JOSHI, Raviraj. Better to ask in english? Evaluating factual accuracy of multilingual LLMms in english and low-resource languages. Available at: <https://arxiv.org/abs/2504.20022>.

SAHOO, Pranab; SINGH, Ayush Kumar; SAHA, Sriparna; JAIN, Vinija; MONDAL, Samrat; CHADHA, Aman. A systematic survey of prompt engineering in large language models: techniques and applications. 2024. Available at: <https://arxiv.org/abs 2402.07927>.

SCHUT, Lisa; GAL, Yarin; FARQUHAR, Sebastian. Do multilingual LLMs think in english? 2025. Available at: <https://arxiv.org/abs/2502.15603>.

SCLAR, Melanie; CHOI, Yejin; TSVETKOV, Yulia; SUHR, Alane. Quantifying language models’ sensitivity to spurious features in prompt design or: how I learned to start worrying about prompt formatting. 2023. Available at: <https://arxiv.org/abs/2310.11324>.

SEGER, Elizabeth. What do we mean when we talk about “AI democratisation”. 2023. Available at: <https://www.governance.ai/analysis/what-do-we-mean-when-we-talk-about-ai-democratisation>.

SUN, Xiaofei; LI, Xiaoya; LI, Jiwei; WU, Fei; GUO, Shangwei; ZHANG, Tianwei; WANG, Guoyin. Text classification via large language models. 2023. Available at: <https://arxiv.org/abs/2305.08377>.

SUPREME COURT OF ARGENTINA. Statistics For The Year 2024: Technical report. Supreme Court of Argentina, 2024.

WAGH, Vedangi; KHANDVE, Snehal; JOSHI, Isha; WANI, Apurva; KALE, Geetanjali; JOSHI, Raviraj. Comparative study of long document classification. In: TENCON 2021-2021 IEEE Region 10 Conference (TENCON). 2021. p. 732–737.

WEIDINGER, Laura; et al. Sociotechnical safety evaluation of generative AI systems. 2023. Available at: <https://arxiv.org/abs/2310.11986>.

YE, Junjie; XU, Nuo; WANG, Yikun; ZHOU, Jie; ZHANG, Qi; GUI, Tao; HUANG, Xuanjing. LLM-DA: Data augmentation via large language models for few-shot named entity recognition. 2024. Available at: <https://arxiv.org/abs/2402.14568>.

YOON, Sungwook Yoon; et. al. Digital innovation in public administration through intelligent public sector automation (IPSA): Strategies and challenges. Journal of Multimedia Information System, v. 11, issue 4, p. 249-260, 2024.

ZHANG, Shuai; GU, Xiaodong; CHEN, Yuting; SHEN, Beijun. Infere: step-by-step regex generation via chain of inference. In: 38th IEEE/ACM International Conference on Automated Software Engineering (ASE). IEEE, 2023. p. 1505–1515.

Publicado

25.07.2026

Como Citar

CORVALÁN, Juan Gustavo et al. Predictive pattern analysis in administrative and tax litigation through generative AI and regular expressions. International Journal of Digital Law, Belo Horizonte, v. 7, p. e701, 2026. DOI: 10.47975/ijdl.v7.1305. Disponível em: https://journal.nuped.com.br/index.php/revista/article/view/1305. Acesso em: 25 jul. 2026.

Edição

Seção

Artigos originais

Categorias

Artigos Semelhantes

1 2 3 4 5 > >> 

Você também pode iniciar uma pesquisa avançada por similaridade para este artigo.