Domain Adaptation of Foundation Language Models for Legal Document Generation: A Systematic Review and Implications for Analogous Regulated Institutional Contexts
DOI:
https://doi.org/10.22481/recic.v8i1.19945Palavras-chave:
Large Language Models, Foundation Models, Legal NLP, Domain Adaptation, Parameter Efficient Fine-Tuning, Retrieval Augmented Generation, Systematic Literature ReviewResumo
The adaptation of foundation language models for legal document generation has become a relevant problem in Legal Natural Language Processing, particularly because generated texts must preserve factuality, normative coherence, traceability and control over sensitive information. This systematic literature review maps, classifies and critically analyzes primary studies on domain adaptation techniques for foundation language models applied to legal and normative-argumentative document generation. It also examines the implications of the resulting evidence for analogous regulated institutional contexts that share requirements related to reliability, privacy, auditability and institutional control. The review followed guidelines for systematic reviews in Software Engineering and the applicable items of PRISMA 2020. Searches were conducted in Scopus, Web of Science, IEEE Xplore and ACM Digital Library, covering studies published from January 2021 to December 2025 and complemented by backward snowballing. From 200 initial records, 32 primary studies were included. The corpus is predominantly concentrated in the legal domain, with only one study directly addressing document production in Public Security. The results show a recurrent use of hybrid adaptation strategies involving parameter-efficient fine-tuning, retrieval-augmented generation and structured inference, although the literature does not establish a universally superior combination. Data engineering is marked by a trade-off between scale and control, with frequent use of semi-synthetic datasets. Traditional Natural Language Processing metrics remain useful as initial indicators, but are insufficient to assess factuality, normative validity and argumentative consistency. Ethical safeguards, privacy controls and bias mitigation metrics remain underdeveloped. The findings indicate that the institutional adoption of foundation language models for legal document generation should be treated as a problem of architecture, evaluation and governance, rather than only textual performance. These conclusions are directly grounded in the Legal NLP literature, while their transfer to analogous regulated institutional contexts should be understood as a functionally motivated implication requiring domain-specific empirical validation.
Downloads
Referências
R. Al-Qaesm, M. Hendi and B. Tantour, “Alkafi-llama3: Fine-tuning LLMs for precise legal understanding in Palestine,” Discover Artificial Intelligence, vol. 5, no. 1, p. 107, 2025, doi: 10.1007/s44163-025-00313- w.
P. Colombo, T. Pires, M. Boudiaf, R. Melo, D. Culver, S. Morgado, E. Malaboeuf, G. Hautreux, J. Charpentier and M. Desa, “SaulLM-54B & SaulLM-141B: Scaling up domain adaptation for the legal domain,” Adv. Neural Inf. Process. Syst., vol. 37, 2024.
P. S. Garc´ıa-Montero, P. Vizca´ıno, I. G. Reyes-Chac´on and M. E. Morocho-Cayamcela, “Legal AI for all: Reducing perplexity and boosting accuracy in normative texts with fine-tuned LLMs and RAG,” IEEE Access, vol. 13, pp. 179759–179775, 2025, doi: 10.1109/ACCESS.2025.3622138.
N. T. Ha, T.-P. Nguyen, K. T. Trung, H.-L. Le, L. T. V. Huong, C. T. Nguyen and M.-T. Nguyen, “Vietnamese legal question answering: An experimental study,” in Proc. 2024 16th Int. Conf. Knowl. Syst. Eng. (KSE), 2024, pp. 440–446, doi: 10.1109/KSE63888.2024.11063637.
A. Halterman, “Synthetically generated text for supervised text analysis,” Political Analysis, vol. 33, no. 3, pp. 181–194, 2025, doi: 10.1017/pan.2024.31.
S. Ghosh, D. Verma, B. Ganesan, P. Bindal, V. Kumar and V. Bhatnagar, “InLegalLLaMA: Indian legal knowledge enhanced large language model,” CEUR Workshop Proc., vol. 3818, pp. 37–46, 2024.
H. Kim, D. Kim, J. Lee, C. Yoon, D. Choi, M. Gim and J. Kang, “LAPIS: Language model-augmented police investigation system,” in Proc. Int. Conf. Inf. Knowl. Manage., 2024, pp. 4637–4644, doi: 10.1145/3627673.3680044.
C. Liu, Y. Kang, S. Wang, L. Qing, F. Zhao, C. Wu, C. Sun, K. Kuang and F. Wu, “More than catastrophic forgetting: Integrating general capabilities for domain-specific LLMs,” in Proc. Conf. Empir. Methods Nat. Lang. Process. (EMNLP), 2024, pp. 7531–7548, doi: 10.18653/v1/2024.emnlp-main.429.
Z. Liu, Y. Zhu and M. Lu, “Enhancing legal expertise in large language models through composite model integration: The development and evaluation of Law-Neo,” in Proc. Nat. Leg. Lang. Process. Workshop (NLLP), 2024, pp. 33–41.
S. Huang, Y. Liu, J. Qi, H. Yang, Z. Luan and D. Qian, “Learning to follow domain-specific instruction with verifiable rewards,” in Proc. 2025 Int. Joint Conf. Neural Netw. (IJCNN), 2025, pp. 1–8, doi: 10.1109/IJCNN64981.2025.11228901.
H. D. Pimpale, A. Raut, Y. Patil, G. Parpol, P. Yadav and J. Sangoi, “LEGALMIND: A fine-tuned Gemma-2-based legal assistant for Indian judiciary with RAG and embedding integration,” Journal of Engineering and Technology for Industrial Applications, vol. 11, no. 55, pp. 105–117, 2025, doi: 10.5935/jetia.v11i55.1925.
D. Sinha and O. Sharma, “Generating legal arguments using LLM and vector database to support precedents,” in Proc. 2025 Int. Conf. Next Generation Inf. Syst. Eng. (NGISE), vol. 1, 2025, pp. 1–5, doi: 10.1109/NGISE64126.2025.11085307.
Y. Song, Y. Qin, R. Huang, Y. Chen and C. Lin, “Legal text summarization via judicial syllogism with large language models,” Journal of King Saud University Computer and Information Sciences, vol. 37, no. 5, 2025, doi: 10.1007/s44443-025-00113-3.
W. Su, B. Yue, Q. Ai, Y. Hu, J. Li, C. Wang, K. Zhang, Y. Wu and Y. Liu, “JuDGE: Benchmarking judgment document generation for Chinese legal system,” in Proc. 48th Int. ACM SIGIR Conf. Research and Development in Information Retrieval, 2025, pp. 3573–3583, doi: 10.1145/3726302.3730295.
N. Nagesh, Z. Wang and A. M. Rahmani, “FairCauseSyn: Towards causally fair LLM-augmented synthetic data generation,” in Proc. Annu. Int. Conf. IEEE Engineering in Medicine and Biology Society (EMBC), 2025, pp. 1–6, doi: 10.1109/EMBC58623.2025.11252705.
R. Ta, R. Salunke, R. R, R. Nv and S.R. Upadhyaya, “Fine-tuning a large language model for the Indian legal system,” in Proc. Int. Symp. INFOTEH-JAHORINA (INFOTEH), 2025, doi: 10.1109/INFOTEH64129.2025.10959207.
U. Ujwal, S. S. H. Surampudi, S. Mitra and T. Saha, “Reasoning before responding: Towards legal long-form question answering with interpretability,” in Proc. Int. Conf. Inf. Knowl. Manage., 2024, pp. 4922–4930, doi: 10.1145/3627673.3680082.
F. Valerio, P. Basile and M. de Gemmis, “Adapting a large language model to the legal domain: A case study in Italian,” CEUR Workshop Proc., vol. 3877, 2024.
N. Xie, Y. Bai, H. Gao, F. Fang, Q. Zhao, Z. Li, Z. Xue, L. Zhu, S. Ni and M. Yang, “DeliLaw: A Chinese legal counselling system based on a large language model,” in Proc. 33rd ACM Int. Conf. Inf. Knowl. Manage., 2024, pp. 5299–5303, doi: 10.1145/3627673.3679219.
Y. Yu, “Domain-specific legal language modeling through knowledge fusion and structured attention,” in Proc. Int. Conf. Electr. Inf. Technol. Comput. Eng. (EITCE), 2025, pp. 685–689, doi: 10.1145/3766671.3766790.
H. Le, N. Luu, T. Nguyen, T. Dao and S. Dinh, “Optimizing answer generator in Vietnamese legal question answering systems using language models,” ACM Trans. Asian Low-Resource Lang. Inf. Process., vol. 24, no. 6, 2025, doi: 10.1145/3732938.
D. Shu, H. Zhao, X. Liu, D. Demeter, M. Du and Y. Zhang, “LawLLM: Law large language model for the US legal system,” in Proc. 33rd ACM Int. Conf. Inf. Knowl. Manage., 2024, pp. 4882–4889, doi: 10.1145/3627673.3680020.
S. Yao, Q. Ke, Q. Wang, K. Li and J. Hu, “Lawyer GPT: A legal large language model with enhanced domain knowledge and reasoning capabilities,” in Proc. 2024 3rd Int. Symp. Robotics, Artificial Intelligence and Information Engineering, 2024, pp. 108–112, doi: 10.1145/3689299.3689319.
J. Cui, Z. Li, Y. Yan, B. Chen and L. Yuan, “Chatlaw: A multi-agent collaborative legal assistant with knowledge graph enhanced mixture-of-experts large language model,” 2023.
Z. Zhou, J.-X. Shi, P.-X. Song, X.-W. Yang, Y.-X. Jin, L.-Z. Guo and Y.-F. Li, “LawGPT: A Chinese legal knowledge-enhanced large language model,” arXiv preprint arXiv:2406.04614, 2024.
M. Dahl, V. Magesh, M. Suzgun and D. E. Ho, “Large legal fictions: Profiling legal hallucinations in large language models,” Journal of Legal Analysis, vol. 16, no. 1, pp. 64–93, 2024, doi: 10.1093/jla/laae003.
C. Jiang and X. Yang, “Legal syllogism prompting: Teaching large language models for legal judgment prediction,” in Proc. Nineteenth Int. Conf. Artificial Intelligence and Law, 2023, pp. 417–421, doi: 10.1145/3594536.3595170.
F. Yu, L. Quartey and F. Schilder, “Legal prompting: Teaching a language model to think like a lawyer,” in Proc. 2022 Conf. Empir. Methods Nat. Lang. Process. (EMNLP), 2022, doi: 10.48448/sdt7-nc75.
J. Cui, Z. Li, Y. Yan, B. Chen and L. Yuan, “ChatLaw: Open-source legal large language model with integrated external knowledge bases”, arXiv preprint arXiv:2306.16092, 2023.
W. Deng, J. Pei, K. Kong, Z. Chen, F. Wei, Y. Li, Z. Ren, Z. Chen and P. Ren, “Syllogistic reasoning for legal judgment analysis,” in Proc. 2023 Conf. Empir. Methods Nat. Lang. Process. (EMNLP), 2023, pp. 10283–10298.
C. Xiao, X. Hu, Z. Liu, C. Tu and M. Sun, “Lawformer: A pre-trained language model for Chinese legal long documents,” AI Open, vol. 2, pp. 79–84, 2021, doi: 10.1016/j.aiopen.2021.06.003.
N. Guha, D. E. Ho, J. Nyarko and C. R´e, “LegalBench: Prototyping a collaborative benchmark for legal reasoning,” arXiv preprint arXiv:2209.06120, 2022.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser and I. Polosukhin, “Attention is all you need,” in Proc. 31st Int. Conf. Neural Information Processing Systems, 2017, pp. 6000–6010.
J. Crawford, “Systematic literature reviews: Why I rejected your review,” Journal of University Teaching & Learning Practice, vol. 22, no. 2, 2025, doi: 10.53761/10vb5076.
H. Touvron et al., “LLaMA: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023.
E. J. Hu et al., “LoRA: Low-rank adaptation of large language models,” in Proc. Int. Conf. Learning Representations, 2022.
T. Dettmers, A. Pagnoni, A. Holtzman and L. Zettlemoyer, “QLoRA: Efficient finetuning of quantized LLMs,” in Proc. 37th Int. Conf. Neural Information Processing Systems, 2023.
J. Popay, H. Roberts, A. Sowden, M. Petticrew, L. Arai, M. Rodgers, et al., Guidance on the Conduct of Narrative Synthesis in Systematic Reviews. ESRC Methods Programme, 2006.
J. P. T. Higgins, J. Thomas, J. Chandler, M. Cumpston, T. Li, M. J. Page, et al., Cochrane Handbook for Systematic Reviews of Interventions, 2nd ed. Hoboken, NJ: John Wiley & Sons, 2019, doi: 10.1002/9781119536604.
B. Kitchenham and S. Charters, “Guidelines for performing systematic literature reviews in software engineering,” Keele University and Durham University, Tech. Rep. EBSE-2007-01, 2007.
P. Lewis et al., “retrieval augmented generation for knowledge-intensive NLP tasks,” Adv. Neural Inf. Process. Syst., vol. 33, 2020.
M. J. Page, J. E. McKenzie, P. M. Bossuyt, I. Boutron, T. C. Hoffmann, C. D. Mulrow, et al., “The PRISMA 2020 statement: An updated guideline for reporting systematic reviews,” BMJ, vol. 372, Art. no. n71, 2021, doi: 10.1136/bmj.n71.
M. Ouzzani, H. Hammady, Z. Fedorowicz and A. Elmagarmid, “Rayyan: A web and mobile app for systematic reviews,” Systematic Reviews, vol. 5, no. 210, 2016, doi: 10.1186/s13643-016-0384-4.
R. Rafailov et al., “Direct preference optimization: Your language model is secretly a reward model,” Adv. Neural Inf. Process. Syst., 2023.
C. Wohlin, “Guidelines for snowballing in systematic literature studies and a replication in software engineering,” in Proc. 18th Int. Conf. Evaluation and Assessment in Software Engineering, 2014, doi: 10.1145/2601248.2601268.
A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, et al., “Qwen3 technical report,” arXiv preprint arXiv:2505.09388, 2025.
J. Lai, W. Gan, J. Wu, Z. Qi and P. S. Yu, “Large language models in law: A survey,” AI Open, vol. 5, pp. 181–196, 2024, doi: 10.1016/j.aiopen.2024.09.002.
Z. Hou, Z. Ye, N. Zeng, T. Hao and K. Zeng, “Large language models meet legal artificial intelligence: A survey,” arXiv preprint arXiv:2509.09969, 2025.
F. Dehghani, R. Dehghani, Y. Naderzadeh Ardebili and S. Rahnamayan, “Large language models in legal systems: A survey,” Humanities and Social Sciences Communications, vol. 12, no. 1, Art. no. 1977, 2025, doi: 10.1057/s41599-025-05924-3
Downloads
Publicado
Como Citar
Edição
Seção
Licença
Copyright (c) 2026 Revista de Ciência da Computação

Este trabalho está licenciado sob uma licença Creative Commons Attribution 4.0 International License.