Large language models in high-stakes Arabic-to-English legal translation: a contextual terminology study of ChatGPT 5.2 and DeepSeek
Publication Type
Original research
Authors

This study examines large language models in high-stakes Arabic-to-English legal translation through an evaluation of ChatGPT 5.2 and DeepSeek-V3. Focusing on Palestinian criminal judgments, the research tests how models translate specialized terminology. From a 120,000-word corpus, 100 recurrent terms were extracted based on their doctrinal significance and procedural function. Context was operationalized by providing the models with the complete authentic sentence in which each target term was originally embedded, rather than isolated vocabulary lists. Both models translated the identical dataset under a zero-shot, single-output protocol. Two accredited translators evaluated the semantic fidelity and institutional validity of the translated terms using a three-point accuracy scale. This quantitative scoring was directly combined with qualitative error coding to categorize the specific types of legal translation failures, such as conceptual distortion or literalism. Inter-rater reliability was established via Cohen’s Kappa, and differences were analyzed using Wilcoxon signed-rank testing. Results indicate statistically significant performance differences. ChatGPT 5.2 achieved higher overall accuracy without critical errors, while DeepSeek generated notably more conceptually distorted outputs. The findings demonstrate that when legal terms are processed within complete sentences, the models display distinct operational vulnerabilities – such as algorithmic paraphrasing or structural literalism – and that model performance differs considerably across systems.

Journal
Title
Naeem Salameh
Publisher
De Gruyter Brill
Publisher Country
Germany
Indexing
Thomson Reuters
Impact Factor
2.5
Publication Type
Prtinted only
Volume
--
Year
2026
Pages
--