This study examines large language models in high-stakes Arabic-to-English legal translation through an evaluation of ChatGPT 5.2 and DeepSeek-V3. Focusing on Palestinian criminal judgments, the research tests how models translate specialized terminology. From a 120,000-word corpus, 100 recurrent terms were extracted based on their doctrinal significance and procedural function. Context was operationalized by providing the models with the complete authentic sentence in which each target term was originally embedded, rather than isolated vocabulary lists. Both models translated the identical dataset under a zero-shot, single-output protocol. Two accredited translators evaluated the semantic fidelity and institutional validity of the translated terms using a three-point accuracy scale. This quantitative scoring was directly combined with qualitative error coding to categorize the specific types of legal translation failures, such as conceptual distortion or literalism. Inter-rater reliability was established via Cohen’s Kappa, and differences were analyzed using Wilcoxon signed-rank testing. Results indicate statistically significant performance differences. ChatGPT 5.2 achieved higher overall accuracy without critical errors, while DeepSeek generated notably more conceptually distorted outputs. The findings demonstrate that when legal terms are processed within complete sentences, the models display distinct operational vulnerabilities – such as algorithmic paraphrasing or structural literalism – and that model performance differs considerably across systems.
