The Use of Artificial Intelligence in the Rendition of Certain Articles of Ḥanbalī Sharīʿa Code into English: A Comparison of Automatic and Human Evaluation of Translation Output Quality of ChatGPT and Gemini in Arabic–English Legal Translation
摘要
Based on Alwazna [1], the present paper seeks to assess the translation output quality of both ChatGPT and Gemini using certain articles of Ḥanbalī Sharīʿa Code through both automatic evaluation with its TER, BLEU and ChrF3 metrics and human evaluation with its adequacy and fluency criteria. It carries out automatic evaluation followed by human evaluation to test the reliability of the former’s results in evaluating the two AI-generated translations and whether or not such AI-generated translations are reliable within the Arabic–English legal translation context. The present paper argues that AI tools including ChatGPT and Gemini have made a considerable advancement in Arabic–English legal translation, though they still encounter certain challenges, particularly in the rendition of legal terms. The paper also claims that the use of automatic evaluation metrics cannot be depended on individually in assessing the translation output quality of AI tools with regard to Arabic–English legal translation. However, human evaluation should be carried out after undertaking automatic evaluation to test whether or not the scores of automatic evaluation are reliable. This paper offers a baseline for assessing the translation output quality of AI tools in Arabic–English legal translation, which may have implications for other similar language-pair contexts.