نوع مقاله : علمی-پژوهشی
عنوان مقاله English
نویسندگان English
The rise of digital communication tools, such as social networks and messaging platforms, has led to an expansion of diverse writing styles in Persian, including colloquial and non-standard (broken) forms. Converting such texts into standard Persian writing, especially based on the latest orthographic guidelines of the Academy of Persian Language and Literature, requires advanced natural language processing tools. This study compares the performance of five large language models, namely ChatGPT, Gemini, Perplexity, Claude, and DeepSeek, by automatically converting colloquial Persian into standard written form. The dataset utilized in this study was compiled and corrected from two major corpora. A sub-corpus has been created through random sampling and resulted in 1,025 sentences and 11,939 tokens. Outputs from the models were compared to gold-standard annotations, focusing on both positive correction (e.g., correct use of half-spaces, Ezafe, hamza, and standardizing informal forms) and negative errors (e.g., word replacement, deletion of clitics, and stylistic shifts). Findings indicate that the language model in Claude achieved the highest alignment with the gold-standard (51.39%) at word and sentence level, while Perplexity performed the worst. The results highlight that, despite their strengths, large language models still face challenges in accurately standardizing Persian text in line with formal orthographic conventions.
کلیدواژهها English