نوع مقاله : علمی-پژوهشی
عنوان مقاله English
نویسندگان English
Abstract
The rise of digital communication tools, such as social networks and messaging platforms, has led to an expansion of diverse writing styles in Persian, including colloquial and non-standard (broken) forms. Converting such texts into standard Persian writing, especially based on the latest orthographic guidelines of the Academy of Persian Language and Literature, requires advanced natural language processing tools. This study compares the performance of five large language models, namely ChatGPT, Gemini, Perplexity, Claude, and DeepSeek, by automatically converting colloquial Persian into standard written form. The dataset utilized in this study is compiled and correct from two major corpora. A sub-corpus has been created through random sampling and resulted in 1,025 sentences and 11,939 tokens. Outputs from the models are compared to the gold standard annotations, focusing on both positive correction (e.g., correct use of pseudo-spaces, Ezafe, hamza, and standardizing informal forms) and negative errors (e.g., word replacement, deletion of clitics, and stylistic shifts). Findings indicate that the language model in Claude achieved the highest alignment with the gold standard (51.39%) at word and sentence levels, while Perplexity performed the worst. The results highlight that, despite their strengths, large language models still face challenges in accurately standardizing Persian text based on the formal orthographic conventions.
Keywords: colloquial Persian, non-standard (broken) writing, standardization, large language models, Persian orthography
1. Introduction
The widespread adoption of digital communication technologies, such as messaging applications and social media platforms, has significantly expanded public participation in written communication. As a result, diverse forms of writing have emerged, reflecting the linguistic practices of users from various educational, social, and cultural backgrounds. In Persian, this trend has led to the increasing use of non-standard and colloquial writing styles that closely mirror everyday speech rather than established orthographic conventions. While these forms demonstrate the dynamism and creativity of language in digital environments and provide valuable material for linguistic research, they pose challenges when the conversion of such texts into standardized writing is required. Therefore, the normalization of non-standard texts according to the official orthographic guidelines established by the Academy of Persian Language and Literature is of both practical and scholarly importance.
The emergence of Large Language Models (LLM) since 2020 has provided new possibilities for processing and standardizing Persian writing. These models, relying on deep learning and extensive text data, have the ability to convert non-standard and colloquial writing styles into standard writing. This paper contributes to convert non-standard spelling varieties into the standard ones by using LLM tools and to measure the degree to which the LLMs’ output conforms to formal writing. To this end, the present study evaluates the performance of five LLM tools, including ChatGPT, Gemini, Perplexity, Claude, and DeepSeek. The conversion is compared with the latest version of orthographic guidelines of the Academy of Persian Language and Literature (2023). The results of this research can be effective in understanding the capacities and limitations of LLMs and evaluating their role in standardizing Persian writing in the digital space.
Research Question(s)
How the performance of the LLM tools is comparable to convert a colloquial and non-standard texts into standard ones?
2. Literature Review
There are a number of domestic and international research on the automatic conversion of colloquial or non-standard texts into standard ones. To this end, investigations show that sequential machine translation (Mansfield et al., 2019; Arora and Kansal, 2020), deep neural models, rule-based methods (Ariffin & Tiun, 2020), and statistical approaches are the most important solutions used in this field. International studies in English, Malay (Arora & Kansal, 2020), Kazakh (Kozhirbayev & Yessenbayev, 2020) and medical texts (Liu et al., 2021) have reported that the use of sub-words, encoder-decoder networks, BERT-based models, and combining rules with machine learning can significantly increase the accuracy of standardization and improve the performance of natural language processing systems, including sentiment analysis (Arora & Kansal, 2020) and grammatical labeling (Ariffin & Tiun, 2020).
In Persian, from the rule-based methods (Armin and Shamsfard, 2010) to transformer-based deep learning models (Rasooli et al., 2020; AbdiKhojasteh et al., 2020; Adibian & Momtazi, 2022), extensive efforts have been made to convert colloquial and non-standard texts into the standard ones. In addition to designing automatic converters, such as a spell-checker (Tajalli et al., 2023), these studies have focused on collecting parallel formal-informal corpora, using crowdsourcing to generate data (Masoumi et al., 2020) and analyze the features of rewriting in social networks (Ghayoomi & Mesgarkhoyi, 2023). The results show that a large part of cyberspace words are compatible with the standard spelling, but broken writing remains as a serious challenge. Overall, the available evidence suggests that combining large data sets, standard script, and modern deep learning models can provide an effective path for the automatic standardization of Persian writing and the improvement of natural language processing tools.
3. Methodology
To achieve the research goals, the web version of the LLM tools, including ChatGPT, Gemini, Perplexity, Claude, and DeepSeek, are used to exploit their latest features and capabilities. The input of the tools is non-standard spelling of Persian and it is expected to be converted into standard spelling rules by the tools. To evaluate the performance, the gold standard labels of the input is required. To this end, two Persian corpora containing colloquial and non-standard Persian texts, prepared by Masoumi et al. (2020) and Tajali et al. (2023), are used. The first corpus contains 4500 sentences and the second one contains 50,011 sentences. By concatenating the two corpora, a corpus with 54,511 sentences and 584,418 words is obtained. Due to the limitations of using web versions of LLM tools, 1025 sentences are randomly selected from the combined corpus, which includes 11,939 tokens and 4016 type words. After comparing the output of the input with the gold standard labels, the changes, either corrections or mistakes, are manually categorized by experts into “positive” or “negative” categories. The “positive” performance includes:
1) Inserting the Ezafe clitic as “ی” /ye/, such as “خانه ما” /xāne.ye ma/ ‘our home’ which becomes “خانهی ما”;
2) Inserting the Ezafe clitic as “ ۀ ” /ye/, such as “خانه ما” /xāne.ye ma/ ‘our home’ which becomes “خانۀ ما”;
3) Inserting the pseudo-space, such as “ساقه هایش” /sāqe hāyaš/ ‘its stalks’ which becomes “ساقههایش”;
4) Correctly inserting the hamza, such as “تاثیر” /ta?sir/ ‘impact’ which becomes “تأثیر”;
5) Correctly inserting the tanvin, such as “کلا” /kollan/ ‘generally’ which becomes “کلاً”;
6) Correctly converting the colloquial and non-standard writing into the standard form, such as “دستتون” /dastetun/ ‘your hand’ which becomes “دستتان” /dastetān/;
7) Correctly inserting the mediating consonants, such as “توش” /tuš/ ‘in it’ which becomes “تویش” /tuyaš/;
8) Correcting the spelling or unintentional errors in writing words, such as “ینی” /yani/ ‘meaning’, which becomes “یعنی”;
9) Substituting the standard spelling for fancy writing, such as “عسیس” /?asis/ ‘dear’, which becomes “عزیز” /?aziz/.
The negative changes in the model output that are not consistent with the gold standard labels are as follow:
1) Changing the word order in a sentence, such as the sentence “بیا دوست باشیم دیگر” /biyā dust bāšim digar/ ‘Let's be friends’ which can be changed to “بیا دیگر دوست باشیم”;
2) Replacing a word, such as changing “مگر” /mage/ ‘unless’ to “آیا”
/?āyā/ [the question word for yes/no questions];
3) Deleting clitics and replacing them with the corresponding simple word, such as “خودشانند” /xodešānand/ ‘they are themselves’, which can be changed to “خودشان هستند” /xodešān hastand/;
4) Not converting colloquial or non-standard texts into the standard one, such as “یه” /ye/ ‘one’, whose correct form is “یک” /yek/ but without a conversion, no changes have been done;
5) Incorrect insertion of Ezafe, such as “مجسمههای” /moǰassamehāye/ ‘statues’ which is incorrectly changed to “مجسمۀهای”;
6) Incorrect insertion of pseudo-space, such as “برایتان” /barāyetān/ ‘for you’ while the correct writing is “برایتان”;
7) Deleting a word, such as “ماروپلهاش رو بازی کنه” /māropelle?as ro bāzi kone/ ‘he/she plays his snake and ladder game’ which can be converted “ماروپلهاش بازی کنه”. In this example, “رو” /ro/ which is the correct form of “را” /rā/, is deleted from the sentence;
8) Inserting a new word, such as “میخواستم بروم داخل” /mixāstam beravam dāxel/ ‘I wanted to enter’ which is converted into “میخواستم بروم به داخل” /mixāstam beravam be dāxel/ ‘I wanted to enter to’. In the conversion the preposition “به” /be/ ‘to’ is added;
9) Inserting a space incorrectly, such as “اون که” /?un ke/ ‘one that’ which is changed to “آنکه” /?ānke/ while the correct form is ‘آن که’;
10) No insertion of a proper space, such as “توهم” /toham/ ‘you too’, while the correct form is “تو هم” /to ham/;
11) Replacing the wrong word, such as “تو” /tu/ ‘in’ in the example “تو خانه” ‘at home’ which is converted into “در خانه” /dar xāne/
In general, any change that causes the text to deviate from the expected standard and correct form falls into this category of negative errors.
4. Results
According to the obtained results in Table 1, GPT and Perplexity performed the most and the least positive and negative impacts at the word level, respectively. The most positive impact of the language models had been made by Claude, and Perplexity performed the worst. DeepSeek, GPT, and Gemini stay in the second to fourth ranks, respectively. Among the positive changes, inserting a pseudo-space had the most impact in converting non-standard text to standard. Inserting the Ezafe clitic applied the least change to non-standard data.
Furthermore, according to the negative impacts of the LLM tools on the data, Perplexity and GPT applied the least and the most negative impact, respectively. Among the negative changes, the conversion of colloquial or broken writing to standard writing had the most negative changes, such that Perplexity had the least and Claude had the most changes. Inserting a wrong hamza had the least negative effect on data conversion. In this category, DeepSeek and Gemini did not have any negative effect on data conversion. Word substitution was the most negative change applied by GPT.
Table 1
distribution of positive and negative changes of LLMs in sample data at the word level
GPT
Gemini
Perplexity
Claude
DeepSeek
Positive
12.36
11.49
6.48
15.40
14.44
Negative
34.13
17.50
14.12
23.13
22.48
Next, we examined the performance of the LLM tools against the gold standard data in terms of exact matching at the sentence level for the input and output data. The results are reported in Table1. Based on the results, Claude obtained the highest matching rate (%51.39) and Perplexity the lowest matching rate (%90.7) with the gold standard data at the sentence level. The other language models, namely DeepSeek, Gemini, and GPT, ranked second to fourth respectively among the models. In examining LLMs with each other, without considering the gold standard data, it was concluded that Claude and DeepSeek perform similarly, and GPT and Perplexity had the lowest performance.
Table 2
comparing performance of LLMs with gold standard data at the sentence level
Gold
GPT
Gemini
Perplexity
Claude
DeepSeek
GPT
98
Gemini
244
84
Perplexity
81
47
183
Claude
405
136
293
97
DeepSeek
337
118
231
108
368
In the further studying the performances of LLMs, the models were compared from the perspective of syntax. First, the sentences obtained from LLMs and the gold standard data were part-of-speech tagged. Then, the content and functional words were separated from each other and the performances of the models were compared. The gold standard data contained %50.76 of content words and %50.23 of functional words. The results of the analysis are reported in Table 3. Based on the results obtained, Claude and Perplexity were able to obtain the largest and smallest volume of content and functional words, respectively. The ratio of content word changes in GPT (70.29%) was higher than other models, while Claude had decreased this ratio to %67.73. The ratio of functional word change in comparing GPT (%29.71) and Claude (%32.27) was the opposite such that Claude had a better performance in converting non-standard functional words into standard ones.
Table 3
distribution of content and functional words in LLMs compared to gold standard data
Content word
Functional word
GPT
4.85
1.44
Gemini
14.33
4.42
Perplexity
4.00
1.32
Claude
25.23
8.14
DeepSeek
21.15
6.84
5. Conclusion
The general conclusion that can be drawn from the present study is that the difference in the results obtained from different LLMs indicates that due to the difference in their algorithms and training data of the models, the difference in their performance is obvious. Therefore, it is not possible to choose a single LLM and limit a research to that model. In a comprehensive study, it is necessary to examine various LLMs. Examining the differences can be a guide to choose the right tool in future research.
Bibliography
AbdiKhojasteh, H., Ansari, E., & Bohlouli, M. (2020). LSCP: Enhanced large scale colloquial Persian language understanding. Proceedings of the 12th International Conference on Language Resources and Evaluation, Marseille, France: European Language Resources Association, pp. 6323–6327.
Academy of Persian Language and Literature (2023). The Approved Orthography Rules of the Academy of Persian Language and Literature. Tehran: Persian Language and Literature Academy. [In Persian]
Adibian, M., & Momtazi, S. (2012). Converting Persian colloquial text to formal text using neural networks based on converters. Journal of Language and Linguistics, 18(35), 45–67. [In Persian]
Armin, N., & Shamsfard, M. (2010). Converting Persian colloquial text to formal text using n-grams. Proceedings of the 16th Annual National Conference of the Iranian Computer Association. [In Persian]
Ariffin, S. N. A. N., & Tiun, S. (2020). Rule-based text normalization for Malay social media texts. International Journal of Advanced Computer Science and Applications, 11.
Arora, M., & Kansal, V. (2020). Character level embedding with deep convolutional neural network for text normalization of unstructured data for Twitter sentiment analysis. International Journal of Advanced Computer Science and Applications, 11.
Ghayoomi, M., & Mesgarkhoyi, M. (2023). Analysis of the corpus of Persian language data in cyberspace. Language and Linguistics, 19, 37: 117–136. [In Persian]
Kozhirbayev, Z., & Yessenbayev, Z. (2020). Kazakh text normalization using machine translation approaches. CEUR Workshop Proceedings, 2780.
Liu, Y., Ji, B., Yu, J., Tan, Y., Ma, J., & Wu, Q. (2021). An advanced ICD-9 terminology standardization method based on BERT and text similarity. Advances in Natural Computation, Fuzzy Systems and Knowledge Discovery. ICNC-FSKD 2020. Lecture Notes on Data Engineering and Communications Technologies, eds. Meng, H., Lei, T., Li, M., Li, K., Xiong, N., & Wang, L., Speringer, vol. 88, pp. 1868–1879.
Mansfield, C., Sun, M., Liu, Y., Gandhe, A., & Hoffmeister, B. (2019). Neural text normalization with subword units. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, eds. Loukina, A., Morales, M., & Kumar, R., Minneapolis, Minnesota: Association for Computational Linguistics, vol. 2, pp. 190–196.
Masoumi, V., Salehi, M., Veisi, H., Haddadian, G., Ranjbar, V., & Sahebdel, M. (2020). TeleCrowd: A crowdsourcing approach to create informal to formal text corpora. Arxiv.
https://arxiv.org/abs/2004.11771.
Rasooli, M. S., Bakhtyari, F., Shafiei, F., Ravanbakhsh, M., & Callison-Burch, C. (2020). Automatic standardization of colloquial Persian. Arxiv. https://arxiv.org/abs/2012.05879
Tajalli, V., Kalantari, F., & Shamsfard, M. (2023). Developing an informal formal Persian corpus. Arxiv. https://arxiv.org/abs/2308.05336.
کلیدواژهها English