تحلیلی بر عملکرد مدل‌های زبانی بزرگ در تبدیل خودکار گفتاری‌نویسی به نوشتار معیار براساس دستور خط مصوب فارسی

نوع مقاله : علمی-پژوهشی

نویسندگان
1 پژوهشکده زبان‌شناسی، پژوهشگاه علوم انسانی و مطالعات فرهنگی
2 پژوهشکده زبانشناسی، پژوهشگاه علوم انسانی و مطالعات فرهنگی
3 دانشکده کامپیوتر، دانشگاه صنعتی امیرکبیر
10.30465/lsi.2026.52503.1820
چکیده
رشد ارتباطات دیجیتال، مانند شبکه‌های اجتماعی و پیام‌رسان‌ها، موجب گسترش شکل‌های متنوعی از نوشتار در زبان فارسی، ازجمله گفتاری‌نویسی و شکسته‌نویسی، شده‌است. تبدیل این گونه‌های نوشتاری به نوشتار معیار، به‌ویژه براساس دستور خط مصوب فرهنگستان زبان و ادب فارسی، نیازمند بهره‌گیری از ابزارهای پیشرفته پردازش زبان طبیعی است. در پژوهش حاضر، عملکرد پنج مدل زبانی بزرگ شامل ChatGPT، Gemini، Perplexity، Claudeو DeepSeek در تبدیل خودکار گفتاری‌نویسی فارسی به نوشتار معیار براساس دستور خط مصوب ارزیابی شده‌است. داده‌های پژوهش از اصلاح و ترکیب دو پیکرۀ حاصل از فضای مجازی حاصل شده و به‌صورت تصادفی یک زیرپیکره شامل ۱۰۲۵ جمله که حاوی ۱۱۹۳۹ واژه است تهیه شده‌است. تحلیل خروجی مدل‌ها بر مبنای برچسب‌های طلایی معیار و با تمرکز بر تغییرات مثبت (مانند درج نیم‌فاصله، کسرۀ اضافه، همزه و اصلاح شکسته‌نویسی) و تغییرات منفی (مانند حذف واژه‌بست، جایگزینی واژه و تغییر سبک و جابه‌جایی واژه‌ها) انجام شده‌است. نتایج نشان می‌دهد مدل Claude با بیشترین تطابق با برچسب طلایی (51/39٪) در سطح واژه و جمله عملکرد بهتر و Perplexity ضعیف‌ترین عملکرد را از این منظر داشته‌است. این پژوهش نشان می‌دهد که مدل‌های زبانی بزرگ، علی‌رغم توانایی بالا، همچنان با چالش‌هایی در معیارسازی نوشتار فارسی، به‌ویژه در رعایت دقیق دستور خط، مواجه‌است.
کلیدواژه‌ها

عنوان مقاله English

Analyzing the Performance of Large Language Models in the Automatic Conversion of Colloquial Persian to Standard Writing based on the Confirmed Persian Orthography Grammar

نویسندگان English

Masood Ghayoomi 1
Fatemeh Mohammadi 2
Ali Jahan 3
1 Faculty of Linguistics, Institute for Humanities and Cultural Studies
2 Faculty of Linguistics, Institute for Humanities and Cultural Studies
3 Department of Computer Engineering, Amirkabir University of Technology
چکیده English

Abstract
The rise of digital communication tools, such as social networks and messaging platforms, has led to an expansion of diverse writing styles in Persian, including colloquial and non-standard (broken) forms. Converting such texts into standard Persian writing, especially based on the latest orthographic guidelines of the Academy of Persian Language and Literature, requires advanced natural language processing tools. This study compares the performance of five large language models, namely ChatGPT, Gemini, Perplexity, Claude, and DeepSeek, by automatically converting colloquial Persian into standard written form. The dataset utilized in this study is compiled and correct from two major corpora. A sub-corpus has been created through random sampling and resulted in 1,025 sentences and 11,939 tokens. Outputs from the models are compared to the gold standard annotations, focusing on both positive correction (e.g., correct use of pseudo-spaces, Ezafe, hamza, and standardizing informal forms) and negative errors (e.g., word replacement, deletion of clitics, and stylistic shifts). Findings indicate that the language model in Claude achieved the highest alignment with the gold standard (51.39%) at word and sentence levels, while Perplexity performed the worst. The results highlight that, despite their strengths, large language models still face challenges in accurately standardizing Persian text based on the formal orthographic conventions.
Keywords: colloquial Persian, non-standard (broken) writing, standardization, large language models, Persian orthography
1. Introduction
The widespread adoption of digital communication technologies, such as messaging applications and social media platforms, has significantly expanded public participation in written communication. As a result, diverse forms of writing have emerged, reflecting the linguistic practices of users from various educational, social, and cultural backgrounds. In Persian, this trend has led to the increasing use of non-standard and colloquial writing styles that closely mirror everyday speech rather than established orthographic conventions. While these forms demonstrate the dynamism and creativity of language in digital environments and provide valuable material for linguistic research, they pose challenges when the conversion of such texts into standardized writing is required. Therefore, the normalization of non-standard texts according to the official orthographic guidelines established by the Academy of Persian Language and Literature is of both practical and scholarly importance.
The emergence of Large Language Models (LLM) since 2020 has provided new possibilities for processing and standardizing Persian writing. These models, relying on deep learning and extensive text data, have the ability to convert non-standard and colloquial writing styles into standard writing. This paper contributes to convert non-standard spelling varieties into the standard ones by using LLM tools and to measure the degree to which the LLMs’ output conforms to formal writing. To this end, the present study evaluates the performance of five LLM tools, including ChatGPT, Gemini, Perplexity, Claude, and DeepSeek. The conversion is compared with the latest version of orthographic guidelines of the Academy of Persian Language and Literature (2023). The results of this research can be effective in understanding the capacities and limitations of LLMs and evaluating their role in standardizing Persian writing in the digital space.
Research Question(s)
How the performance of the LLM tools is comparable to convert a colloquial and non-standard texts into standard ones?
2. Literature Review
There are a number of domestic and international research on the automatic conversion of colloquial or non-standard texts into standard ones. To this end, investigations show that sequential machine translation (Mansfield et al., 2019; Arora and Kansal, 2020), deep neural models, rule-based methods (Ariffin & Tiun, 2020), and statistical approaches are the most important solutions used in this field. International studies in English, Malay (Arora & Kansal, 2020), Kazakh (Kozhirbayev & Yessenbayev, 2020) and medical texts (Liu et al., 2021) have reported that the use of sub-words, encoder-decoder networks, BERT-based models, and combining rules with machine learning can significantly increase the accuracy of standardization and improve the performance of natural language processing systems, including sentiment analysis (Arora & Kansal, 2020) and grammatical labeling (Ariffin & Tiun, 2020).
In Persian, from the rule-based methods (Armin and Shamsfard, 2010) to transformer-based deep learning models (Rasooli et al., 2020; AbdiKhojasteh et al., 2020; Adibian & Momtazi, 2022), extensive efforts have been made to convert colloquial and non-standard texts into the standard ones. In addition to designing automatic converters, such as a spell-checker (Tajalli et al., 2023), these studies have focused on collecting parallel formal-informal corpora, using crowdsourcing to generate data (Masoumi et al., 2020) and analyze the features of rewriting in social networks (Ghayoomi & Mesgarkhoyi, 2023). The results show that a large part of cyberspace words are compatible with the standard spelling, but broken writing remains as a serious challenge. Overall, the available evidence suggests that combining large data sets, standard script, and modern deep learning models can provide an effective path for the automatic standardization of Persian writing and the improvement of natural language processing tools.
3. Methodology
To achieve the research goals, the web version of the LLM tools, including ChatGPT, Gemini, Perplexity, Claude, and DeepSeek, are used to exploit their latest features and capabilities. The input of the tools is non-standard spelling of Persian and it is expected to be converted into standard spelling rules by the tools. To evaluate the performance, the gold standard labels of the input is required. To this end, two Persian corpora containing colloquial and non-standard Persian texts, prepared by Masoumi et al. (2020) and Tajali et al. (2023), are used. The first corpus contains 4500 sentences and the second one contains 50,011 sentences. By concatenating the two corpora, a corpus with 54,511 sentences and 584,418 words is obtained. Due to the limitations of using web versions of LLM tools, 1025 sentences are randomly selected from the combined corpus, which includes 11,939 tokens and 4016 type words. After comparing the output of the input with the gold standard labels, the changes, either corrections or mistakes, are manually categorized by experts into “positive” or “negative” categories. The “positive” performance includes:
1) Inserting the Ezafe clitic as “ی” /ye/, such as “خانه ما” /xāne.ye ma/ ‘our home’ which becomes “خانه‌ی ما”;
2) Inserting the Ezafe clitic as “ ۀ ” /ye/, such as “خانه ما” /xāne.ye ma/ ‘our home’ which becomes “خانۀ ما”;
3) Inserting the pseudo-space, such as “ساقه هایش” /sāqe hāyaš/ ‘its stalks’ which becomes “ساقه‌هایش”;
4) Correctly inserting the hamza, such as “تاثیر” /ta?sir/ ‘impact’ which becomes “تأثیر”;
5) Correctly inserting the tanvin, such as “کلا” /kollan/ ‘generally’ which becomes “کلاً”;
6) Correctly converting the colloquial and non-standard writing into the standard form, such as “دستتون” /dastetun/ ‘your hand’ which becomes “دستتان” /dastetān/;
7) Correctly inserting the mediating consonants, such as “توش” /tuš/ ‘in it’ which becomes “تویش” /tuyaš/;
8) Correcting the spelling or unintentional errors in writing words, such as “ینی” /yani/ ‘meaning’, which becomes “یعنی”;
9) Substituting the standard spelling for fancy writing, such as “عسیس” /?asis/ ‘dear’, which becomes “عزیز” /?aziz/.
The negative changes in the model output that are not consistent with the gold standard labels are as follow:
1) Changing the word order in a sentence, such as the sentence “بیا دوست باشیم دیگر” /biyā dust bāšim digar/ ‘Let's be friends’ which can be changed to “بیا دیگر دوست باشیم”;
2) Replacing a word, such as changing “مگر” /mage/ ‘unless’ to “آیا
/?āyā/ [the question word for yes/no questions];

3) Deleting clitics and replacing them with the corresponding simple word, such as “خودشانند” /xodešānand/ ‘they are themselves’, which can be changed to “خودشان هستند” /xodešān hastand/;
4) Not converting colloquial or non-standard texts into the standard one, such as “یه” /ye/ ‘one’, whose correct form is “یک” /yek/ but without a conversion, no changes have been done;
5) Incorrect insertion of Ezafe, such as “مجسمه‌های” /moǰassamehāye/ ‘statues’ which is incorrectly changed to “مجسمۀهای”;
6) Incorrect insertion of pseudo-space, such as “برای‌تان” /barāyetān/ ‘for you’ while the correct writing is “برایتان”;
7) Deleting a word, such as “ماروپله‌اش رو بازی کنه” /māropelle?as ro bāzi kone/ ‘he/she plays his snake and ladder game’ which can be converted “ماروپله‌اش بازی کنه”. In this example, “رو” /ro/ which is the correct form of “را” /rā/, is deleted from the sentence;
8) Inserting a new word, such as “می‌خواستم بروم داخل” /mixāstam beravam dāxel/ ‘I wanted to enter’ which is converted into “می‌خواستم بروم به داخل” /mixāstam beravam be dāxel/ ‘I wanted to enter to’. In the conversion the preposition “به” /be/ ‘to’ is added;
9) Inserting a space incorrectly, such as “اون که” /?un ke/ ‘one that’ which is changed to “آنکه” /?ānke/ while the correct form is ‘آن که’;
10) No insertion of a proper space, such as “توهم” /toham/ ‘you too’, while the correct form is “تو هم” /to ham/;
11) Replacing the wrong word, such as “تو” /tu/ ‘in’ in the example “تو خانه” ‘at home’ which is converted into “در خانه” /dar xāne/
In general, any change that causes the text to deviate from the expected standard and correct form falls into this category of negative errors.
4. Results
According to the obtained results in Table 1, GPT and Perplexity performed the most and the least positive and negative impacts at the word level, respectively. The most positive impact of the language models had been made by Claude, and Perplexity performed the worst. DeepSeek, GPT, and Gemini stay in the second to fourth ranks, respectively. Among the positive changes, inserting a pseudo-space had the most impact in converting non-standard text to standard. Inserting the Ezafe clitic applied the least change to non-standard data.
Furthermore, according to the negative impacts of the LLM tools on the data, Perplexity and GPT applied the least and the most negative impact, respectively. Among the negative changes, the conversion of colloquial or broken writing to standard writing had the most negative changes, such that Perplexity had the least and Claude had the most changes. Inserting a wrong hamza had the least negative effect on data conversion. In this category, DeepSeek and Gemini did not have any negative effect on data conversion. Word substitution was the most negative change applied by GPT.
Table 1
distribution of positive and negative changes of LLMs in sample data at the word level







GPT


Gemini


Perplexity


Claude


DeepSeek




Positive


12.36


11.49


6.48


15.40


14.44




Negative


34.13


17.50


14.12


23.13


22.48





Next, we examined the performance of the LLM tools against the gold standard data in terms of exact matching at the sentence level for the input and output data. The results are reported in Table1. Based on the results, Claude obtained the highest matching rate (%51.39) and Perplexity the lowest matching rate (%90.7) with the gold standard data at the sentence level. The other language models, namely DeepSeek, Gemini, and GPT, ranked second to fourth respectively among the models. In examining LLMs with each other, without considering the gold standard data, it was concluded that Claude and DeepSeek perform similarly, and GPT and Perplexity had the lowest performance.
Table 2
comparing performance of LLMs with gold standard data at the sentence level







Gold


GPT


Gemini


Perplexity


Claude


DeepSeek




GPT


98



















Gemini


244


84
















Perplexity


81


47


183













Claude


405


136


293


97










DeepSeek


337


118


231


108


368







In the further studying the performances of LLMs, the models were compared from the perspective of syntax. First, the sentences obtained from LLMs and the gold standard data were part-of-speech tagged. Then, the content and functional words were separated from each other and the performances of the models were compared. The gold standard data contained %50.76 of content words and %50.23 of functional words. The results of the analysis are reported in Table 3. Based on the results obtained, Claude and Perplexity were able to obtain the largest and smallest volume of content and functional words, respectively. The ratio of content word changes in GPT (70.29%) was higher than other models, while Claude had decreased this ratio to %67.73. The ratio of functional word change in comparing GPT (%29.71) and Claude (%32.27) was the opposite such that Claude had a better performance in converting non-standard functional words into standard ones.
Table 3
distribution of content and functional words in LLMs compared to gold standard data







Content word


Functional word




GPT


4.85


1.44




Gemini


14.33


4.42




Perplexity


4.00


1.32




Claude


25.23


8.14




DeepSeek


21.15


6.84




5. Conclusion
The general conclusion that can be drawn from the present study is that the difference in the results obtained from different LLMs indicates that due to the difference in their algorithms and training data of the models, the difference in their performance is obvious. Therefore, it is not possible to choose a single LLM and limit a research to that model. In a comprehensive study, it is necessary to examine various LLMs. Examining the differences can be a guide to choose the right tool in future research.
Bibliography
AbdiKhojasteh, H., Ansari, E., & Bohlouli, M. (2020). LSCP: Enhanced large scale colloquial Persian language understanding. Proceedings of the 12th International Conference on Language Resources and Evaluation, Marseille, France: European Language Resources Association, pp. 6323–6327.
Academy of Persian Language and Literature (2023). The Approved Orthography Rules of the Academy of Persian Language and Literature. Tehran: Persian Language and Literature Academy. [In Persian]
Adibian, M., & Momtazi, S. (2012). Converting Persian colloquial text to formal text using neural networks based on converters. Journal of Language and Linguistics, 18(35), 45–67. [In Persian]
Armin, N., & Shamsfard, M. (2010). Converting Persian colloquial text to formal text using n-grams. Proceedings of the 16th Annual National Conference of the Iranian Computer Association. [In Persian]
Ariffin, S. N. A. N., & Tiun, S. (2020). Rule-based text normalization for Malay social media texts. International Journal of Advanced Computer Science and Applications, 11.
Arora, M., & Kansal, V. (2020). Character level embedding with deep convolutional neural network for text normalization of unstructured data for Twitter sentiment analysis. International Journal of Advanced Computer Science and Applications, 11.
Ghayoomi, M., & Mesgarkhoyi, M. (2023). Analysis of the corpus of Persian language data in cyberspace. Language and Linguistics, 19, 37: 117–136. [In Persian]
Kozhirbayev, Z., & Yessenbayev, Z. (2020). Kazakh text normalization using machine translation approaches. CEUR Workshop Proceedings, 2780.
Liu, Y., Ji, B., Yu, J., Tan, Y., Ma, J., & Wu, Q. (2021). An advanced ICD-9 terminology standardization method based on BERT and text similarity. Advances in Natural Computation, Fuzzy Systems and Knowledge Discovery. ICNC-FSKD 2020. Lecture Notes on Data Engineering and Communications Technologies, eds. Meng, H., Lei, T., Li, M., Li, K., Xiong, N., & Wang, L., Speringer, vol. 88, pp. 1868–1879.
Mansfield, C., Sun, M., Liu, Y., Gandhe, A., & Hoffmeister, B. (2019). Neural text normalization with subword units. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, eds. Loukina, A., Morales, M., & Kumar, R., Minneapolis, Minnesota: Association for Computational Linguistics, vol. 2, pp. 190–196.
Masoumi, V., Salehi, M., Veisi, H., Haddadian, G., Ranjbar, V., & Sahebdel, M. (2020). TeleCrowd: A crowdsourcing approach to create informal to formal text corpora. Arxiv.
https://arxiv.org/abs/2004.11771.
Rasooli, M. S., Bakhtyari, F., Shafiei, F., Ravanbakhsh, M., & Callison-Burch, C. (2020). Automatic standardization of colloquial Persian. Arxiv. https://arxiv.org/abs/2012.05879
Tajalli, V., Kalantari, F., & Shamsfard, M. (2023). Developing an informal formal Persian corpus. Arxiv. https://arxiv.org/abs/2308.05336.

کلیدواژه‌ها English

Colloquial Persian
non-standard (broken) writing
standardization
large language models
Persian orthography
آرمین، نادیه و شمس‌فرد، مهرنوش. (1389)، تبدیل متن محاوره‌ای فارسی به رسمی به کمک n-gramها. مجموعه مقالات شانزدهمین کنفرانس ملی سالانه انجمن کامپیوتر ایران.
https://www.researchgate.net/publication/314239155_tbdyl_mtn_mhawrh_ay_farsy_bh_kmk_N_gram_ha
ادیبیان، مجید، و ممتازی، سعیده. (1401). تبدیل متن محاوره به رسمی فارسی با استفاده از شبکه‌های عصبی مبتنی بر مبدل. زبان و زبان‌شناسی. 18 (35)، 67-45.
بی‌جن‌خان، محمود. (1383) نقش پیکره زبانی در نوشتن دستور زبان: معرفی یک نرم‌افزار رایانه‌ای. زبان‌شناسی. 19(37): 48-67.
فرهنگستان زبان و ادب فارسی. (1389). دستور خط مصوب فرهنگستان. تهران: فرهنگستان زبان و ادب فارسی.
فرهنگستان زبان و ادب فارسی. (1402). دستور خط مصوب فرهنگستان. تهران: فرهنگستان زبان و ادب فارسی.
صادقی، علی‌اشرف و زندی‌مقدم، زهرا. (1394). فرهنگ املایی خطّ فارسی مصوب فرهنگستان زبان و ادب فارسی. تهران: فرهنگستان زبان و ادب فارسی (نشر آثار).
طبیب‌زاده، امید. (1398). مبانی و دستور خط فارسی شکسته براساس صد سال آثار داستانی و نمایشی. تهران: پژوهشگاه علوم انسانی و مطالعات فرهنگی.
قیومی، مسعود و مسگرخویی، مریم. (۱۴۰۲). تحلیل بر پیکره حاصل از داده‌های زبان فارسی در فضای مجازی. زبان و زبان‌شناسی. ۱۹(۳۷)، ۱۱۷۱۳۶.