Corpus-based typology and word order in Persian

Document Type : .

10.30465/lsi.2026.54514.1844
Abstract
Abstract
This study examines word order typology in Persian through the lens of corpus-based typology, challenging traditional grammar-based classifications. While grammar-based studies have characterized Persian as a ‘mixed’ language—exhibiting both verb-final and verb-initial features—they fail to capture the gradient nature of word order variation. Reviewing recent quantitative corpus studies across multiple syntactic parameters, this research demonstrates that Persian shows considerable disorder in head-dependent ordering but remarkable stability in subject-object ordering (over 90% subject-before-object). The presence of differential object marking () does not enable flexible subject-object order, contrary to expectations. Furthermore, Persian places nominal objects before verbs in approximately 82% of cases, pronominal objects before verbs in 92% of cases, and adjectives after nouns in 95% of cases. The study advances areal explanations for these mixed properties, situating Persian within the Western Asia Transition Zone, where centuries of contact between Turkic (verb-final) and Semitic (verb-initial) languages have shaped Iranian languages. This regional perspective explains why Persian exhibits apparently contradictory typological features without necessarily being in transition toward a different type. The findings highlight the superiority of corpus-based approaches for revealing systematic patterns in word order variation and emphasize the need for expanded Persian corpora.
Keywords: Areal Typology, Corpus, Persian, Verb-Final, Verb-Initial, Word Order
1. Introduction
This article presents a comprehensive investigation into the typology of word order in the Persian language, with a specific focus on the emerging paradigm of corpus-based typology. Traditional grammar-based typological studies, following the foundational work of Greenberg (1963) and Dryer (1992), have typically categorized languages into discrete typological classes based on the relative order of core syntactic elements, particularly the positioning of the verb with respect to subject and object. Persian has posed a persistent challenge to such categorical classifications, as it exhibits mixed characteristics that do not align neatly with either verb-initial or verb-final language types.
Research Question(s)
Does Persian conform to a consistent word order typology, and if not, what factors explain its mixed nature? Through a systematic review of corpus-based typological studies, this article demonstrates that Persian's word order properties are more complex and gradient than previously recognized, showing high disorder in some syntactic domains and remarkable stability in others.
2. Literature Review
Greenberg's (1963) seminal work established implicational universals that link word order patterns across languages. Dryer's (1992) subsequent elaboration identified strong correlations between verb-object positioning and other syntactic properties, allowing languages to be classified as either verb-initial/head-initial or verb-final/head-final. Persian was found to present an inconsistent profile when evaluated against these correlation pairs.
Building on Dabir Moghaddam's (2013, 2018) comprehensive application of Dryer's parameters to Persian, the article shows that Persian demonstrates: Verb-final features in only four parameters, including the placement of adpositional phrases, manner adverbs, the copula, and wh-words. Verb-initial features in nine parameters, including relative clauses, genitive constructions, complementizers, and interrogative particles. Intermediate or mixed behavior in four additional parameters (comparative constructions, auxiliary placement, tense/mood markers, and object positioning). This distribution reveals that Persian cannot be confidently assigned to either typological category. Dabir-Moghaddam (2013) proposed two competing hypotheses: Persian is either undergoing a slow typological shift from verb-final to verb-initial, or it has inherently free word order. Examination of historical data from Old and Middle Persian supports the shift hypothesis, showing a gradual but consistent movement toward verb-initial patterns across historical periods.
3. Methodology
The last two decades have witnessed a paradigm shift in typological research, driven by the proliferation of linguistic corpora. This development has given rise to what is now termed corpus-based typology, token-based typology, or gradient typology, in contrast to traditional feature-based or grammar-based approaches (Haig et al., 2022; Levshina, 2019, 2021; Levshina et al., 2023). Key innovations of this approach include: direct access to primary data, moving beyond second-hand judgments; context-sensitive analysis grounded in actual language use; quantitative measurements that capture the gradient nature of linguistic features; Discovery of new cross-linguistic generalizations and re-evaluation of established universals; enhanced explanatory power for cognitive, functional, diachronic, and areal accounts.
Several corpus types have proven valuable for typological investigations, including parallel corpora (e.g., Universal Declaration of Human Rights translations, Europarl, film subtitle corpora), comparable corpora (e.g., Multi-CAST), and corpora with unified annotation (e.g., Universal Dependencies). These resources enable systematic cross-linguistic comparison and quantitative analysis of word order patterns.
4. Results
Head-Dependent Ordering: Futrell et al. (2015) examined dependency length minimization across 34 languages, measuring the disorder (entropy) in head-dependent ordering. Persian occupies a position in the lower half of their entropy scale, indicating considerable inconsistency in whether it is consistently head-initial or head-final across different construction types. While the language places verbs sentence-finally, it cannot be characterized as consistently head-final. Notably, both Turkish (head-final) and Arabic (head-initial) show approximately half the disorder of Persian, demonstrating greater typological consistency.
Subject-Object Ordering: Remarkably, despite its mixed head-dependent ordering, Persian exhibits high stability in subject-object ordering. Futrell et al. (2015) and Levshina et al. (2023) demonstrate that Persian places subjects before objects in over 90% of cases, placing it among languages with rigid subject-object order. This finding is particularly significant given Persian's differential object marking (DOM) system employing the marker *rā* (Rasekh-Mahand & Parizadeh, 2024). Contrary to expectations that case marking enables freer word order, Persian's DOM does not result in flexible subject-object ordering, aligning with the cross-linguistic pattern that languages with robust case systems may nonetheless maintain fixed word order.
Interaction between Case Marking and Word Order: Levshina (2019) and Levshina et al. (2023) establish a nuanced relationship between case marking and word order flexibility. Languages with comprehensive case systems (e.g., Lithuanian, Hungarian, Latin) tend to show greater word order flexibility, while languages with reduced or differential marking exhibit more rigid order. Persian's DOM system places it in an intermediate position: it has some case marking but maintains relatively fixed word order. This pattern confirms functional-adaptive explanations for word order properties: languages tend to maintain sufficient disambiguation strategies, either through case marking or through fixed word order, but these strategies are not necessarily mutually enabling (Levshina, 2019, p. 547).
Agreement and Word Order: Spirgath (2025) examines the interaction between verb agreement and word order entropy, showing a negative correlation: languages with stronger agreement tend to show more rigid word order. Persian, which has robust agreement marking, confirms this pattern, showing relatively low word order disorder (approximately 20% flexibility) despite its rich agreement system. This finding distinguishes agreement from case marking, as the two systems serve complementary functions in information management.
Multiple Word Order Parameters: Gerdes et al. (2021) provide detailed quantitative analysis of several word order parameters: Object-Verb Order: Persian places both nominal and pronominal objects before the verb in over 90% of cases, showing remarkable consistency comparable to Turkish and contrasting sharply with Arabic and English. Adjective-Noun Order: Persian overwhelmingly (approximately 95%) places adjectives after nouns, confirming Greenberg's Universal 19 and showing greater consistency than expected in a mixed typology language. Adverb Placement: Persian places lexical adverbs before the verb in over 96% of cases, but adverbial clauses appear postverbally in approximately 17% of cases, revealing variable behavior depending on adverbial type. Adposition Type: Persian is consistently prepositional, aligning with Arabic and contrasting with Turkish, demonstrating one of the most stable parameters in the language.
Post-Predicate Elements: Focusing on post-predicate constituents, Rasekh-Mahand et al. (2024) and Haig et al. (2024a) demonstrate that goal and caused goal arguments show the highest frequency of postverbal placement, with Persian placing 84% of goal arguments after the verb. This behavior is not random but systematically conditioned by semantic role, with animacy, style, and constituent weight playing secondary roles. Most importantly, this phenomenon cannot be attributed to information structure but reflects a structural property of colloquial Persian.
5. Discussion
The article advances areal-typological explanations for Persian's mixed word order properties, drawing on the concept of the Western Asia Transition Zone (WATZ) proposed by Haig (2017) and Haig and Khan (2018). This region, encompassing parts of Iran, eastern Turkey, and northern Iraq, has historically been a contact zone between Turkic (verb-final, postpositional), Semitic (verb-initial, prepositional), Indo-European (primarily Iranian), and Kartvelian language families.
Stilo (2005, 2009) pioneered the analysis of how Persian and other Iranian languages have been shaped by their position between these typologically distinct families. His research demonstrates that Iranian languages in the north and east (adjacent to Turkic and Caucasian languages) show more verb-final characteristics, while those in the west and south (adjacent to Semitic languages) show more verb-initial features. Haig et al. (2025) provide corpus-based confirmation of these patterns across 35 language varieties in the WATZ. Key findings include:
Object Position: While Iranian and Turkic languages historically maintain object-before-verb order in over 90% of instances, Semitic languages in the region show variable behavior. Some Neo-Aramaic varieties, surrounded by Iranian languages, have adopted object-before-verb order, demonstrating contact-induced change.
Postverbal Goal Arguments: The widespread tendency to place goal arguments after the verb extends across language families in the WATZ, with Iranian and Turkic languages showing rates of 60-100% postverbal goal placement. This pattern is strongest in western Iranian languages (closer to Semitic contact) and weaker in eastern varieties (further from Semitic contact).
Adposition Type: Iranian languages show a gradient from predominantly postpositional in the north and east (closer to Turkic) to predominantly prepositional in the west and south (closer to Semitic), with many languages exhibiting both.
These areal patterns challenge purely historical explanations for Persian's mixed typology. Rather than representing an inherent property or a unidirectional diachronic shift, Persian's mixed characteristics reflect centuries of contact with typologically diverse language families. This perspective explains why Persian can be simultaneously verb-final (in object-verb order) and prepositional (patterned with verb-initial languages) without indicating transitional status or free word order.
6. Conclusion
This review of corpus-based typological research reveals several crucial insights about Persian word order. Persian's typological properties are not binary but exist on a continuum, showing high disorder in some domains (head-dependent ordering, postverbal argument placement) and high stability in others (subject-object ordering, adjective-noun order, adposition type). The interaction between case marking, agreement, and word order supports functional explanations: languages maintain sufficient disambiguation mechanisms without unnecessary redundancy. Persian's mixed properties cannot be attributed solely to internal historical development but reflect prolonged contact with neighboring language families in the Western Asia Transition Zone. Corpus-based approaches provide more accurate and nuanced descriptions than grammar-based typology, revealing systematic patterns in apparent disorder.
Bibliography:
Csató, É. Á., Isaksson, B., & Jahani, C. (Eds.) (2005). Linguistic convergence and areal diffusion: Case studies from Iranian, Semitic and Turkic. Routledge.
Dabir Moghaddam, M. (2018). Typological approaches and dialects. In A. Sedighi & P. Shabani-Jadidi (Eds.), The Oxford handbook of Persian linguistics (pp. 52–88). Oxford University Press.
Dryer, M. S. (1992). The Greenbergian word order correlations. Language, 68, 81-138.
Dryer, M. S. (1997). On the six-way word order typology. Studies in Language, 21(2), pp. 69–103.
Dryer, M. S. (2013). Order of subject, object and verb. In M. S. Dryer & M. Haspelmath (Eds.), WALS Online (v2020.4) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13950591. Retrieved from http://wals.info/chapter/81.
Dryer, M. S., & Haspelmath, M. (Eds.) (2013). WALS Online (v2020.4) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13950591. Retrieved from https://wals.info.
Frommer, P. (1981). Post-verbal phenomena in colloquial Persian syntax (Doctoral dissertation), University of Southern California.
Futrell, R., Mahowald, K., Edward Gibson, E. (2015). Large-scale evidence of dependency length minimization in 37 languages. PNAS, 112(33), pp. 10336–10341.
Gerdes, K., Kahane S., and Chen, X. (2021). Typometrics: From Implicational to Quantitative Universals in Word Order Typology. Glossa: a journal of general linguistics 6(1): 17. 1–31. DOI: https://doi. org/10.5334/gjgl.764
Goldhahn, D., Eckart, Th., Quasthoff, U. (2012). Building Large Monolingual Dictionaries at the Leipzig Corpora Collection: From 100 to 200 Languages. In: Calzolari, N., Choukri, Kh., Declerck, Th. et al. (Eds.), Proceedings of the Eighth International Conference on Language Resources and Evaluation, pp. 759–765. Istanbul: ELRA.
Greenberg, J. (1963/1966). Some Universals of Grammar with Particular Reference to the Order of Meaningful Elements. Universals of Language, London: MIT Press, pp. 110-113.
Haig, G. (2017). Western Asia: East Anatolia as a transition zone. In Hickey, Raymond (ed.), The Cambridge Handbook of Areal Linguistics. Cambridge: CUP. pp. 396-423.
Haig, G. & Khan, G. (2018). Introduction. In: Haig, Geoffrey & Geoffrey Khan (eds.) The languages and linguistics of Western Asia. An areal perspective. Berlin: Mouton de Gruyter. Pp. 1-29.
Haig, G. & Rasekh-Mahand, M. (2019). Post-predicate constituents in Iranian and neighboring languages. Position paper for the project Post-predicate constituents in Iranian and neighboring languages. https://www.uni-bamberg.de/fileadmin/aspra/05_Events/2019_Post
Haig, G., & Rasekh-Mahand, M. (Eds.). (2022). HamBam: The Hamedan-Bamberg corpus of contemporary spoken Persian. University of Bamberg. http://multicast.aspra.uni-bamberg.de/resources/hambam/
Haig, G., & Schnell, S. (2014). Annotations using GRAID (Grammatical Relations and Animacy in Discourse): Introduction and guidelines for annotators (Version 7.0). https://multicast.aspra.uni-bamberg.de/
Haig, G., & Schnell, S. (2015). Multi-CAST: Multilingual corpus of annotated spoken texts. University of Bamberg. https://multicast.aspra.uni-bamberg.de/
Haig, G., & Schnell, S. (Eds.). (2023). Multi-CAST: Multilingual corpus of annotated spoken texts (Version 2311). University of Bamberg. http://multicast.aspra.uni-bamberg.de/
Haig, G., Schnell, S., & Seifart, F. (Eds.). (2022). Doing corpus-based typology with spoken language corpora: State of the art (Language Documentation & Conservation Special Publication No. 25). University of Hawai‘i Press. https://doi.org/10.5167/uzh-217096
Haig, G., & Rasekh-Mahand, M. (Eds.). (2022). HamBam: The Hamedan-Bamberg corpus of contemporary spoken Persian. University of Bamberg. http://multicast.aspra.uni-bamberg.de/resources/hambam/
Haig, G., & Schnell, S. (2014). Annotations using GRAID (Grammatical Relations and Animacy in Discourse): Introduction and guidelines for annotators (Version 7.0). https://multicast.aspra.uni-bamberg.de/
Haig, G., & Schnell, S. (2015). Multi-CAST: Multilingual corpus of annotated spoken texts. University of Bamberg. https://multicast.aspra.uni-bamberg.de/
Haig, G., & Schnell, S. (Eds.). (2023). Multi-CAST: Multilingual corpus of annotated spoken texts (Version 2311). University of Bamberg. http://multicast.aspra.uni-bamberg.de/
Haig, G., Schnell, S., & Schiborr, N. N. (2021). Universals of reference in discourse and grammar: Evidence from the Multi-CAST collection of spoken corpora. In G. Haig, S. Schnell, & F. Seifart (Eds.), Doing corpus-based typology with spoken language corpora (Language Documentation & Conservation Special Publication No. 25). University of Hawai‘i Press.
Haig, G., Vollmer, M., & Thiele, H. (2019). Multi-CAST Northern Kurdish. In G. Haig & S. Schnell (Eds.), Multi-CAST: Multilingual corpus of annotated spoken texts (Version 2311). University of Bamberg. http://multicast.aspra.uni-bamberg.de/#nkurd
Haig, G., Stilo, D., Schiborr, N. N., & Dogan, M. (Eds.). (2022). WOWA: Word Order in Western Asia: A spoken-language-based corpus for investigating areal effects in word order variation. University of Bamberg. http://multicast.aspra.uni-bamberg.de/resources/wowa/
Haig, G., Rasekh-Mahand, M., Stilo, D., Schreiber, L., & Schiborr, N. N. (Eds.). (2024a). Post-predicate elements in the Western Asian Transition Zone: A corpus-based approach to areal typology (Contact and Multilingualism 8). Language Science Press. https://langsci-press.org/catalog/book/395
Haig, G., Rasekh-Mahand, M., Stilo, D., Schreiber, L., & Schiborr, N. (2024b). Post-predicate elements in the Western Asian Transition Zone: Data, theory, and methods. In G. Haig, M. Rasekh-Mahand, D. Stilo, L. Schreiber, & N. Schiborr (Eds.), Post-predicate elements in the Western Asian Transition Zone: A corpus-based approach to areal typology (pp. 1–48). Language Science Press. https://langsci-press.org/catalog/book/395
Haig, G., Noorlander, P., & Schiborr, N. (2025). Which word order features are stable in a contact setting? Corpus-based evidence from the Western Asian Transition Zone. In J. Darquennes, J. C. Salmons, & W. Vandenbussche (Eds.), Language contact, an international handbook (Vol. 2, pp. 159–184). De Gruyter Mouton. https://doi.org/10.1515/9783110443011-010
Levshina, N. (2015). European analytic causatives as a comparative concept: Evidence from a parallel corpus of film subtitles. Folia Linguistica, 49(2), 487–520.
Levshina, N. (2019). Token-based typology and word order entropy. Linguistic Typology, 23(3), 533–572. https://doi.org/10.1515/lingty-2019-0025
Levshina, N. (2022a). Corpus-based typology: Applications, challenges and some solutions. Linguistic Typology, 26(1), 129–160.
Levshina, N. (2022b). Communicative efficiency: Language structure and use. Cambridge University Press.
Levshina, N., Nambodiripad, S., Allassonnière-Tang, M., Kramer, M., Talamo, L., Verkerk, A., Wilmoth, S., Garrido Rodriguez, G., Gupton, T. M., Kidd, E., Liu, Z., Naccarato, C., Nordlinger, R., Panova, A., & Stoynova, N. (2023). Why we need a gradient approach to word order. Linguistics, 61(4), 825–883. https://doi.org/10.1515/ling-2021-0098
Parizadeh, M., & Rasekh-Mahand, M. (2024). Post-predicate elements in Early New Persian. In G. Haig, M. Rasekh-Mahand, D. Stilo, L. Schreiber, & N. Schiborr (Eds.), Post-predicate elements in the Western Asian Transition Zone: A corpus-based approach to areal typology (pp. 184–196). Language Science Press. https://langsci-press.org/catalog/book/395
Rasekh-Mahand, M., & Parizadeh, M. (2024). Different functions of ‘rā’ in New Persian: A semantic map analysis. Journal of Historical Linguistics, 14(1), 31–57. https://doi.org/10.1075/jhl.21056.ras
Rasekh-Mahand, M., Izadi, E., Parizadeh, M., Haig, G., & Schiborr, N. (2024). Post-predicate elements in colloquial Persian: A multi-variate analysis. In G. Haig, M. Rasekh-Mahand, D. Stilo, L. Schreiber, & N. Schiborr (Eds.), Post-predicate elements in the Western Asian Transition Zone: A corpus-based approach to areal typology (pp. 151–182). Language Science Press. https://langsci-press.org/catalog/book/395
Schnell, S., & Schiborr, N. N. (2022). Crosslinguistic corpus studies in linguistic typology. Annual Review of Linguistics, 8, 171–191. https://doi.org/10.1146/annurev-linguistics-031120-104629
Spirgath, A. (2025). The interaction of word order entropy and verb agreement: A token-based approach. Linguistics in Amsterdam, 16(1), 30–50.
Stilo, D. (1994). Phonological systems in contact in Iran and Transcaucasia. In M. Marashi (Ed.), Persian studies in North America: Studies in honor of Mohammad Ali Jazayery. Iranbooks.
Stilo, D. (2005). Iranian as buffer zone between the universal typologies of Turkic and Semitic. In É. Á. Csató, B. Isaksson, & C. Jahani (Eds.), Linguistic convergence and areal diffusion: Case studies from Iranian, Semitic and Turkic (pp. 35–63). Routledge.
Stilo, D. (2009a). Circumpositions as an areal response: The case study of the Iranian zone. Turkic Languages, 13, 3–33.
Stilo, D. (2009b). Case in Iranian: From reduction and loss to innovation and renewal. In A. L. Malchukov & A. Spencer (Eds.), The Oxford handbook of case. Oxford University Press. https://doi.org/10.1093/oxfordhb/9780199206476.013.0049
Stilo, D. (2012). Intersection zones, overlapping isoglosses, and ‘Fade-out/Fade-in’ phenomena in Central Iran. In B. Aghaei & M. R. Ghanoonparvar (Eds.), Iranian languages and culture: Essays in honor of Gernot Ludwig Windfuhr (pp. 3–33). Mazda Publishers.
Tiedemann, J. (2012). Parallel data, tools and interfaces in OPUS. In N. Calzolari, K. Choukri, T. Declerck, M. U. Doğan, B. Maegaard, J. Mariani, A. Moreno, J. Odijk, & S. Piperidis (Eds.), Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC-2012) (pp. 2214–2218). European Language Resources Association (ELRA).
Vatanen, T., Väyrynen, J. J., & Virpioja, S. (2010). Language identification of short text segments with n-gram models. In Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC’10) (pp. 3423–3430). European Language Resources Association (ELRA).
Verkerk, A. (2014). The evolutionary dynamics of motion event encoding [Doctoral dissertation, Radboud University Nijmegen].
Zeman, D., et al. (2024). Universal Dependencies 2.15 [Data set]. LINDAT/CLARIAH-CZ digital library, Institute of Formal and Applied Linguistics (ÚFAL), Faculty of Mathematics and Physics, Charles University. http://hdl.handle.net/11234/1-5787
Keywords
Subjects

Anonby, C. V. D. W. (2015). A Grammar of Kumzari: A mixed Perso-Arabian language of Oman. Leiden: Leiden University PhD Dissertation.
Croft, W. (2003). Typology and universals, 2nd ed. Cambridge: Cambridge University Press.
Csató, É. Á, Isaksson, B. & Jahani, C. (eds.) (2005). Linguistic Convergence and Areal Diffusion: Case studies from Iranian, Semitic and Turkic. London, New York: Routledge.
Dabir Moghaddam, M. (2018). Typological approaches and dialects. In Anousha Sedighi & Pouneh Shabani-Jadidi (eds.), The Oxford handbook of Persian linguistics, 5288. Oxford: Oxford University Press.
Dryer, M. S. (1992). The Greenbergian Word Order Correlations. Language 68: 81-138.
Dryer, M. S. (1997). On the six-way word order typology. Studies in Language, 21(2), pp. 69–103.
Dryer, M. S. (2013). Order of Subject, Object and Verb. In Dryer, M. S. & Haspelmath, M. (Eds.), WALS Online (v2020.4) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13950591. Retrieved from http://wals.info/chapter/81.
Dryer, M. S., & Haspelmath, M. (Eds.) (2013). WALS Online (v2020.4) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13950591. Retrieved from https://wals.info.
Ebert, C., Bickel, B., & Widmer, P. (2024). "Areal and phylogenetic dimensions of word order variation in Indo-European languages" Linguistics, vol. 62, no. 5, pp. 1085-1116. https://doi.org/10.1515/ling-2022-0146
Frommer, P. (1981). Post-verbal phenomena in colloquial Persian syntax. Los Angeles: University of Southern California. (Doctoral dissertation).
Futrell, R., Mahowald, K., Edward Gibson, E. (2015). Large-scale evidence of dependency length minimization in 37 languages. PNAS, 112(33), pp. 10336–10341. 22
Gerdes, K., Sylvain K., & Xinying, C. (2021). Typometrics: From Implicational to Quantitative Universals in Word Order Typology. Glossa: a journal of general linguistics 6(1): 17. 1–31. DOI: https://doi. org/10.5334/gjgl.764 
Goldhahn, D., Eckart, Th., & Quasthoff, U. (2012). Building Large Monolingual Dictionaries at the Leipzig Corpora Collection: From 100 to 200 Languages. In: Calzolari, N., Choukri, Kh., Declerck, Th. et al. (Eds.), Proceedings of the Eighth International Conference on Language Resources and Evaluation, pp. 759–765. Istanbul: ELRA.
Greenberg, J. (1963/1966). Some Universals of Grammar with Particular Reference to the Order of Meaningful Elements. Universals of Language, London: MIT Press, pp. 110-113.
Haig, G. (2017). Western Asia: East Anatolia as a transition zone. In Hickey, Raymond (ed.), The Cambridge Handbook of Areal Linguistics, 396-423. Cambridge: CUP.
Haig, G., & Khan, G. (2018). Introduction. In: Haig, Geoffrey & Geoffrey Khan (eds.) The languages and linguistics of Western Asia. An areal perspective, 1-29. Berlin: Mouton de Gruyter.
Haig, G., & Rasekh-Mahand, M. (2019). Post-predicate constituents in Iranian and neighboring languages. Position paper for the project Post-predicate constituents in Iranian and neighboring languages. https://www.uni-bamberg.de/fileadmin/aspra/05_Events/2019_Post
Haig, G., & Rasekh-Mahand, M. (eds.). (2022). HamBam: The Hamedan-Bamberg Corpus of Contemporary Spoken Persian. Bamberg: University of Bamberg. (multicast.aspra.uni-bamberg.de/resources/hambam/).
Haig, G., & Schnell, S. (2014). Annotations using GRAID (Grammatical Relations and Animacy in Discourse): Introduction and guidelines for annotators. Version 7.0. https://multicast.aspra.uni-bamberg.de/
Haig, G., & Schnell, S. (2015). Multi-CAST: Multilingual Corpus of Annotated Spoken Texts. Bamberg: University of Bamberg. https://multicast.aspra.uni-bamberg.de/
Haig, G., & Schnell, S. (eds.). (2023). Multi-CAST: Multilingual Corpus of Annotated Spoken Texts (Version 2311). Bamberg: University of Bamberg. multicast.aspra.uni-bamberg.de/
Haig, G., Schnell, S., & Seifart, F. (eds.). (2022). Doing Corpus-Based Typology With Spoken Language Corpora: State of the Art. Language Documentation & Conservation Special Publication 25. Honolulu: University of Hawai‘i Press. https://doi.org/10.5167/uzh-217096
Haig, G., Schnell, S. & Schiborr, N. N. (2021). Universals of reference in discourse and grammar: Evidence from the Multi-CAST collection of spoken corpora. In G. Haig, S. Schnell & F. Seifart (eds.), Doing Corpus-Based Typology with Spoken Language Corpora. Language Documentation & Conservation Special Publication 25. Honolulu: University of Hawai’i Press.
Haig, G., Vollmer, M., & Thiele, H. (2019). Multi-CAST Northern Kurdish. In G. Haig & S. Schnell (eds.), Multi-CAST: Multilingual Corpus of Annotated Spoken Texts (Version 2311). Bamberg: University of Bamberg. multicast.aspra.uni-bamberg.de/#nkurd
Haig, G., Stilo, D. Schiborr, N. N., & Dogan, M. (eds.). (2022). WOWA: Word Order in Western Asia: A spoken-language-based corpus for investigating areal effects in word order variation. Bamberg: University of Bamberg. (multicast.aspra.uni-bamberg.de/resources/wowa/).
Haig, G., Rasekh-Mahand, M., Stilo, D., Schreiber, L. & Schiborr, N. N. (eds.). (2024a). Post-predicate elements in the Western Asian Transition Zone: A corpus-based approach to areal typology (Contact and Multilingualism 8). Berlin: Language Science Press. https://langsci-press.org/catalog/book/395 
Haig, G., Rasekh-Mahand, M., Stilo, D., Schreiber, L. & Schiborr, N. N. (eds.). (2024b). Post-predicate elements in the Western Asian Transition Zone: Data, theory, and methods. In Geoffrey Haig, Mohammad Rasekh-Mahand, Donald Stilo, Laurentia Schreiber & Nils Schiborr (eds.). 2024. Post-predicate elements in the Western Asian Transition Zone: a corpus-based approach to areal typology. Berlin: Language Science Press. 1-48. https://langsci-press.org/catalog/book/395
Haig, G., Noorlander, P., & Schiborr, N. N. (2025). "10. Which word order features are stable in a contact setting? Corpus-based evidence from the Western Asian Transition Zone". Language contact, an international handbook, Volume 2, edited by Jeroen Darquennes, Joseph C. Salmons and Wim Vandenbussche, Berlin, Boston: De Gruyter Mouton. 159-184. https://doi.org/10.1515/9783110443011-010
Ivani, J., & Bickel, B. (In press). Databases for comparative syntactic research. To appear in The Cambridge Handbook of Comparative Syntax, ed. Barbiers, S., N. Corver & M. Polinsky
Jing, Y., Blasi, D. E., & Bickel, B. (2022). Dependency-length minimization and its limits: A possible role for a probabilistic version of the final-over-final condition. Language, 98(3), pp. 397–418. DOI https://doi.org/10.1353/lan.0.0267.
Kiparsky, P. (1997). The rise of positional licensing. In Ans von Kemenade and Nigel Vincent, editors, Parameters of morphosyntactic change, pages 460– 494. Cambridge University Press.
Levshina, N. (2015). European analytic causatives as a comparative concept. Evidence from a parallel corpus of film subtitles. Folia Linguistica 49(2). 487–520.
Levshina, N. (2019). Token-based typology and word order entropy. Linguistic Typology, 23(3), pp. 533–572. DOI https://doi.org/10.1515/lingty-2019-0025.
Levshina N (2021). Cross-Linguistic Trade-Offs and Causal Relationships Between Cues to Grammatical Subject and Object, and the Problem of Efficiency-Related Explanations. Front. Psychol. 12:648200. doi: 10.3389/fpsyg.2021.648200
Levshina, N. (2022). Corpus-based typology: Applications, challenges and some solutions. Linguistic Typology 26(1). 129–160.
Levshina, N. (2022). Communicative Efficiency: Language Structure and Use. Cambridge: Cambridge University Press.
Levshina, N., Nambodiripad, S., Allassonnière-Tang, M., Kramer, M., Talamo, L., Verkerk, A., Wilmoth, S., Garrido Rodriguez, G., Gupton, T.M., Kidd, E., Liu, Z., Naccarato, Ch., Nordlinger, R., Panova, A., & Stoynova, N. (2023). Why we need a gradient approach to word order. Linguistics, 61(4), pp. 825–883. https://doi.org/10.1515/ling-2021-0098.
Mohammadirad, M. (2020). Predicative possession across Western Iranian languages, Folia Linguistica, 54(3): 497-526. https://doi.org/10.1515/flin-2020-2038
Parizadeh, M., & Rasekh-Mahand, M. (2024). Post-predicate elements in Early New Persian. In Geoffrey Haig, Mohammad Rasekh-Mahand, Donald Stilo, Laurentia Schreiber & Nils Schiborr (eds.). Post-predicate elements in the Western Asian Transition Zone: a corpus-based approach to areal typology. Berlin: Language Science Press. 184-196. https://langsci-press.org/catalog/book/395
Rasekh-Mahand, M., & Parizadeh, M. (2024). Different functions of ‘rā’ in New Persian A semantic map analysis. Journal of Historical Linguistics14(1): 31 – 57.  https://doi.org/10.1075/jhl.21056.ras
Rasekh-Mahand, M., Izadi, E., Parizadeh, M., Haig, G., & Schiborr, N. N. (2024). Post-predicate elements in colloquial Persian; a multi-variate analysis; in Geoffrey Haig, Mohammad Rasekh-Mahand, Donald Stilo, Laurentia Schreiber & Nils Schiborr (eds.). 2024. Post-predicate elements in the Western Asian Transition Zone: a corpus-based approach to areal typology. Berlin: Language Science Press. 151-182. https://langsci-press.org/catalog/book/395
Schnell, S., & Schiborr, N. N. (2022). Crosslinguistic Corpus Studies in Linguistic Typology. Annual Review Linguistics. 8:171-191. https://doi.org/10.1146/annurev-linguistics-031120-104629
Spirgath, A. (2025). The interaction of word order entropy and verb agreement: A token-based approach. Linguistics in Amsterdam 16 (1): 30-50.
Stilo, D. (1994). Phonological Systems in Contact in Iran and Transcaucasia, [in] Mehdi, Marashi, Ed., Persian Studies in North America: Studies in Honor of Mohammad Ali Jazayery, Bethesda, Maryland: Iranbooks,1994.
Stilo, D. (2005). Iranian as buffer zone between the universal typologies of Turkic and Semitic. In Csató, Éva Ágnes, Isaksson, Bo & Jahani, Carina (eds.), Linguistic Convergence and Areal Diffusion: Case studies from Iranian, Semitic and Turkic, 35-63. London, New York: Routledge.
Stilo, D. (2009a). Circumpositions as an areal response: The case study of the Iranian zone. Turkic Languages 13. 3-33.
Stilo, D. (2009b). ' Case In Iranian: From Reduction and Loss to Innovation and Renewal', in Andrej L. Malchukov, and Andrew Spencer (eds), The Oxford Handbook Case.  https://doi.org/10.1093/oxfordhb/9780199206476.013.0049, accessed 22 Dec. 2025.
Stilo, D. (2012). Intersection zones, overlapping isoglosses, and ‘Fade-out/Fade-in’ phenomena in Central Iran. In: Behrad Aghaei & M. R. Ghanoonparvar (eds.), Iranian Languages and Culture: Essays in Honor of Gernot Ludwig Windfuhr, 3-33. Costa Mesa: Mazda Publishers.
Tiedemann, J. (2012). Parallel data, tools and interfaces in OPUS. In Nicoletta Calzolari (Conference Chair), Khalid Choukri, Thierry Declerck, Mehmet Uğur Doğan, Bente Maegaard, Joseph Mariani, Asuncion Moreno, Jan Odijk & Stelios Piperidis (eds.), Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC-2012), 2214–2218. Istanbul: European Language Resources Association (ELRA).
Vatanen, T., Väyrynen, J. J., & Virpioja, S. (2010). Language identification of short text segments with n-gram models. In Proceedings of the seventh international conference on language resources and evaluation (LREC’10), 3423–3430. Malta: European Language Resources Association (ELRA).
Verkerk, A. (2014). The evolutionary dynamics of motion event encoding. PhD dissertation. Radboud University Nijmegen.
Zeman, D., et al. (2024). Universal Dependencies 2.15, LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL), Faculty of Mathematics and Physics, Charles University. Retrieved from http://hdl.handle.net/11234/1-5787.[1]
استناد به این مقاله: راسخ‌مهند، محمد. (1404). رده‌شناسی پیکره‌بنیاد و توالی کلمات در زبان فارسی. زبان و زبان‌شناسی، 21(42)، 1- 60.
 doi: 10.30465/lsi.2026.54514.1844