<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE ArticleSet PUBLIC "-//NLM//DTD PubMed 2.7//EN" "https://dtd.nlm.nih.gov/ncbi/pubmed/in/PubMed.dtd">
<ArticleSet>
<Article>
<Journal>
				<PublisherName>انجمن زبان شناسی ایران</PublisherName>
				<JournalTitle>زبان و زبان‌شناسی</JournalTitle>
				<Issn>23223847</Issn>
				<Volume>18</Volume>
				<Issue>35</Issue>
				<PubDate PubStatus="epublish">
					<Year>2022</Year>
					<Month>06</Month>
					<Day>22</Day>
				</PubDate>
			</Journal>
<ArticleTitle>The Corpus of ATU Papers, Theses and Dissertations Abstracts</ArticleTitle>
<VernacularTitle>پیکره چکیده‌های مقالات و پایان‌نامه‌های دانشگاهی دانشگاه علامه طباطبائی</VernacularTitle>
			<FirstPage>127</FirstPage>
			<LastPage>145</LastPage>
			<ELocationID EIdType="pii">8798</ELocationID>
			
<ELocationID EIdType="doi">10.30465/lsi.2023.43633.1644</ELocationID>
			
			<Language>FA</Language>
<AuthorList>
<Author>
					<FirstName>آزاده</FirstName>
					<LastName>میرزائی</LastName>
<Affiliation>دانشگاه علامه طباطبائی</Affiliation>
<Identifier Source="ORCID">0000-0001-9956-556X</Identifier>

</Author>
<Author>
					<FirstName>فاطمه</FirstName>
					<LastName>صدقی</LastName>
<Affiliation>کارشناسی ارشد هوش مصنوعی دانشگاه الزهرا</Affiliation>

</Author>
</AuthorList>
				<PublicationType>Journal Article</PublicationType>
			<History>
				<PubDate PubStatus="received">
					<Year>2022</Year>
					<Month>11</Month>
					<Day>22</Day>
				</PubDate>
			</History>
		<Abstract>This study explains how to develop the corpus of “ATU Papers, Theses and Dissertations Abstracts” and introduces the different characteristics and features of the corpus. The corpus contains ten thousand thesis abstracts and 9538 article abstracts from the scientific journals of Allameh Tabatabai University with a volume of more than three and a half million tokens. Academic abstracts as brief authored texts with scientific content can depict special linguistic features and therefore, they are valuable documents. In this article, to express the importance of access to such data and to examine some features of the corpus, the word content of a part of the data has been examined and presented according to the concept of keyness and n-grams. The results showed that the lexical content of this corpus could lead researchers to propose some hypotheses. Also, the exploring n-grams of this corpus showed that the language of science has specific word clusters that can depict a particular type of language.</Abstract>
			<OtherAbstract Language="FA">این مقاله از نحوۀ شکل‌گیری پیکرۀ «چکیده‌های مقالات و پایان‌نامه‌های دانشگاهی دانشگاه علامه طباطبائی» و همچنین از ویژگی‌ها و امکانات آن می‌گوید. داده‌های این پیکره شامل ده هزار چکیده پایان‌نامه و 9538 چکیده مقاله (برگرفته از نشریات علمی دانشگاه علامه طباطبائی) با حجمی در حدود سه و نیم میلیون موردواژه است که در قالب طرح پژوهشی گردآوری شده‌اند. اهمیت داده‌های این پیکره یعنی چکیده‌های دانشگاهی از آن جهت است که این نوع داده‌ها به عنوان متون تألیفیِ فشرده و با محتوای علمی می‌توانند تصویرگر ویژگی‌های خاص زبان علم به عنوان گونه‌ای از زبان باشند. در این نوشتار برای بیان اهمیت دسترسی به چنین داده‌هایی و به جهت بررسی امکانات پیکره، محتوایِ واژه‌ایِ بخشی از داده با توجه به مفهوم کلیدی‌بودگی و فهرست چندپشته‌ها مورد بررسی قرار گرفت. بررسی‌ها نشان داد محتوای واژگانی این پیکره می‌تواند پژوهشگران را به سوی طرح برخی فرضیه‌ها سوق دهد. همچنین بررسی چندپشته‌های داده‌های علمی نشان داد که زبان علم دارای توالی‌های واژه‌ای مشخصی است که می‌تواند تصویرگر نوع خاصی از زبان باشد.</OtherAbstract>
		<ObjectList>
			<Object Type="keyword">
			<Param Name="value">پیکره</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">زبان علم</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">کلیدی‌بودگی</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">چندپشته</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">زبان فارسی</Param>
			</Object>
		</ObjectList>
<ArchiveCopySource DocType="pdf">https://lsi-linguistics.ihcs.ac.ir/article_8798_6a0724f1d3e80fd5f761cacb7efe8593.pdf</ArchiveCopySource>
</Article>
</ArticleSet>
