Developing, Compiling and Annotating Corpora for the Persian Language
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AZLV5VWC8" target="_blank" >RIV/00216208:11320/26:ZLV5VWC8 - isvavai.cz</a>
Result on the web
<a href="http://dx.doi.org/10.1007/978-3-031-98989-6_2" target="_blank" >http://dx.doi.org/10.1007/978-3-031-98989-6_2</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/978-3-031-98989-6_2" target="_blank" >10.1007/978-3-031-98989-6_2</a>
Alternative languages
Result language
angličtina
Original language name
Developing, Compiling and Annotating Corpora for the Persian Language
Original language description
In this chapter, we briefly overview the general criteria that have to be taken into consideration while developing a corpus. Developing a corpus for the Persian language is challenging. In this chapter, the challenges are discussed and categorized. Then, we discuss the steps that have to be taken to make the developed corpus usable for natural language processing techniques. Furthermore, we explain how the data can be annotated. In the rest of this chapter, the sketch of supervised, semi-supervised, and unsupervised machine learning methods for data annotation is briefly explained. We collect 167 research papers that developed a corpus for their study; then, we categorize them based on the task and the corpus size. The annotated data has to be standardized. We briefly introduce the major standards used for structuring data. © 2025 The Editor(s) (if applicable) and The Author(s), under exclusive license to Springer Nature Switzerland AG.
Czech name
—
Czech description
—
Classification
Type
C - Chapter in a specialist book
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
—
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Book/collection name
New Frontiers in Corpus Based Studies of Persian: Challenges, Innovations and Applications
ISBN
978-3-031-98989-6
Number of pages of the result
38
Pages from-to
25-62
Number of pages of the book
263
Publisher name
Springer Science+Business Media
Place of publication
—
UT code for WoS chapter
—