Vše

Co hledáte?

Vše
Projekty
Výsledky výzkumu
Subjekty

Rychlé hledání

  • Projekty podpořené TA ČR
  • Významné projekty
  • Projekty s nejvyšší státní podporou
  • Aktuálně běžící projekty

Chytré vyhledávání

  • Takto najdu konkrétní +slovo
  • Takto z výsledků -slovo zcela vynechám
  • “Takto můžu najít celou frázi”

Developing, Compiling and Annotating Corpora for the Persian Language

Identifikátory výsledku

  • Kód výsledku v IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AZLV5VWC8" target="_blank" >RIV/00216208:11320/26:ZLV5VWC8 - isvavai.cz</a>

  • Výsledek na webu

    <a href="http://dx.doi.org/10.1007/978-3-031-98989-6_2" target="_blank" >http://dx.doi.org/10.1007/978-3-031-98989-6_2</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1007/978-3-031-98989-6_2" target="_blank" >10.1007/978-3-031-98989-6_2</a>

Alternativní jazyky

  • Jazyk výsledku

    angličtina

  • Název v původním jazyce

    Developing, Compiling and Annotating Corpora for the Persian Language

  • Popis výsledku v původním jazyce

    In this chapter, we briefly overview the general criteria that have to be taken into consideration while developing a corpus. Developing a corpus for the Persian language is challenging. In this chapter, the challenges are discussed and categorized. Then, we discuss the steps that have to be taken to make the developed corpus usable for natural language processing techniques. Furthermore, we explain how the data can be annotated. In the rest of this chapter, the sketch of supervised, semi-supervised, and unsupervised machine learning methods for data annotation is briefly explained. We collect 167 research papers that developed a corpus for their study; then, we categorize them based on the task and the corpus size. The annotated data has to be standardized. We briefly introduce the major standards used for structuring data. © 2025 The Editor(s) (if applicable) and The Author(s), under exclusive license to Springer Nature Switzerland AG.

  • Název v anglickém jazyce

    Developing, Compiling and Annotating Corpora for the Persian Language

  • Popis výsledku anglicky

    In this chapter, we briefly overview the general criteria that have to be taken into consideration while developing a corpus. Developing a corpus for the Persian language is challenging. In this chapter, the challenges are discussed and categorized. Then, we discuss the steps that have to be taken to make the developed corpus usable for natural language processing techniques. Furthermore, we explain how the data can be annotated. In the rest of this chapter, the sketch of supervised, semi-supervised, and unsupervised machine learning methods for data annotation is briefly explained. We collect 167 research papers that developed a corpus for their study; then, we categorize them based on the task and the corpus size. The annotated data has to be standardized. We briefly introduce the major standards used for structuring data. © 2025 The Editor(s) (if applicable) and The Author(s), under exclusive license to Springer Nature Switzerland AG.

Klasifikace

  • Druh

    C - Kapitola v odborné knize

  • CEP obor

  • OECD FORD obor

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Návaznosti výsledku

  • Projekt

  • Návaznosti

Ostatní

  • Rok uplatnění

    2025

  • Kód důvěrnosti údajů

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Údaje specifické pro druh výsledku

  • Název knihy nebo sborníku

    New Frontiers in Corpus Based Studies of Persian: Challenges, Innovations and Applications

  • ISBN

    978-3-031-98989-6

  • Počet stran výsledku

    38

  • Strana od-do

    25-62

  • Počet stran knihy

    263

  • Název nakladatele

    Springer Science+Business Media

  • Místo vydání

  • Kód UT WoS kapitoly