Vše

Co hledáte?

Vše
Projekty
Výsledky výzkumu
Subjekty

Rychlé hledání

  • Projekty podpořené TA ČR
  • Významné projekty
  • Projekty s nejvyšší státní podporou
  • Aktuálně běžící projekty

Chytré vyhledávání

  • Takto najdu konkrétní +slovo
  • Takto z výsledků -slovo zcela vynechám
  • “Takto můžu najít celou frázi”

Lemmatization of Czech and Croatian Noun Clusters for Terminology Extraction

Identifikátory výsledku

  • Kód výsledku v IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216224%3A14330%2F25%3A00142890" target="_blank" >RIV/00216224:14330/25:00142890 - isvavai.cz</a>

  • Výsledek na webu

    <a href="https://nlp.fi.muni.cz/raslan/2025/" target="_blank" >https://nlp.fi.muni.cz/raslan/2025/</a>

  • DOI - Digital Object Identifier

Alternativní jazyky

  • Jazyk výsledku

    angličtina

  • Název v původním jazyce

    Lemmatization of Czech and Croatian Noun Clusters for Terminology Extraction

  • Popis výsledku v původním jazyce

    During terminology extraction, terms discovered in corpora are presented in their canonical form. Lemmatization of multi-word terms consisting of noun clusters can be ambiguous due to the lack of information on their internal structure. In this paper, we show that grammatical case alone is often not sufficient for the construction of canonical forms of noun clusters. We focus on two-noun clusters in the genitive, which are the most frequent type with ambiguous parsing. Based on corpus research, we design rules that make use of multiple morphological categories to improve the lemmatization of noun clusters found in Czech and Croatian corpora. In addition to case, we also take note of gender, animacy, and whether the noun is a proper noun. The improvements lead to more accurate and more unified forms of the terms produced during terminology extraction for these two languages in Sketch Engine.

  • Název v anglickém jazyce

    Lemmatization of Czech and Croatian Noun Clusters for Terminology Extraction

  • Popis výsledku anglicky

    During terminology extraction, terms discovered in corpora are presented in their canonical form. Lemmatization of multi-word terms consisting of noun clusters can be ambiguous due to the lack of information on their internal structure. In this paper, we show that grammatical case alone is often not sufficient for the construction of canonical forms of noun clusters. We focus on two-noun clusters in the genitive, which are the most frequent type with ambiguous parsing. Based on corpus research, we design rules that make use of multiple morphological categories to improve the lemmatization of noun clusters found in Czech and Croatian corpora. In addition to case, we also take note of gender, animacy, and whether the noun is a proper noun. The improvements lead to more accurate and more unified forms of the terms produced during terminology extraction for these two languages in Sketch Engine.

Klasifikace

  • Druh

    D - Stať ve sborníku

  • CEP obor

  • OECD FORD obor

    10200 - Computer and information sciences

Návaznosti výsledku

  • Projekt

    <a href="/cs/project/LM2023062" target="_blank" >LM2023062: Digitální výzkumná infrastruktura pro jazykové technologie, umění a humanitní vědy</a><br>

  • Návaznosti

    P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)

Ostatní

  • Rok uplatnění

    2025

  • Kód důvěrnosti údajů

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Údaje specifické pro druh výsledku

  • Název statě ve sborníku

    Recent Advances in Slavonic Natural Language Processing, RASLAN 2025

  • ISBN

    9788026318583

  • ISSN

    2336-4289

  • e-ISSN

  • Počet stran výsledku

    8

  • Strana od-do

    165-172

  • Název nakladatele

    Tribun EU

  • Místo vydání

    Brno, Czech Republic

  • Místo konání akce

    Kouty nad Desnou, Česká Republika

  • Datum konání akce

    1. 1. 2025

  • Typ akce podle státní příslušnosti

    WRD - Celosvětová akce

  • Kód UT WoS článku