Vše

Co hledáte?

Vše
Projekty
Výsledky výzkumu
Subjekty

Rychlé hledání

  • Projekty podpořené TA ČR
  • Významné projekty
  • Projekty s nejvyšší státní podporou
  • Aktuálně běžící projekty

Chytré vyhledávání

  • Takto najdu konkrétní +slovo
  • Takto z výsledků -slovo zcela vynechám
  • “Takto můžu najít celou frázi”

Hyph-bench: Benchmark Dataset of Hyphenated Words for Generating Hyphenation Patterns

Identifikátory výsledku

  • Kód výsledku v IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216224%3A14330%2F25%3A00142895" target="_blank" >RIV/00216224:14330/25:00142895 - isvavai.cz</a>

  • Výsledek na webu

    <a href="https://nlp.fi.muni.cz/raslan/2025/paper16.pdf" target="_blank" >https://nlp.fi.muni.cz/raslan/2025/paper16.pdf</a>

  • DOI - Digital Object Identifier

Alternativní jazyky

  • Jazyk výsledku

    angličtina

  • Název v původním jazyce

    Hyph-bench: Benchmark Dataset of Hyphenated Words for Generating Hyphenation Patterns

  • Popis výsledku v původním jazyce

    The hyphenation algorithm, based on hyphenation patterns and developed primarily for TeX, is used almost exclusively by typesetting systems, web browsers, and other applications that require breaking text lines. The essence of minimal pattern generation lies in the NP-complete task of optimizing size and coverage in pattern generation from a word list of hyphenated words, while maintaining near 100% accuracy. The problem of optimising pattern generation has already been studied for several Slavic languages; however, the heuristic setting of parameters for the pattern generation process is based on the quality and quantity of the hyphenated word list. We have designed and collected the datasets of a hyphenated word list of Slavic and non-Slavic languages with the primary goal of benchmarking and optimization of pattern generation by Patgen. We have set and computed baselines for included languages. We have prepared a benchmark dataset with baselines that often beat the currently widely used patterns in precision and/or recall. We are paving the road towards more accurate, smaller patterns with better and consistent coverage of word hyphenation in several Slavic and non-Slavic languages. The dataset also enables experiments in generating segmentation or hyphenations for language families, or even the universal, syllabic hyphenation patterns.

  • Název v anglickém jazyce

    Hyph-bench: Benchmark Dataset of Hyphenated Words for Generating Hyphenation Patterns

  • Popis výsledku anglicky

    The hyphenation algorithm, based on hyphenation patterns and developed primarily for TeX, is used almost exclusively by typesetting systems, web browsers, and other applications that require breaking text lines. The essence of minimal pattern generation lies in the NP-complete task of optimizing size and coverage in pattern generation from a word list of hyphenated words, while maintaining near 100% accuracy. The problem of optimising pattern generation has already been studied for several Slavic languages; however, the heuristic setting of parameters for the pattern generation process is based on the quality and quantity of the hyphenated word list. We have designed and collected the datasets of a hyphenated word list of Slavic and non-Slavic languages with the primary goal of benchmarking and optimization of pattern generation by Patgen. We have set and computed baselines for included languages. We have prepared a benchmark dataset with baselines that often beat the currently widely used patterns in precision and/or recall. We are paving the road towards more accurate, smaller patterns with better and consistent coverage of word hyphenation in several Slavic and non-Slavic languages. The dataset also enables experiments in generating segmentation or hyphenations for language families, or even the universal, syllabic hyphenation patterns.

Klasifikace

  • Druh

    D - Stať ve sborníku

  • CEP obor

  • OECD FORD obor

    10200 - Computer and information sciences

Návaznosti výsledku

  • Projekt

  • Návaznosti

    S - Specificky vyzkum na vysokych skolach

Ostatní

  • Rok uplatnění

    2025

  • Kód důvěrnosti údajů

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Údaje specifické pro druh výsledku

  • Název statě ve sborníku

    Recent Advances in Slavonic Natural Language Processing, RASLAN 2025

  • ISBN

    9788026318583

  • ISSN

    2336-4289

  • e-ISSN

  • Počet stran výsledku

    13

  • Strana od-do

    173-185

  • Název nakladatele

    Tribun EU

  • Místo vydání

    Brno, Czech Republic

  • Místo konání akce

    Kouty nad Desnou, Česká Republika

  • Datum konání akce

    1. 1. 2025

  • Typ akce podle státní příslušnosti

    WRD - Celosvětová akce

  • Kód UT WoS článku