TrTok: A Fast and Trainable Tokenizer for Natural Languages
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F12%3A10129918" target="_blank" >RIV/00216208:11320/12:10129918 - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.2478/v10108-012-0010-0" target="_blank" >http://dx.doi.org/10.2478/v10108-012-0010-0</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.2478/v10108-012-0010-0" target="_blank" >10.2478/v10108-012-0010-0</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
TrTok: A Fast and Trainable Tokenizer for Natural Languages
Popis výsledku v původním jazyce
We present a universal data-driven tool for segmenting and tokenizing text. The presented tokenizer lets the user define where token and sentence boundaries should be considered. These instances are then judged by a classifier which is trained from provided tokenized data. The features passed to the classifier are also defined by the user making, e.g., the inclusion of abbreviation lists trivial. This level of customizability makes the tokenizer a versatile tool which we show is capable of sentence detection in English text as well as word segmentation in Chinese text. In the case of English sentence detection, the system outperforms previous methods. The software is available as an open-source project on GitHub
Název v anglickém jazyce
TrTok: A Fast and Trainable Tokenizer for Natural Languages
Popis výsledku anglicky
We present a universal data-driven tool for segmenting and tokenizing text. The presented tokenizer lets the user define where token and sentence boundaries should be considered. These instances are then judged by a classifier which is trained from provided tokenized data. The features passed to the classifier are also defined by the user making, e.g., the inclusion of abbreviation lists trivial. This level of customizability makes the tokenizer a versatile tool which we show is capable of sentence detection in English text as well as word segmentation in Chinese text. In the case of English sentence detection, the system outperforms previous methods. The software is available as an open-source project on GitHub
Klasifikace
Druh
J<sub>x</sub> - Nezařazeno - Článek v odborném periodiku (Jimp, Jsc a Jost)
CEP obor
AI - Jazykověda
OECD FORD obor
—
Návaznosti výsledku
Projekt
—
Návaznosti
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2012
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
The Prague Bulletin of Mathematical Linguistics
ISSN
0032-6585
e-ISSN
—
Svazek periodika
98
Číslo periodika v rámci svazku
1
Stát vydavatele periodika
CZ - Česká republika
Počet stran výsledku
11
Strana od-do
75-85
Kód UT WoS článku
—
EID výsledku v databázi Scopus
—