TrTok: A Fast and Trainable Tokenizer for Natural Languages
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F12%3A10129918" target="_blank" >RIV/00216208:11320/12:10129918 - isvavai.cz</a>
Result on the web
<a href="http://dx.doi.org/10.2478/v10108-012-0010-0" target="_blank" >http://dx.doi.org/10.2478/v10108-012-0010-0</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.2478/v10108-012-0010-0" target="_blank" >10.2478/v10108-012-0010-0</a>
Alternative languages
Result language
angličtina
Original language name
TrTok: A Fast and Trainable Tokenizer for Natural Languages
Original language description
We present a universal data-driven tool for segmenting and tokenizing text. The presented tokenizer lets the user define where token and sentence boundaries should be considered. These instances are then judged by a classifier which is trained from provided tokenized data. The features passed to the classifier are also defined by the user making, e.g., the inclusion of abbreviation lists trivial. This level of customizability makes the tokenizer a versatile tool which we show is capable of sentence detection in English text as well as word segmentation in Chinese text. In the case of English sentence detection, the system outperforms previous methods. The software is available as an open-source project on GitHub
Czech name
—
Czech description
—
Classification
Type
J<sub>x</sub> - Unclassified - Peer-reviewed scientific article (Jimp, Jsc and Jost)
CEP classification
AI - Linguistics
OECD FORD branch
—
Result continuities
Project
—
Continuities
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Others
Publication year
2012
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
The Prague Bulletin of Mathematical Linguistics
ISSN
0032-6585
e-ISSN
—
Volume of the periodical
98
Issue of the periodical within the volume
1
Country of publishing house
CZ - CZECH REPUBLIC
Number of pages
11
Pages from-to
75-85
UT code for WoS article
—
EID of the result in the Scopus database
—