All

What are you looking for?

All

Projects

Results

Organizations

Quick search

Projects supported by TA ČR
Excellent projects
Projects with the highest public support
Current projects

Smart search

That is how I find a specific +word
That is how I leave the -word out of the results
“That is how I can find the whole phrase”

EN

Čeština English

RIV/00216224:14330/22:00127488

Utok: The Fast Rule-based Tokenizer

The result's identifiers

Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216224%3A14330%2F22%3A00127488" target="_blank" >RIV/00216224:14330/22:00127488 - isvavai.cz</a>
Result on the web
<a href="https://nlp.fi.muni.cz/raslan/2022/paper24.pdf" target="_blank" >https://nlp.fi.muni.cz/raslan/2022/paper24.pdf</a>
DOI - Digital Object Identifier
—

Alternative languages

Result language
angličtina
Original language name
Utok: The Fast Rule-based Tokenizer
Original language description
Tokenization is one of the first processing steps in most natural language processing applications. The papper introduces a new tokenizer Utok which follows the Unitok tokenizer in the form of simplicity of configuration for different languages and is much faster in processing speed.
Czech name
—
Czech description
—

Classification

Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
10200 - Computer and information sciences

Result continuities

Project
<a href="/en/project/LM2018101" target="_blank" >LM2018101: Digital Research Infrastructure for the Language Technologies, Arts and Humanities</a><br>
Continuities
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)

Others

Publication year
2022
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

Article name in the collection
Proceedings of the Sixteenth Workshop on Recent Advances in Slavonic Natural Languages Processing, RASLAN 2022
ISBN
9788026317524
ISSN
2336-4289
e-ISSN
—
Number of pages
6
Pages from-to
149-154
Publisher name
Tribun EU
Place of publication
Brno
Event location
Brno
Event date
Jan 1, 2022
Type of event by nationality
CST - Celostátní akce
UT code for WoS article
—

Similar results(10)

Where are we Still Split on Tokenization?Trainable Tokenizer UNLT: Urdu Natural Language Toolkit