Computational Linguistics for Ho Language
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F25%3AS6GHRNH9" target="_blank" >RIV/00216208:11320/25:S6GHRNH9 - isvavai.cz</a>
Result on the web
<a href="https://www.scopus.com/inward/record.uri?eid=2-s2.0-85205121593&doi=10.1007%2f978-981-97-5204-1_7&partnerID=40&md5=8f1fab9a5fe894630c22fe09ec52f834" target="_blank" >https://www.scopus.com/inward/record.uri?eid=2-s2.0-85205121593&doi=10.1007%2f978-981-97-5204-1_7&partnerID=40&md5=8f1fab9a5fe894630c22fe09ec52f834</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/978-981-97-5204-1_7" target="_blank" >10.1007/978-981-97-5204-1_7</a>
Alternative languages
Result language
angličtina
Original language name
Computational Linguistics for Ho Language
Original language description
Computational linguistics uses computer science methods to analyze and synthesize language and speech, crucial for understanding human languages. In India, 8.6% of the population is tribal, with different languages, cultures, and customs. English is widely used for communication, but many are deficient in English. Machine translations help overcome language barriers. The Ho language is spoken by the Ho, Kolha, Kol, and Munda tribes. It has a distinct culture, literature, and script. This is the first attempt to create computational linguistics for the Ho language, allowing both Ho and non-Ho speakers to learn and gain knowledge. In our research work, we have prepared a parallel corpus (English-Ho), which contains small sentences and words. The corpus size of the Ho language is twelve thousand five hundred sentence pairs, and four thousand three-hundred word pairs were used for training and testing. We have used Statistical Machine Translation (SMT) and Neural Machine Translation (NMT) to experiment, having an accuracy of around 81%. The accuracy was measured by using the BLEU and ChrF scores. We examined automatic text summarization of Ho language using machine learning algorithms to reduce long texts into small sections of relevant sentences. © The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2024.
Czech name
—
Czech description
—
Classification
Type
C - Chapter in a specialist book
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
—
Others
Publication year
2024
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Book/collection name
Intelligent Technologies
ISBN
978-981-9752-03-4
Number of pages of the result
23
Pages from-to
139-161
Number of pages of the book
395
Publisher name
Springer Singapore
Place of publication
—
UT code for WoS chapter
—