All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

NER for Albanian Language: A Manually Annotated Corpus and Machine Learning Models

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3A34R253AY" target="_blank" >RIV/00216208:11320/26:34R253AY - isvavai.cz</a>

  • Result on the web

    <a href="http://dx.doi.org/10.1007/978-3-031-87769-8_14" target="_blank" >http://dx.doi.org/10.1007/978-3-031-87769-8_14</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1007/978-3-031-87769-8_14" target="_blank" >10.1007/978-3-031-87769-8_14</a>

Alternative languages

  • Result language

    angličtina

  • Original language name

    NER for Albanian Language: A Manually Annotated Corpus and Machine Learning Models

  • Original language description

    Recent advancements in artificial intelligence (AI) have significantly enhanced tasks like named entity recognition (NER), enabling the identification of people, organizations, products, events, places, and dates. This paper introduces an Albanian NER corpus with 1,003,836 tokens (56,595 sentences), including 89,850 labelled tokens, annotated with 10 NER tags. The corpus, sourced from well-known Albanian news platforms, are used to train and evaluate 10 models using algorithms such as Naïve Bayes, Logistic Regression, SVM, Random Forest, Gradient Boosting, Extreme Gradient Boosting, and Multi-Layer Perceptron variants. Among these, Extra Trees and Random Forest achieved the best results, with approximately 96% accuracy and a 95% F1 score. As the largest NER corpus in Albanian, this resource advances linguistic research and AI applications, enhancing NER tasks and advance natural language processing (NLP) developments for the Albanian language. © The Author(s), under exclusive license to Springer Nature Switzerland AG 2025.

  • Czech name

  • Czech description

Classification

  • Type

    J<sub>SC</sub> - Article in a specialist periodical, which is included in the SCOPUS database

  • CEP classification

  • OECD FORD branch

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Result continuities

  • Project

  • Continuities

Others

  • Publication year

    2025

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • Name of the periodical

    Lecture Notes on Data Engineering and Communications Technologies

  • ISSN

    23674512

  • e-ISSN

  • Volume of the periodical

    247

  • Issue of the periodical within the volume

    2025

  • Country of publishing house

    US - UNITED STATES

  • Number of pages

    13

  • Pages from-to

    153-165

  • UT code for WoS article

  • EID of the result in the Scopus database

    2-s2.0-105003097414