Vše

Co hledáte?

Vše
Projekty
Výsledky výzkumu
Subjekty

Rychlé hledání

  • Projekty podpořené TA ČR
  • Významné projekty
  • Projekty s nejvyšší státní podporou
  • Aktuálně běžící projekty

Chytré vyhledávání

  • Takto najdu konkrétní +slovo
  • Takto z výsledků -slovo zcela vynechám
  • “Takto můžu najít celou frázi”

Japanese Author Attribution Using BERT Finetuning with Stylometric Features

Identifikátory výsledku

  • Kód výsledku v IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AYZ8DATAH" target="_blank" >RIV/00216208:11320/26:YZ8DATAH - isvavai.cz</a>

  • Výsledek na webu

    <a href="http://dx.doi.org/10.1007/978-981-96-5123-8_20" target="_blank" >http://dx.doi.org/10.1007/978-981-96-5123-8_20</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1007/978-981-96-5123-8_20" target="_blank" >10.1007/978-981-96-5123-8_20</a>

Alternativní jazyky

  • Jazyk výsledku

    angličtina

  • Název v původním jazyce

    Japanese Author Attribution Using BERT Finetuning with Stylometric Features

  • Popis výsledku v původním jazyce

    This study investigates author attribution (AA) in Japanese texts through fine-tuning the pre-trained BERT model “cl-tohoku/bert-large-japanese-v2” with Japanese-specific stylometric features. Experiments explored combinations of these features with classifiers such as LR, SVM, and RF, across varying author counts from 5 to 75. Focusing solely on native Japanese compositions, the study utilized the “Composition Bilingual Database” from the National Institute for Japanese Language and Linguistics to maintain linguistic consistency. The BERT model combined with LR achieved the highest accuracy of 96.3% for 5 authors, demonstrating deep learning’s potential in Japanese AA. However, high-dimensional stylistic features introduced noise when integrated, highlighting challenges in feature alignment. Future work will explore advanced non-linear models like XGBoost, LightGBM, and CatBoost for improved feature integration, and low-resource classification methods such as prototypical networks to enhance performance without extensive dataset expansion. Additionally, further testing of alternative Japanese pre-trained language models will be conducted to capture linguistic nuances more effectively. © The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025.

  • Název v anglickém jazyce

    Japanese Author Attribution Using BERT Finetuning with Stylometric Features

  • Popis výsledku anglicky

    This study investigates author attribution (AA) in Japanese texts through fine-tuning the pre-trained BERT model “cl-tohoku/bert-large-japanese-v2” with Japanese-specific stylometric features. Experiments explored combinations of these features with classifiers such as LR, SVM, and RF, across varying author counts from 5 to 75. Focusing solely on native Japanese compositions, the study utilized the “Composition Bilingual Database” from the National Institute for Japanese Language and Linguistics to maintain linguistic consistency. The BERT model combined with LR achieved the highest accuracy of 96.3% for 5 authors, demonstrating deep learning’s potential in Japanese AA. However, high-dimensional stylistic features introduced noise when integrated, highlighting challenges in feature alignment. Future work will explore advanced non-linear models like XGBoost, LightGBM, and CatBoost for improved feature integration, and low-resource classification methods such as prototypical networks to enhance performance without extensive dataset expansion. Additionally, further testing of alternative Japanese pre-trained language models will be conducted to capture linguistic nuances more effectively. © The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025.

Klasifikace

  • Druh

    D - Stať ve sborníku

  • CEP obor

  • OECD FORD obor

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Návaznosti výsledku

  • Projekt

  • Návaznosti

Ostatní

  • Rok uplatnění

    2025

  • Kód důvěrnosti údajů

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Údaje specifické pro druh výsledku

  • Název statě ve sborníku

    Commun. Comput. Info. Sci.

  • ISBN

    978-981-96-5122-1

  • ISSN

  • e-ISSN

  • Počet stran výsledku

    15

  • Strana od-do

    293-307

  • Název nakladatele

    Springer Science and Business Media Deutschland GmbH

  • Místo vydání

  • Místo konání akce

    Beijing

  • Datum konání akce

    1. 1. 2026

  • Typ akce podle státní příslušnosti

    WRD - Celosvětová akce

  • Kód UT WoS článku