All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

Japanese Author Attribution Using BERT Finetuning with Stylometric Features

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AYZ8DATAH" target="_blank" >RIV/00216208:11320/26:YZ8DATAH - isvavai.cz</a>

  • Result on the web

    <a href="http://dx.doi.org/10.1007/978-981-96-5123-8_20" target="_blank" >http://dx.doi.org/10.1007/978-981-96-5123-8_20</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1007/978-981-96-5123-8_20" target="_blank" >10.1007/978-981-96-5123-8_20</a>

Alternative languages

  • Result language

    angličtina

  • Original language name

    Japanese Author Attribution Using BERT Finetuning with Stylometric Features

  • Original language description

    This study investigates author attribution (AA) in Japanese texts through fine-tuning the pre-trained BERT model “cl-tohoku/bert-large-japanese-v2” with Japanese-specific stylometric features. Experiments explored combinations of these features with classifiers such as LR, SVM, and RF, across varying author counts from 5 to 75. Focusing solely on native Japanese compositions, the study utilized the “Composition Bilingual Database” from the National Institute for Japanese Language and Linguistics to maintain linguistic consistency. The BERT model combined with LR achieved the highest accuracy of 96.3% for 5 authors, demonstrating deep learning’s potential in Japanese AA. However, high-dimensional stylistic features introduced noise when integrated, highlighting challenges in feature alignment. Future work will explore advanced non-linear models like XGBoost, LightGBM, and CatBoost for improved feature integration, and low-resource classification methods such as prototypical networks to enhance performance without extensive dataset expansion. Additionally, further testing of alternative Japanese pre-trained language models will be conducted to capture linguistic nuances more effectively. © The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025.

  • Czech name

  • Czech description

Classification

  • Type

    D - Article in proceedings

  • CEP classification

  • OECD FORD branch

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Result continuities

  • Project

  • Continuities

Others

  • Publication year

    2025

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • Article name in the collection

    Commun. Comput. Info. Sci.

  • ISBN

    978-981-96-5122-1

  • ISSN

  • e-ISSN

  • Number of pages

    15

  • Pages from-to

    293-307

  • Publisher name

    Springer Science and Business Media Deutschland GmbH

  • Place of publication

  • Event location

    Beijing

  • Event date

    Jan 1, 2026

  • Type of event by nationality

    WRD - Celosvětová akce

  • UT code for WoS article