Japanese Author Attribution Using BERT Finetuning with Stylometric Features
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AYZ8DATAH" target="_blank" >RIV/00216208:11320/26:YZ8DATAH - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.1007/978-981-96-5123-8_20" target="_blank" >http://dx.doi.org/10.1007/978-981-96-5123-8_20</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/978-981-96-5123-8_20" target="_blank" >10.1007/978-981-96-5123-8_20</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Japanese Author Attribution Using BERT Finetuning with Stylometric Features
Popis výsledku v původním jazyce
This study investigates author attribution (AA) in Japanese texts through fine-tuning the pre-trained BERT model “cl-tohoku/bert-large-japanese-v2” with Japanese-specific stylometric features. Experiments explored combinations of these features with classifiers such as LR, SVM, and RF, across varying author counts from 5 to 75. Focusing solely on native Japanese compositions, the study utilized the “Composition Bilingual Database” from the National Institute for Japanese Language and Linguistics to maintain linguistic consistency. The BERT model combined with LR achieved the highest accuracy of 96.3% for 5 authors, demonstrating deep learning’s potential in Japanese AA. However, high-dimensional stylistic features introduced noise when integrated, highlighting challenges in feature alignment. Future work will explore advanced non-linear models like XGBoost, LightGBM, and CatBoost for improved feature integration, and low-resource classification methods such as prototypical networks to enhance performance without extensive dataset expansion. Additionally, further testing of alternative Japanese pre-trained language models will be conducted to capture linguistic nuances more effectively. © The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025.
Název v anglickém jazyce
Japanese Author Attribution Using BERT Finetuning with Stylometric Features
Popis výsledku anglicky
This study investigates author attribution (AA) in Japanese texts through fine-tuning the pre-trained BERT model “cl-tohoku/bert-large-japanese-v2” with Japanese-specific stylometric features. Experiments explored combinations of these features with classifiers such as LR, SVM, and RF, across varying author counts from 5 to 75. Focusing solely on native Japanese compositions, the study utilized the “Composition Bilingual Database” from the National Institute for Japanese Language and Linguistics to maintain linguistic consistency. The BERT model combined with LR achieved the highest accuracy of 96.3% for 5 authors, demonstrating deep learning’s potential in Japanese AA. However, high-dimensional stylistic features introduced noise when integrated, highlighting challenges in feature alignment. Future work will explore advanced non-linear models like XGBoost, LightGBM, and CatBoost for improved feature integration, and low-resource classification methods such as prototypical networks to enhance performance without extensive dataset expansion. Additionally, further testing of alternative Japanese pre-trained language models will be conducted to capture linguistic nuances more effectively. © The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Commun. Comput. Info. Sci.
ISBN
978-981-96-5122-1
ISSN
—
e-ISSN
—
Počet stran výsledku
15
Strana od-do
293-307
Název nakladatele
Springer Science and Business Media Deutschland GmbH
Místo vydání
—
Místo konání akce
Beijing
Datum konání akce
1. 1. 2026
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—