Providing Web Archive News Articles as Corpus Data
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AE4WVXG3V" target="_blank" >RIV/00216208:11320/26:E4WVXG3V - isvavai.cz</a>
Výsledek na webu
<a href="http://openhumanitiesdata.metajnl.com/articles/10.5334/johd.281/" target="_blank" >http://openhumanitiesdata.metajnl.com/articles/10.5334/johd.281/</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.5334/johd.281" target="_blank" >10.5334/johd.281</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Providing Web Archive News Articles as Corpus Data
Popis výsledku v původním jazyce
While the huge data repositories of web archives carry big potential for knowledge production in academia, researchers have described significant challenges when trying to access and make use of web archives in research. This article describes the creation of a “Web News Collection” where content from the National Library of Norway’s web archive has been made available for computational text analysis, in a manner that facilitates access for research and beyond – aligning with FAIR principles, while also accounting for copyright restrictions. Developing the warc2corpus pipeline, we detail the processes for extracting natural language from WARC files, curating content, and enhancing metadata for analytical purposes. This structured collection — consisting of 1.5 million news articles accessible via a REST API —enables distant reading of news from the web, with tools for building corpora, word frequencies and collocations. To support usage, both programming interfaces and user-friendly web apps are offered, representing a significant step forward in making web archives usable and valuable for digital scholars.
Název v anglickém jazyce
Providing Web Archive News Articles as Corpus Data
Popis výsledku anglicky
While the huge data repositories of web archives carry big potential for knowledge production in academia, researchers have described significant challenges when trying to access and make use of web archives in research. This article describes the creation of a “Web News Collection” where content from the National Library of Norway’s web archive has been made available for computational text analysis, in a manner that facilitates access for research and beyond – aligning with FAIR principles, while also accounting for copyright restrictions. Developing the warc2corpus pipeline, we detail the processes for extracting natural language from WARC files, curating content, and enhancing metadata for analytical purposes. This structured collection — consisting of 1.5 million news articles accessible via a REST API —enables distant reading of news from the web, with tools for building corpora, word frequencies and collocations. To support usage, both programming interfaces and user-friendly web apps are offered, representing a significant step forward in making web archives usable and valuable for digital scholars.
Klasifikace
Druh
J<sub>imp</sub> - Článek v periodiku v databázi Web of Science
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
Journal of Open Humanities Data
ISSN
2059-481X
e-ISSN
—
Svazek periodika
11
Číslo periodika v rámci svazku
2025-01-23
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
15
Strana od-do
2
Kód UT WoS článku
001412688800002
EID výsledku v databázi Scopus
2-s2.0-85217456978