All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

Scalable Similarity Joins for Fast and Accurate Record Deduplication in Big Data

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216275%3A25530%2F24%3A39921108" target="_blank" >RIV/00216275:25530/24:39921108 - isvavai.cz</a>

  • Result on the web

    <a href="https://link.springer.com/chapter/10.1007/978-3-031-60328-0_18" target="_blank" >https://link.springer.com/chapter/10.1007/978-3-031-60328-0_18</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1007/978-3-031-60328-0_18" target="_blank" >10.1007/978-3-031-60328-0_18</a>

Alternative languages

  • Result language

    angličtina

  • Original language name

    Scalable Similarity Joins for Fast and Accurate Record Deduplication in Big Data

  • Original language description

    Record linkage is the process of matching records from multiple data sources that refer to the same entities. When applied to a single data source, this process is known as deduplication. With the increasing size of data source, recently referred to as big data, the complexity of the matching process becomes one of the major challenges for record linkage and deduplication. In recent decades, several blocking, indexing and filtering techniques have been developed. Their purpose is to reduce the number of record pairs to be compared by removing obvious non-matching pairs in the deduplication process, while maintaining high quality of matching. Currently developed algorithms and traditional techniques are not efficient, using methods that still lose significant proportion of true matches when removing comparison pairs. This paper proposes more efficient algorithms for removing non-matching pairs, with an explicitly proven mathematical lower bound on recently used stateof-the-art approximate string matching method - Fuzzy Jaccard Similarity. The algorithm is also much more efficient in classification using Density-based spatial clustering of applications with noise (DBSCAN) in log-linear time complexity O(|E| log(|E|)).

  • Czech name

  • Czech description

Classification

  • Type

    D - Article in proceedings

  • CEP classification

  • OECD FORD branch

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Result continuities

  • Project

  • Continuities

    S - Specificky vyzkum na vysokych skolach

Others

  • Publication year

    2024

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • Article name in the collection

    Good Practices and New Perspectives in Information Systems and Technologies : WorldCIST 2024, Volume 6

  • ISBN

    978-3-031-60327-3

  • ISSN

    2367-3370

  • e-ISSN

    2367-3389

  • Number of pages

    11

  • Pages from-to

    "181 "- 191

  • Publisher name

    Springer Nature Switzerland AG

  • Place of publication

    Cham

  • Event location

    Lodž

  • Event date

    Mar 26, 2024

  • Type of event by nationality

    EUR - Evropská akce

  • UT code for WoS article

    001267244400018