Please use this identifier to cite or link to this item: https://hdl.handle.net/10216/15175
Author(s): Luís Sarmento
Alexander Kehlenbeck
Eugénio Oliveira
Lyle Ungar
Title: Efficient clustering of web-derived data sets
Issue Date: 2009
Abstract: Many data sets derived from the web are large, high-dimensional, sparse and have a Zipfian distribution of both classes and features. On such data sets, current scalable clustering methods such as streaming clustering suffer from fragmentation. where large classes are incorrectly divided into many smaller clusters. and computational efficiency drops significantly. We present a new clustering algorithm based on connected components that addresses these issues and so works well oil web-type data.
Subject: Informática, Ciências da computação e da informação
Informatics, Computer and information sciences
Scientific areas: Ciências exactas e naturais::Ciências da computação e da informação
Natural sciences::Computer and information sciences
DOI: 10.1007/978-3-642-03070-3_30
URI: https://repositorio-aberto.up.pt/handle/10216/15175
Source: Machine Learning and Data Mining in Pattern Recognition
Document Type: Artigo em Livro de Atas de Conferência Internacional
Rights: openAccess
License: https://creativecommons.org/licenses/by-nc/4.0/
Appears in Collections:FEUP - Artigo em Livro de Atas de Conferência Internacional

Files in This Item:
File Description SizeFormat 
57542.pdf276.64 kBAdobe PDFThumbnail
View/Open


This item is licensed under a Creative Commons License Creative Commons