Utilize este identificador para referenciar este registo: https://hdl.handle.net/10216/15175
Autor(es): Luís Sarmento
Alexander Kehlenbeck
Eugénio Oliveira
Lyle Ungar
Título: Efficient clustering of web-derived data sets
Data de publicação: 2009
Resumo: Many data sets derived from the web are large, high-dimensional, sparse and have a Zipfian distribution of both classes and features. On such data sets, current scalable clustering methods such as streaming clustering suffer from fragmentation. where large classes are incorrectly divided into many smaller clusters. and computational efficiency drops significantly. We present a new clustering algorithm based on connected components that addresses these issues and so works well oil web-type data.
Assunto: Informática, Ciências da computação e da informação
Informatics, Computer and information sciences
Áreas do conhecimento: Ciências exactas e naturais::Ciências da computação e da informação
Natural sciences::Computer and information sciences
DOI: 10.1007/978-3-642-03070-3_30
URI: https://repositorio-aberto.up.pt/handle/10216/15175
Fonte: Machine Learning and Data Mining in Pattern Recognition
Tipo de Documento: Artigo em Livro de Atas de Conferência Internacional
Condições de Acesso: openAccess
Licença: https://creativecommons.org/licenses/by-nc/4.0/
Aparece nas coleções:FEUP - Artigo em Livro de Atas de Conferência Internacional

Ficheiros deste registo:
Ficheiro Descrição TamanhoFormato 
57542.pdf276.64 kBAdobe PDFThumbnail
Ver/Abrir


Este registo está protegido por Licença Creative Commons Creative Commons