Please use this identifier to cite or link to this item:
https://hdl.handle.net/10216/137327| Author(s): | João Adriano Portela de Matos Silva |
| Title: | Ensembles de OCRs para aplicações médicas |
| Issue Date: | 2021-10-14 |
| Abstract: | With the increasing use of new technologies in research, there is an enormous advantage in extracting data and knowledge stored in traditional media such as books and written records into databases. This transition is necessary because it facilitates all processes that involve the handling and processing of data on a large scale. One of the cases where this transition is necessary is the case of the "Child and Youth Health Bolentins", which is the case that this dissertation will focus on. These documents contain the information of users from birth to 20 years old. The information is stored in a table per document, and in the following pages the same information is shown in graphs. There is a great deal of information contained in these bulletins that is of interest to the scientific community to be transposed to digital media, so that it can be used in pediatric studies. What was aimed in the dissertation is to achieve an automated process through an Optical Character Recognition (OCR) system, associated with machine learning, data mining and also using Ensembles methods, in order to collect the data contained in the bulletins, obtaining the best possible predictive performance of the algorithms used. |
| Description: | Com a crescente utilização de novas tecnologias na investigação, torna-se extremamente vantajoso a extracção automática de dados e conhecimento disponível (explícita ou implícita) na informação guardada em meios tradicionais, como livros e registos escritos, para bases de dados. Esta transição é necessária pois facilita todos os processos que impliquem o tratamento e processamento de dados à grande escala. Um dos casos em que essa transição é necessária, é o caso dos "Bolentins de Saúde Infantil e Juvenil", sendo este o caso que a dissertação vai incidir. Estes documentos contêm a informação de utentes desde que nascem até aos 20 anos. A informação é guardada numa tabela por documento e, nas páginas seguintes, a mesma informação é demonstrada em gráficos. Existe uma grande quantidade de informação contida nestes boletins que é do interesse da comunidade científica que seja transposta para o meio digital, de forma a poder ser utilizada em estudos pediátricos. O que foi almejado na dissertação é conseguir um processo automatizado através de um sistema de reconhecimento de caracteres óptico (Optical Character Recognition - OCR), associado a aprendizagem computacional (Machine Learning), mineração de dados e também utilizando métodos de Ensembles, de forma a recolher os dados contidos nos boletins, obtendo a melhor performance preditiva dos algoritmos utilizados possível. |
| Subject: | Engenharia electrotécnica, electrónica e informática Electrical engineering, Electronic engineering, Information engineering |
| Scientific areas: | Ciências da engenharia e tecnologias::Engenharia electrotécnica, electrónica e informática Engineering and technology::Electrical engineering, Electronic engineering, Information engineering |
| DOI: | 10.34626/dmvs-1x48 |
| TID identifier: | 202820912 |
| URI: | https://hdl.handle.net/10216/137327 |
| Document Type: | Dissertação |
| Rights: | openAccess |
| Appears in Collections: | FEUP - Dissertação |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| 512320.pdf | Ensembles de OCRs para aplicações médicas | 5.35 MB | Adobe PDF | ![]() View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.
