Utilize este identificador para referenciar este registo:
https://hdl.handle.net/10216/163698Registo completo
| Campo DC | Valor | Idioma |
|---|---|---|
| dc.creator | Ali, Felermino | |
| dc.creator | Cardoso, Henrique Lopes | |
| dc.creator | Sousa-Silva, Rui | |
| dc.date.accessioned | 2024-12-13T00:16:45Z | - |
| dc.date.available | 2024-12-13T00:16:45Z | - |
| dc.date.issued | 2024 | |
| dc.identifier.other | sigarra:698921 | |
| dc.identifier.uri | https://hdl.handle.net/10216/163698 | - |
| dc.description.abstract | The accurate identification of loanwords within a given text holds significant potential as a valuable tool for addressing data augmentation and mitigating data sparsity issues. Such identification can improve the performance of various natural language processing tasks, particularly in the context of low-resource languages that lack standardized spelling conventions.This research proposes a supervised method to identify loanwords in Emakhuwa, borrowed from Portuguese. Our methodology encompasses a two-fold approach. Firstly, we employ traditional machine learning algorithms incorporating handcrafted features, including language-specific and similarity-based features. We build upon prior studies to extract similarity features and propose utilizing two external resources: a Sequence-to-Sequence model and a dictionary. This innovative approach allows us to identify loanwords solely by analyzing the target word without prior knowledge about its donor counterpart. Furthermore, we fine-tune the pre-trained CANINE model for the downstream task of loanword detection, which culminates in the impressive achievement of the F1-score of 93%. To the best of our knowledge, this study is the first of its kind focusing on Emakhuwa, and the preliminary results are promising as they pave the way to further advancements. | |
| dc.language.iso | eng | |
| dc.relation.ispartof | Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) | |
| dc.rights | openAccess | |
| dc.title | Detecting loanwords in Emakhuwa: an extremely low-resource bantu language exhibiting significant borrowing from portuguese | |
| dc.type | Artigo em Livro de Atas de Conferência Internacional | |
| dc.contributor.uporto | Faculdade de Engenharia | |
| dc.contributor.uporto | Faculdade de Letras | |
| Aparece nas coleções: | FEUP - Artigo em Livro de Atas de Conferência Internacional FLUP - Artigo em Livro de Atas de Conferência Internacional | |
Ficheiros deste registo:
| Ficheiro | Descrição | Tamanho | Formato | |
|---|---|---|---|---|
| 698921.pdf | 396.21 kB | Adobe PDF | ![]() Ver/Abrir |
Todos os registos no repositório estão protegidos por leis de copyright, com todos os direitos reservados.
