Please use this identifier to cite or link to this item:
https://hdl.handle.net/10216/166467| Author(s): | Correia, PG Henrique Lopes Cardoso |
| Title: | Towards Explaining Shortcut Learning Through Attention Visualization and Adversarial Attacks |
| Issue Date: | 2023 |
| Abstract: | Since its introduction, the attention-based Transformer architecture has become the de facto standard for building models with state-of-the-art performance on many Natural Language Processing tasks. However, it seems that the success of these models might have to do with their exploitation of dataset artifacts, rendering them unable to generalize to other data and vulnerable to adversarial attacks. On the other hand, the attention mechanism present in all models based on the Transformer, such as BERT-based ones, has been seen by many as a potential way to explain these deep learning models: by visualizing attention weights, it might be possible to gain insights on the reasons behind these opaque models' decisions. This paper introduces Attentive-BERT, an interactive attention weights visualization tool for diagnosing BERT-based models, focusing on explaining the occurrence of shortcut learning. The distinctive feature of this tool is enabling the visual comparison of attention weights before and after a change to the model's input, in order to visually analyse adversarial attacks. Some illustrations of this use case are explored in this paper. |
| DOI: | 10.1007/978-3-031-34204-2_45 |
| URI: | https://hdl.handle.net/10216/166467 |
| Source: | Communications in Computer and Information Science |
| Document Type: | Artigo em Livro de Atas de Conferência Internacional |
| Rights: | openAccess |
| Appears in Collections: | FEUP - Artigo em Livro de Atas de Conferência Internacional |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| 668667.pdf | 1.37 MB | Adobe PDF | ![]() View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.
