Please use this identifier to cite or link to this item: https://hdl.handle.net/10216/166467
Author(s): Correia, PG
Henrique Lopes Cardoso
Title: Towards Explaining Shortcut Learning Through Attention Visualization and Adversarial Attacks
Issue Date: 2023
Abstract: Since its introduction, the attention-based Transformer architecture has become the de facto standard for building models with state-of-the-art performance on many Natural Language Processing tasks. However, it seems that the success of these models might have to do with their exploitation of dataset artifacts, rendering them unable to generalize to other data and vulnerable to adversarial attacks. On the other hand, the attention mechanism present in all models based on the Transformer, such as BERT-based ones, has been seen by many as a potential way to explain these deep learning models: by visualizing attention weights, it might be possible to gain insights on the reasons behind these opaque models' decisions. This paper introduces Attentive-BERT, an interactive attention weights visualization tool for diagnosing BERT-based models, focusing on explaining the occurrence of shortcut learning. The distinctive feature of this tool is enabling the visual comparison of attention weights before and after a change to the model's input, in order to visually analyse adversarial attacks. Some illustrations of this use case are explored in this paper.
DOI: 10.1007/978-3-031-34204-2_45
URI: https://hdl.handle.net/10216/166467
Source: Communications in Computer and Information Science
Document Type: Artigo em Livro de Atas de Conferência Internacional
Rights: openAccess
Appears in Collections:FEUP - Artigo em Livro de Atas de Conferência Internacional

Files in This Item:
File Description SizeFormat 
668667.pdf1.37 MBAdobe PDFThumbnail
View/Open


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.