Please use this identifier to cite or link to this item: https://hdl.handle.net/10216/176179
Author(s): André Miguel Teixeira Pestana
Title: A Framework for Evaluating the Correctness of Programming Learning Platforms Feedback
Issue Date: 2026-07-21
Abstract: The test oracle problem has gained renewed attention with the emergence of Large Language Models (LLMs), which introduce both opportunities and challenges for automated software testing. We need to understand to what extent problem specifications can be used as oracles to guide test case generation, and whether it is possible to overcome initial difficulties. This analysis is conducted on an online academic platform. A five-step workflow is proposed: (1) collecting problem solutions and their specifications and verifying their acceptance by the platform; (2) generating test cases from the specifications; (3) deriving pre- and post-conditions from the specifications to produce a formal specification; (4)evaluating the performance of the initial prompt across other problem categories; and (5) refining the prompt to achieve improved results. Together, these steps enable a thorough assessment of the effectiveness of this approach. A review of 207 Python solutions revealed that a small fraction, roughly 1.93% (4 out of 207), were mistakenly accepted by the platform despite being incorrect. Of all test cases generated, 1,778 were deemed valid, accounting for 90.99% of the total output. Among the 207 challenges evaluated, 58 (28.0%) contained at least one failing test case, with some challenges accumulating up to 12 failures. Notably, 13 challenges achieved a 0% success rate, meaning every single generated test case failed. Finally, through prompt refinement, the Test Execution Success Rate improved substantially, rising from 54% to 70.9%. In conclusion, this dissertation explores the effectiveness of leveraging problem specifications as test oracles, and assesses whether observed shortcomings, such as high failure rates in certain challenges, can be mitigated by incorporating them as additional context to enhance oracle generation.
Description: A five-step workflow is proposed: (1) collecting problem solutions and their specifications and verifying their acceptance by the platform; (2) generating test cases from the specifications; (3) deriving pre- and post-conditions from the specifications to produce a formal specification; (4)evaluating the performance of the initial prompt across other problem categories; and (5) refining the prompt to achieve improved results. Together, these steps enable a thorough assessment of the effectiveness of this approach.
Subject: Engenharia electrotécnica, electrónica e informática
Electrical engineering, Electronic engineering, Information engineering
Scientific areas: Ciências da engenharia e tecnologias::Engenharia electrotécnica, electrónica e informática
Engineering and technology::Electrical engineering, Electronic engineering, Information engineering
URI: https://hdl.handle.net/10216/176179
Document Type: Dissertação
Rights: openAccess
Appears in Collections:FEUP - Dissertação

Files in This Item:
File Description SizeFormat 
787307.pdfA Framework for Evaluating the Correctness of Programming Learning Platforms Feedback1.3 MBAdobe PDFThumbnail
View/Open


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.