Please use this identifier to cite or link to this item: https://hdl.handle.net/10216/176179
Full metadata record
DC FieldValueLanguage
dc.creatorAndré Miguel Teixeira Pestana
dc.date.accessioned2026-08-07T01:51:48Z-
dc.date.available2026-08-07T01:51:48Z-
dc.date.issued2026-07-21
dc.date.submitted2026-07-31
dc.identifier.othersigarra:787307
dc.identifier.urihttps://hdl.handle.net/10216/176179-
dc.descriptionA five-step workflow is proposed: (1) collecting problem solutions and their specifications and verifying their acceptance by the platform; (2) generating test cases from the specifications; (3) deriving pre- and post-conditions from the specifications to produce a formal specification; (4)evaluating the performance of the initial prompt across other problem categories; and (5) refining the prompt to achieve improved results. Together, these steps enable a thorough assessment of the effectiveness of this approach.
dc.description.abstractThe test oracle problem has gained renewed attention with the emergence of Large Language Models (LLMs), which introduce both opportunities and challenges for automated software testing. We need to understand to what extent problem specifications can be used as oracles to guide test case generation, and whether it is possible to overcome initial difficulties. This analysis is conducted on an online academic platform. A five-step workflow is proposed: (1) collecting problem solutions and their specifications and verifying their acceptance by the platform; (2) generating test cases from the specifications; (3) deriving pre- and post-conditions from the specifications to produce a formal specification; (4)evaluating the performance of the initial prompt across other problem categories; and (5) refining the prompt to achieve improved results. Together, these steps enable a thorough assessment of the effectiveness of this approach. A review of 207 Python solutions revealed that a small fraction, roughly 1.93% (4 out of 207), were mistakenly accepted by the platform despite being incorrect. Of all test cases generated, 1,778 were deemed valid, accounting for 90.99% of the total output. Among the 207 challenges evaluated, 58 (28.0%) contained at least one failing test case, with some challenges accumulating up to 12 failures. Notably, 13 challenges achieved a 0% success rate, meaning every single generated test case failed. Finally, through prompt refinement, the Test Execution Success Rate improved substantially, rising from 54% to 70.9%. In conclusion, this dissertation explores the effectiveness of leveraging problem specifications as test oracles, and assesses whether observed shortcomings, such as high failure rates in certain challenges, can be mitigated by incorporating them as additional context to enhance oracle generation.
dc.language.isoeng
dc.rightsopenAccess
dc.subjectEngenharia electrotécnica, electrónica e informática
dc.subjectElectrical engineering, Electronic engineering, Information engineering
dc.titleA Framework for Evaluating the Correctness of Programming Learning Platforms Feedback
dc.typeDissertação
dc.contributor.uportoFaculdade de Engenharia
dc.subject.fosCiências da engenharia e tecnologias::Engenharia electrotécnica, electrónica e informática
dc.subject.fosEngineering and technology::Electrical engineering, Electronic engineering, Information engineering
thesis.degree.disciplineMestrado em Engenharia de Software
thesis.degree.grantorFaculdade de Engenharia
thesis.degree.grantorUniversidade do Porto
thesis.degree.level1
Appears in Collections:FEUP - Dissertação

Files in This Item:
File Description SizeFormat 
787307.pdfA Framework for Evaluating the Correctness of Programming Learning Platforms Feedback1.3 MBAdobe PDFThumbnail
View/Open


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.