ISEP - DM – Engenharia de Inteligência Artificial
URI permanente para esta coleção:
Navegar
Percorrer ISEP - DM – Engenharia de Inteligência Artificial por Domínios Científicos e Tecnológicos (FOS) "Engenharia e Tecnologia::Engenharia Eletrotécnica, Eletrónica e Informática"
A mostrar 1 - 5 de 5
Resultados por página
Opções de ordenação
- An explainable and privacy-preserving machine learning pipeline for early detection of endometriosis leveraging liquid biopsy and minimally-invasive CPublication . MANESSE, CIRO MIGUEL POÇAS FERREIRA; Martinho, Diogo Emanuel Pereira; Conceição, Luís Manuel SilvaEndometriosis a!ects approximately one in ten women of reproductive age, yet it is diagnosed, on average, seven to eight years after the onset of symptoms, and in some studies up to twelve. This diagnostic delay is associated with prolonged symptoms, uncertainty, repeated healthcare contacts and delayed therapeutic intervention, largely because a definitive diagnosis still depends on an invasive surgical procedure, namely laparoscopy. Recent progress in Machine Learning, together with the growing availability of clinical and molecular data, o!ers a realistic opportunity to shorten this interval. This dissertation investigates whether a non-invasive, explainable and privacy-preserving Machine Learning pipeline can support an earlier clinical suspicion of Endometriosis from micro-RNA (miRNA) measured in a blood sample, that is, a liquid biopsy. Three requirements are addressed: the test must be non-invasive; its decisions must be explainable, so that a clinician can examine and verify them; and the training procedure must protect sensitive patient data. To the best of the author’s knowledge, no previous work brings these three properties together for Endometriosis detection. Working exclusively with real, public data, a leakage-safe pipeline was built for the serummiRNA cohort GSE279435 (127 samples; 67 Endometriosis, 60 benign controls). A central observation in understanding the data is that the measurements are left-censored at the assay’s limit of detection, so that a missing value is itself informative. Additionally, every preprocessing step was fitted strictly inside each cross-validation fold to avoid optimistic bias. Five classifiers were compared under nested, repeated, stratified cross-validation. The two strongest (Random Forest and LightGBM) reached an area under the ROC curve of 0.77– 0.78 with balanced sensitivity and specificity, a performance comparable to that of the previously published eleven-miRNA model evaluated on the same cohort. Explainability was provided with SHAP, which makes each prediction inspectable both globally and at the level of an individual patient, allowing the model to be examined rather than trusted blindly. Five of the ten most influential miRNAs coincide with the published diagnostic panel, which grounds the model’s reasoning in established biology; the remaining five are plausible additional candidates, consistent with the di!erential-expression analysis, that would warrant follow-up in larger studies. On privacy, this is, to the best of the author’s knowledge, the first work to apply both Di!erential Privacy and Federated Learning to non-invasive miRNA-based Endometriosis detection. In a simulated federated setting, created by partitioning the public cohort into four virtual sites, a Federated Learning implementation built with NVIDIA FLARE trained a shared model without moving any raw data and recovered approximately 95% of the accuracy of a model trained on the pooled data, well above what any single site achieved in isolation. Di!erential Privacy, in contrast, reduced accuracy to little above chance at the privacy levels that would meaningfully protect patients, illustrating how costly strong, sample-level privacy guarantees become when the cohort is this small (n = 127). A cross-platform exploratory test on an independent plasma cohort indicated that the general signal — that circulating miRNA carries an Endometriosis-related signal — holds across biofluids, even though the specific serum signature does not transfer directly to plasma. The contribution of this work is therefore not a higher headline accuracy but a pipeline that, using real public data, reaches the performance of the previously published model while adding two properties that model did not have: explanations a clinician can recognise, and training that never centralises patient data. The proposed pipeline is presented not as a medical device, but as a reproducible and ethically framed foundation for the larger, multiinstitutional validation studies needed to support earlier clinical suspicion of Endometriosis.
- Previsão inteligente de falhas microbiológicas em endoscópios descontaminados com quantificação de incertezaPublication . PEREIRA, FILIPE DANIEL NOGUEIRA BARBOSA; Sousa, Laura Luciana Cavalcante de; Martinho, Diogo Manuel PereiraOs endoscópios flexíveis são dispositivos médicos reutilizáveis sujeitos a ciclos repetidos de utilização clínica e descontaminação por High-Level Disinfection (HLD). Apesar dos protocolos estabelecidos, a vigilância microbiológica revela taxas de contaminação residual com impacto direto na segurança do doente. Os métodos convencionais, baseados em culturas microbiológicas realizadas após o reprocessamento, só identificam falhas depois de o equipamento ter reentrado em circulação. Torna-se, por isso, relevante desenvolver abordagens preditivas que estimem o risco de falha antes de cada ciclo de utilização, a partir de variáveis operacionais, de instrumentação e de manutenção. A dissertação propõe e avalia uma pipeline preditiva integrada para a estimação do risco microbiológico em endoscópios descontaminados, enquadrada como prova de conceito com dados sintéticos. Na ausência de datasets clínicos públicos anotados com desfechos microbiológicos, foi construído um dataset sintético causalmente estruturado, com 30 000 registos e 20 dispositivos simulados, cuja cadeia causal modela a progressão desde a instrumentação e o resíduo orgânico até à qualidade de reprocessamento e ao risco microbiológico acumulado. A variável-alvo microbial_failure foi calibrada para uma prevalência de aproximadamente 15%, num cenário com classe positiva minoritária. Foram comparados três modelos de aprendizagem automática: Regressão Logística como referência linear, Random Forest como referência de ensemble e Extreme Gradient Boosting (XGBoost) como modelo principal, pela compatibilidade com TreeSHAP e pelo desempenho em dados tabulares desbalanceados. O XGBoost calibrado por Platt scaling obteve Area Under the Precision-Recall Curve (AUPRC) de 0,400, Area Under the Receiver Operating Characteristic (AUROC) de 0,753, Brier score de 0,110 e Expected Calibration Error (ECE) de 0,009. A análise de limiares identificou 0,12 como o limiar mínimo com recall ≥ 0,70, captando 70,5% das falhas com valor preditivo negativo de 92,9%, em linha com a assimetria de custos entre falsos negativos e falsos positivos. A explicabilidade foi obtida por TreeSHAP e os fatores dominantes foram a intensidade de instrumentação e o protocolo de reprocessamento. A incerteza das previsões foi quantificada por reamostragem bootstrap, evidenciando previsões geralmente estáveis, com variabilidade acrescida em subconjuntos específicos e uma relação exploratória com a posição face ao limiar operacional. Seis estudos de robustez adicionais confirmaram a coerência interna dos resultados. Foi ainda desenvolvido um protótipo funcional de demonstração académica, integrando predição, incerteza e explicabilidade numa interface de apoio à interpretação. A limitação central é a ausência de validação com dados clínicos reais. Os resultados obtidos são evidência exploratória da viabilidade da abordagem em contexto sintético e não podem ser generalizados sem validação prospetiva independente. O trabalho foi desenvolvido como base metodológica para investigação futura com dados hospitalares reais.
- Securing retrieval-augmented generationPublication . PEREIRA, PEDRO EMANUEL SOUSA; Pereira, Isabel Cecília Correia da Silva Praça Gomes; Maia, Eva Catarina GomesRetrieval-Augmented Generation (RAG) systems improve the factual grounding of language models by retrieving external documents before generating an answer. However, this dependence on external knowledge also creates a security risk, if the retrieval corpus is poisoned, the generated response may become incorrect while still appearing evidence-based. This thesis investigates the robustness of RAG pipelines against knowledge-base poisoning attacks. It rst analyzes how retrieval architecture, retrieval depth, database composition, chunking, dataset characteristics, and generator choice in uence poisoning vulnerability. The results show that robustness is a pipeline-level property, dense and graph-based retrieval are generally more resistant than lexical retrieval, but larger top-K values and poisoned multi-database settings increase exposure to adversarial content. The thesis then introduces Micro Collaborative Poisoning, a distributed attack in which several small, plausible poisoned documents collectively support the same false claim. Experiments show that this attack is less obvious at the document level than stronger concentrated poisoning, yet it can still achieve comparable downstream attack success. Overall, the ndings demonstrate that trustworthy RAG systems require defenses that combine robust retrieval, source integrity, cross-document analysis, and cautious generation.
- Topology-dependent privacy risks in decentralized federated learningPublication . GOUVEIA, JOSÉ INÁCIO ANTUNES DE; Pereira, Isabel Cecília Correia da Silva Praça Gomes; Amorim, Ivone de Fátima da CruzFederated Learning (FL) has established itself as a leading paradigm for collaborative machine learning, allowing participants to train models collectively without sharing their private data. Despite its privacy-preserving design, the periodic exchange of model updates leaves these systems vulnerable to information leakage. Notable threats include Membership Inference Attacks (MIA), which exploit model overfitting to determine if specific data samples were used during training, and Gradient Inversion Attacks (GIA), which attempt to reconstruct the exact training images from shared gradients. While existing literature has proposed various active defenses and investigated privacy risks within star and mesh network topologies, a systematic evaluation of how attack effectiveness evolves across training rounds over a diverse spectrum of network topologies remains a critical research gap. This dissertation presents a comprehensive empirical analysis of how different network topologies influence data privacy in centralized and decentralized FL systems over time. By isolating the network topology as the primary variable, we evaluate the vulnerability of six distinct topologies, star, tree, line, ring, full mesh, and partial mesh, against three MIA variants and a GIA. The experiments were conducted using the MNIST and CIFAR10 datasets under both Independent and Identically Distributed and Non-Independent and Identically Distributed (Non-IID) data distributions to capture the temporal evolution of these attacks across the training rounds. Our findings show that network topologies fundamentally dictate the severity and localization of privacy leakage. For instance, intermediate aggregation in the tree topology acts as a native privacy shield, effectively hiding memorized features from a central root node, though it inadvertently shifts the primary risk to intermediate edge nodes. Furthermore, the analysis reveals that extreme data heterogeneity (Non-IID) significantly aggravates vulnerabilities across all topologies, heavily increasing MIA and GIA success rates on complex visual tasks. Moreover, the results establish that restricting an adversary’s awareness of the broader network topology severely impedes their ability to accurately execute GIA, as successful data reconstruction depends heavily on precise knowledge of the aggregation phase in federated systems. Ultimately, this research highlights that network topology can be strategically leveraged as passive defense mechanisms.
- Transfer learning applied to government auditing: A focused approach on financial statements in Maranhão, BrazilPublication . Coelho, Heloisa Guimarães; Marreiros, Maria Goreti CarvalhoSince Brazil’s return to democracy, dozens of laws, decrees and normative instructions have been drafted with the purpose of regulating and improving the mechanisms for controlling and monitoring municipal public resources. These regulations are specifically aimed at the process of accountability by elected officials, who currently rely on the help of accountants responsible for preparing and submitting financial statements to the Courts of Auditors. However, according to data from the TCU (Federal Court of Accounts), in 2023, Maranhão was the Brazilian State with the highest number of rejected accounts. There are several reasons that can lead to these processes being challenged, including incorrect application of resources, flaws in documentation, human errors, among others. In practice, the routine of accountants includes repetitive and mechanical activities that requires considerable time to prepare and review documents, hence often leading to errors in classification and issuing of documentation. In this context, this dissertation investigates the use of Transfer Learning (TL) to improve automation and accuracy in the classification of financial commitment notes, an initial document in the public expenditure cycle, with a specific focus on the context of the state of Maranhão. To this end, BERTimbau, a pre-trained language model for Brazilian Portuguese, was fine-tuned to assist government accountants in reducing classification errors and ensuring compliance with local and national financial regulations. The CRISP-DM methodology, widely used in data science, was adopted to structure the development of the project. The dataset used, consisting of several classifications of commitment notes for the year 2023, was thoroughly analyzed and pre-processed. For the fine-tuning process of the model, two samples with a similar number of data were selected, varying only the number of possible classifications, due to the high degree of imbalance between the classes. Even in a multiclass context with datasets with a reduced number of classes, the results obtained indicate that the BERTimbau model presents strong performance in the classification task, achieving 98% accuracy with an error rate of 0.10 in the test set, highlighting the effectiveness of BERTimbau in public financial auditing applications. These results highlight the effectiveness of BERTimbau for public financial auditing applications. It is therefore concluded that TL models have great potential to optimize and improve financial auditing processes, with positive implications for wider adoption in Brazil.
