Rare diseases (RDs) affect approximately 1 in 20 individuals worldwide, with 80% of these conditions having a genetic origin. The recent advancements in high-throughput sequencing technologies allow the sequencing of the entire genome at lower costs and the identification of millions of genomic variants per patient, including SNVs, INDELs, CNVs, and complex structural variants. Identifying the causative variant among millions is the main challenge, and about 50% of RD patients remain without a definitive diagnosis. Even when the causative variant is eventually identified, it may take many years and involve misdiagnoses. For this reason it is usually called diagnostic odyssey. This PhD Executive project aimed to develop an end-to-end AI-based framework to support geneticists throughout the variant interpretation workflow, by integrating, within both monogenic (Mendelian) and digenic inheritance models, identified genomic variants with the clinical context of the patient, including phenotypic information and family history. This workflow also supports the findings with insights extracted from the published literature. The framework includes three main layers: (i) a Machine-Learning (ML)–based prioritization approach under a Mendelian monogenic hypothesis; (ii) a second ML-based approach that, in addition to phenotypic description and family history, exploits also gene-gene interaction features to prioritize candidate digenic combinations; and (iii) a domain-specific Large Language Model (LLM)-based workflow that extracts, summarizes, and integrates the most relevant literature findings into concise clinical reports that could support the identified diagnosis with controllable and trustworthy sources. The proposed approach aims to shorten the diagnostic gap by including multiple inheritance paradigms in the prioritization framework, to accelerate the interpretation of results by using explainable features that overcome the interpretability limitations of black-box models and to incorporate insights from the scientific literature. The final goal is to reduce the burden on geneticists in identifying the correct genetic diagnosis, providing a short list of candidate variants or variant combinations to be reviewed. These three approaches were implemented, tested, and validated within this thesis work, with feature development, documentation, and performance assessment described in detail in this manuscript. They are currently available through a commercial software solution for genomic data analysis and through an open platform accessible to the scientific community. Together, these approaches aim to move toward an increasingly comprehensive and explainable genomic interpretation workflow, guiding in a smooth and transparent way the adoption of AI-assisted tools in clinical practice.
Le malattie rare (RDs) colpiscono circa 1 individuo su 20 a livello globale, e l’80% di queste patologie ha un’origine genetica. I recenti progressi nelle tecnologie di sequenziamento del DNA consentono di sequenziare l’intero genoma di un individuo a costi sempre più contenuti, permettendo l'identificazione di milioni di varianti genomiche per paziente, incluse SNV, INDEL, CNV ed eventi strutturali più complessi. Identificare, tra milioni di varianti genetiche osservate nel paziente, quella responsabile della patologia rappresenta ancora oggi una delle principali sfide della genetica clinica: circa il 50% dei pazienti non riceve infatti una diagnosi definitiva. Anche quando la variante causativa viene identificata, il processo può richiedere molti anni e includere diversi episodi di diagnosi errate e successive rivalutazioni cliniche. Per questo motivo, viene comunemente definito odissea diagnostica. Questo progetto di dottorato industriale ha avuto come obiettivo lo sviluppo di un framework end-to-end basato su modelli di Intelligenza Artificiale per supportare i genetisti nell'interpretazione delle varianti genetiche, integrando modelli di ereditarietà sia monogenica (Mendeliana), sia digenica con le informazioni cliniche del paziente. In particolare, il framework utilizza la descrizione fenotipica del paziente e la sua storia familiare e l'ipotesi diagnostica viene supportata da evidenze estratte dalla letteratura scientifica pubblicata. La soluzione proposta si articola secondo tre livelli principali: (i) un approccio di prioritizzazione basato su Machine Learning (ML) secondo un modello Mendeliano monogenico; (ii) un secondo approccio di ML che, oltre alla descrizione fenotipica e alla storia familiare, incorpora evidenze di interazione genica per identificare e prioritizzare combinazioni candidate digeniche; e (iii) un workflow basato su un Large Language Model (LLM) specializzato nel dominio genomico che estrae, riassume e integra le evidenze più rilevanti dalla letteratura in report clinici brevi ma informativi, in grado di supportare la diagnosi identificata mediante informazioni estratte da fonti verificabili e affidabili. L'approccio proposto mira a ridurre il divario diagnostico includendo molteplici paradigmi di ereditarietà all'interno di un unico framework di prioritizzazione; l'interpretazione dei risultati è fortemente supportata dall'utilizzo della cosiddetta Explainable AI che permette di superare i limiti di interpretabilità dei modelli black-box e dall'integrazione delle evidenze provenienti dalla letteratura scientifica, con l'obiettivo finale di ridurre il carico di lavoro dei genetisti nell’identificazione della corretta diagnosi genetica, fornendo un numero limitato di varianti o combinazioni di varianti da sottoporre a revisione manuale. Questi tre approcci sono stati implementati, testati e validati nell’ambito del presente lavoro di tesi. Lo sviluppo delle funzionalità, la documentazione e la valutazione delle prestazioni sono descritte in dettaglio nel presente manoscritto. Attualmente, tali strumenti sono disponibili sia attraverso un software commerciale per l’analisi dei dati genomici sia tramite una piattaforma open access a supporto della comunità scientifica. Nel loro insieme, i contributi presentati in questa tesi mirano a sviluppare un workflow di interpretazione genomica end-to-end sempre più completo e facilmente interpretabile, favorendo un'adozione trasparente ed efficace delle tecnologie di Intelligenza Artificiale nella routine della pratica clinica.
Framework di intelligenza artificiale per la diagnosi genetica delle malattie rare: dalla prioritizzazione delle varianti alla refertazione clinica
DE PAOLI, FEDERICA
2026-07-22
Abstract
Rare diseases (RDs) affect approximately 1 in 20 individuals worldwide, with 80% of these conditions having a genetic origin. The recent advancements in high-throughput sequencing technologies allow the sequencing of the entire genome at lower costs and the identification of millions of genomic variants per patient, including SNVs, INDELs, CNVs, and complex structural variants. Identifying the causative variant among millions is the main challenge, and about 50% of RD patients remain without a definitive diagnosis. Even when the causative variant is eventually identified, it may take many years and involve misdiagnoses. For this reason it is usually called diagnostic odyssey. This PhD Executive project aimed to develop an end-to-end AI-based framework to support geneticists throughout the variant interpretation workflow, by integrating, within both monogenic (Mendelian) and digenic inheritance models, identified genomic variants with the clinical context of the patient, including phenotypic information and family history. This workflow also supports the findings with insights extracted from the published literature. The framework includes three main layers: (i) a Machine-Learning (ML)–based prioritization approach under a Mendelian monogenic hypothesis; (ii) a second ML-based approach that, in addition to phenotypic description and family history, exploits also gene-gene interaction features to prioritize candidate digenic combinations; and (iii) a domain-specific Large Language Model (LLM)-based workflow that extracts, summarizes, and integrates the most relevant literature findings into concise clinical reports that could support the identified diagnosis with controllable and trustworthy sources. The proposed approach aims to shorten the diagnostic gap by including multiple inheritance paradigms in the prioritization framework, to accelerate the interpretation of results by using explainable features that overcome the interpretability limitations of black-box models and to incorporate insights from the scientific literature. The final goal is to reduce the burden on geneticists in identifying the correct genetic diagnosis, providing a short list of candidate variants or variant combinations to be reviewed. These three approaches were implemented, tested, and validated within this thesis work, with feature development, documentation, and performance assessment described in detail in this manuscript. They are currently available through a commercial software solution for genomic data analysis and through an open platform accessible to the scientific community. Together, these approaches aim to move toward an increasingly comprehensive and explainable genomic interpretation workflow, guiding in a smooth and transparent way the adoption of AI-assisted tools in clinical practice.| File | Dimensione | Formato | |
|---|---|---|---|
|
PhDThesis_DePaoli_DEF_PDF_A.pdf
embargo fino al 31/01/2028
Descrizione: AN END-TO-END AI FRAMEWORK FOR RARE DISEASE GENETIC DIAGNOSIS: FROM HYPOTHESIS-DRIVEN VARIANT PRIORITIZATION TO CLINICAL REPORTING
Tipologia:
Tesi di dottorato
Dimensione
10.77 MB
Formato
Adobe PDF
|
10.77 MB | Adobe PDF | Visualizza/Apri Richiedi una copia |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


