The skyrocketing development in the field of personal genomics has produced enormous amounts of genetic information yet the issue of how to process this raw material into identifiable and interpretable outcomes is not truly figured out yet. In this paper, the authors present the Genomic Insight-Engine, a rule-based explainable computational framework able to process individual genomic data to make predictions on latent human potentials. As a measure of biological reliability and transparency, the system incorporates scientifically-validated SNP-trait associations in repositories like SNPedia and NHGRI-EBI GWAS Catalog. In contrast to black-box machine learning models, the framework uses deterministic rule logic that explicitly connects all inferred traits with their genetic evidences so that the users can easily map each prediction to its single-nucleotide polymorphism. It is written in Python and Streamlit, a lightweight and privacy-respecting environment to run the engine locally without having to use cloud-based computation. The validity of the system is proven by experimentation, which has shown that the system is capable of working with large genomic datasets and generating interpretable profiles of traits including cognitive, physical and behavioral aspects. The trait generated is not only more informative on the part of the users of genetic predispositions but also creates precedence on ethically and self-aware genomic analysis. The Genomic Insight-Engine combines transparency, efficiency, and privacy to bring about the next era in elucidating bioinformatics concerning personal matters, and lead the way to individualized, evidence-based understanding of human potential in the age of accessible genomics.
Gangula Nithya Sree, Muntha Raju, Yerram Sneha, & Sirisha Veluri. (2026). Genomic Insight-Engine: An Explainable Rule-Based Framework for Profiling Human Potentials from SNP Data. In International Journal of Recent Trends in Technology and Engineering (IJRTTE) (Vol. 5, Issue 3, pp. 1–14). NTL Publisher. https://doi.org/10.5281/zenodo.22094639
Section
Articles
References
Watson DS. Interpretable machine learning for genomics. Hum Genet. 2022;141(9):1499–1513. https://doi.org/10.1007/s00439-021-02387-9
Chen H, Lundberg SM, Lee SI. Explaining a series of models by propagating Shapley values. Nat Commun. 2022;13(1):4512. https://doi.org/10.1038/s41467-022-31384-3
Novakovsky G, Fornes O, Saraswat M, Mostafavi S, Wasserman WW. ExplaiNN: Interpretable and transparent neural networks for genomics. Genome Biol. 2023;24(1):154. https://doi.org/10.1186/s13059-023-02985-y
van Hilten A, Katz S, Saccenti E, Niessen WJ, Roshchupkin GV. Designing interpretable deep learning applications for functional genomics: A quantitative analysis. Brief Bioinform. 2024;25(5):bbae449. https://doi.org/10.1093/bib/bbae449
Wagle MM, Long S, Chen C, Liu C, Yang P. Interpretable deep learning in single-cell omics. Bioinformatics. 2024;40(6):btae374. https://doi.org/10.1093/bioinformatics/btae374
Ahlquist KD, Sugden LA, Ramachandran S. Enabling interpretable machine learning for biological data with reliability scores. PLoS Comput Biol. 2023;19(5):e1011175. https://doi.org/10.1371/journal.pcbi.1011175
Stock M, Van Criekinge W, Boeckaerts D, Taelman S, Van Haeverbeke M, Dewulf P, et al. Hyperdimensional computing: A fast, robust, and interpretable paradigm for biological data. PLoS Comput Biol. 2024;20(9):e1012426. https://doi.org/10.1371/journal.pcbi.1012426
Li Z, Zhang Y, Peng B, Qin S, Zhang Q, Chen Y, et al. A novel interpretable deep learning-based computational framework designed synthetic enhancers with broad cross-species activity. Nucleic Acids Res. 2024;52(21):13447–13468. https://doi.org/10.1093/nar/gkae912
Tang X, Zhang J, He Y, Zhang X, Lin Z, Partarrieu S, et al. Explainable multi-task learning for multi-modality biological data analysis. Nat Commun. 2023;14(1):2546. https://doi.org/10.1038/s41467-023-37477-x
Xiang R, Kelemen M, Xu Y, Harris LW, Parkinson H, Inouye M, et al. Recent advances in polygenic scores: Translation, equitability, methods and FAIR tools. Genome Med. 2024;16(1):33. https://doi.org/10.1186/s13073-024-01304-9
Jermy B, Läll K, Wolford BN, Wang Y, Zguro K, Cheng Y, et al. A unified framework for estimating country-specific cumulative incidence for 18 diseases stratified by polygenic risk. Nat Commun. 2024;15(1):5007. https://doi.org/10.1038/s41467-024-48938-2
Ohta R, Tanigawa Y, Suzuki Y, Kellis M, Morishita S. A polygenic score method boosted by non-additive models. Nat Commun. 2024;15(1):4433. https://doi.org/10.1038/s41467-024-48654-x
Sollis E, Mosaku A, Abid A, Buniello A, Cerezo M, Gil L, et al. The NHGRI-EBI GWAS Catalog: Knowledgebase and deposition resource. Nucleic Acids Res. 2023;51(D1):D977–D985. https://doi.org/10.1093/nar/gkac1010
Cerezo M, Sollis E, Ji Y, Lewis E, Abid A, Bircan KO, et al. The NHGRI-EBI GWAS Catalog: Standards for reusability, sustainability and diversity. bioRxiv. 2024. https://doi.org/10.1101/2024.10.23.619767
Liu X, Tian D, Li C, Tang B, Wang Z, Zhang R, et al. GWAS Atlas: An updated knowledgebase integrating more curated associations in plants and animals. Nucleic Acids Res. 2023;51(D1):D969–D976. https://doi.org/10.1093/nar/gkac924
Zhao Y, Shao J, Asmann YW. Assessment and optimization of explainable machine learning models applied to transcriptomic data. Genomics Proteomics Bioinformatics. 2022;20(5):899–911. https://doi.org/10.1016/j.gpb.2022.07.003
Toussaint PA, Leiser F, Thiebes S, Schlesner M, Brors B, Sunyaev A. Explainable artificial intelligence for omics data: A systematic mapping study. Brief Bioinform. 2023;25(1):bbad453. https://doi.org/10.1093/bib/bbad453
Chen V, Yang M, Cui W, Kim JS, Talwalkar A, Ma J. Applying interpretable machine learning in computational biology: Pitfalls, recommendations and opportunities. Nat Methods. 2024;21(8):1454–1461. https://doi.org/10.1038/s41592-024-02359-7
Kelly CM, McLaughlin RL. Comparison of machine learning methods for genomic prediction of selected Arabidopsis thaliana traits. PLoS One. 2024;19(8):e0308962. https://doi.org/10.1371/journal.pone.0308962
Cao Y, Zhao X, Tang S, Jiang Q, Li S, Li S, et al. scButterfly: A versatile single-cell cross-modality translation method via dual-aligned variational autoencoders. Nat Commun. 2024;15(1):2973. https://doi.org/10.1038/s41467-024-47418-x
Gordillo-Marañón M, Schmidt AF, Warwick A, Tomlinson C, Ytsma C, Engmann J, et al. Disease coverage of human genome-wide association studies and pharmaceutical research and development. Commun Med. 2024;4(1):195. https://doi.org/10.1038/s43856-024-00625-5
Sidak D, Schwarzerová J, Weckwerth W, Waldherr S. Interpretable machine learning methods for predictions in systems biology from omics data. Front Mol Biosci. 2022;9:926623. https://doi.org/10.3389/fmolb.2022.926623
Daoud A, Ben-Hur A. The role of chromatin state in intron retention: A case study in leveraging large-scale deep learning models. PLoS Comput Biol. 2025;21(1):e1012755. https://doi.org/10.1371/journal.pcbi.1012755