Department

Department of Computer Science

Document Type

Article

Publication Date

3-31-2026

Abstract

Genotype imputation enables dense variant coverage for genome-wide association and riskprediction studies, yet conventional reference-panel methods remain limited by ancestry bias and reduced rare-variant accuracy. We present Genotype Bidirectional Encoder Representations from Transformers model (GenoBERT), a transformer-based, reference-free framework that tokenizes phased genotypes and uses self-attention mechanism to capture both short- and long-range linkage disequilibrium (LD) dependencies. Benchmarking on two independent datasets—the Louisiana Osteoporosis Study (LOS) and the 1000 Genomes Project (1KGP)—across ancestry groups and multiple genotypes missing levels (5–50%) shows that GenoBERT achieves the highest overall accuracy compared to the other four baselines (Beagle5.4, SCDA, BiU-Net, and STICI). At practical sparsity levels (≤ 25% missing), GenoBERT attains high overall imputation accuracies (r² ≈ 0.98) across datasets, and even maintains robust performance (r² > 0.90) at 50% missing level. Experiment results across different ancestries confirm consistent gains for both datasets, with resilience to small sample sizes and weak LD. The 128-SNP (single-nucleotide polymorphism) context window (≈ 100Kb) was validated through LD-decay analyses as sufficient to span local correlation structures. By eliminating reference-panel dependence while preserving high accuracy, GenoBERT provides a scalable, and robust solution for genotype imputation and a foundation for downstream genomic modeling.

Journal Title

arXiv

Digital Object Identifier (DOI)

https://doi.org/10.48550/arXiv.2604.00058

Share

COinS