This study presents an explorative population genetic analysis with focus on Indian populations using Principal Component Analysis (PCA) and F-statistic with the goal of investigating genetic structure and admixture patterns. Indian populations have been chosen for this study because they show high genetic diversity due to numerous factors including ancient migration, endogamy and linguistic variations. The analysis was made on genome-wide single nucleotide polymorphism (SNP) datasets from multiple Indian population groups together with individuals from Europe, Middle East and East Asia in order to properly assess historical gene flow between populations. After individual filtering, SNP filtering and linkage disequilibrium pruning, PCA was applied to visualize patterns of genetic clustering and population structure, while f3-statistics were employed to detect admixture signals. The results show gradients of genetic variation consistent with known admixture between Ancestral North Indian (ANI) and Ancestral South Indian (ASI) components, and it has also shown admixture signals that are statistically significant in several populations, supporting historical interactions between caste and regional groups. Finally this study highlights the relevance of population genetics in medical research, in particular in identifying genetic risk factors that are population-specific and understanding the development of complex diseases in Indian populations.

Exploratory Population Genetic Analysis of Indian Populations Using PCA and F-statistics

CARRARO, BEATRICE
2025/2026

Abstract

This study presents an explorative population genetic analysis with focus on Indian populations using Principal Component Analysis (PCA) and F-statistic with the goal of investigating genetic structure and admixture patterns. Indian populations have been chosen for this study because they show high genetic diversity due to numerous factors including ancient migration, endogamy and linguistic variations. The analysis was made on genome-wide single nucleotide polymorphism (SNP) datasets from multiple Indian population groups together with individuals from Europe, Middle East and East Asia in order to properly assess historical gene flow between populations. After individual filtering, SNP filtering and linkage disequilibrium pruning, PCA was applied to visualize patterns of genetic clustering and population structure, while f3-statistics were employed to detect admixture signals. The results show gradients of genetic variation consistent with known admixture between Ancestral North Indian (ANI) and Ancestral South Indian (ASI) components, and it has also shown admixture signals that are statistically significant in several populations, supporting historical interactions between caste and regional groups. Finally this study highlights the relevance of population genetics in medical research, in particular in identifying genetic risk factors that are population-specific and understanding the development of complex diseases in Indian populations.
2025
Exploratory Population Genetic Analysis of Indian Populations Using PCA and F-statistics
population genetics
genetic admixture
indian populations
PCA
F3-statistic
File in questo prodotto:
File Dimensione Formato  
Carraro_Beatrice.pdf

accesso aperto

Dimensione 1.12 MB
Formato Adobe PDF
1.12 MB Adobe PDF Visualizza/Apri

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/111453