Single nucleotide polymorphisms (SNPs) are the most common type of genetic variation among individuals. SNPs are widely used to estimate genomic divergence, population structure, and natural selection, as well as to identify associations between genomic variants and phenotypic traits. However, the identification of SNPs in non-model species or from field-derived samples may still be challenging and would benefit from standardized protocols. The identification of SNPs from raw sequencing data, such as those obtained from a standard Illumina library, involves many processing steps and the use of diverse sets of tools, where the choice of different parameters significantly affects the results. Here, we present a pipeline for a genome-wide identification of SNPs from raw Illumina reads. This pipeline is composed of three major steps that encompass the GATK Best Practices and include quality control of the raw reads, mapping of short reads to a reference genome, recalibration of the alignment, variant calling, quality control for variant calling, and filtering of candidate SNPs. The steps of this pipeline are accompanied by publicly available scripts and datasets used for the identification of SNPs in the genome of the mosquito Aedes aegypti, an invasive species and a major arboviral vector worldwide.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Identification of Single Nucleotide Polymorphism from Insect Genomic Data

  • Alejandro Nabor Lozada-Chávez,
  • Mariangela Bonizzoni

摘要

Single nucleotide polymorphisms (SNPs) are the most common type of genetic variation among individuals. SNPs are widely used to estimate genomic divergence, population structure, and natural selection, as well as to identify associations between genomic variants and phenotypic traits. However, the identification of SNPs in non-model species or from field-derived samples may still be challenging and would benefit from standardized protocols. The identification of SNPs from raw sequencing data, such as those obtained from a standard Illumina library, involves many processing steps and the use of diverse sets of tools, where the choice of different parameters significantly affects the results. Here, we present a pipeline for a genome-wide identification of SNPs from raw Illumina reads. This pipeline is composed of three major steps that encompass the GATK Best Practices and include quality control of the raw reads, mapping of short reads to a reference genome, recalibration of the alignment, variant calling, quality control for variant calling, and filtering of candidate SNPs. The steps of this pipeline are accompanied by publicly available scripts and datasets used for the identification of SNPs in the genome of the mosquito Aedes aegypti, an invasive species and a major arboviral vector worldwide.