Title

1  Introduction

1.1 About SheepVar

SheepVar overview

SheepVar is a comprehensive, variant-centered resource for sheep genomic variation and functional annotation. It integrates high-quality SNPs and InDels derived from large-scale genomic datasets. The current release includes whole-genome resequencing data from 5,603 modern sheep samples, from which approximately 62.71 million SNPs and 3.72 million small InDels were retained after quality control. SheepVar also incorporates genotype data from 7,013 individuals generated using the Illumina OvineSNP50 BeadChip, covering approximately 41K loci, together with 123 sheep epigenomic datasets, including ChIP-seq and ATAC-seq profiles, and paleogenomic data from 100 ancient sheep specimens. Collectively, these datasets provide broad representation of domestic sheep populations and wild relatives worldwide.

SheepVar provides multidimensional annotations for individual variants, including genomic context, evolutionary conservation, inferred ancestral states, population allele frequencies, molecular quantitative trait loci (QTL) associations, epigenomic features, selection signatures, and haplotype blocks.

The platform supports flexible variant queries, integrated visualization of multi-omics data, exploration of population-specific selection signals and haplotype structures, and online genotype imputation. Users can search for variants by genomic coordinates, rsIDs, gene symbols, or SNP array marker IDs and access detailed annotations through the Variant Details page. SheepVar also provides tools for genome browsing, sequence alignment, coordinate conversion, and data download.

Overall, SheepVar provides a curated and standardized platform for browsing, interpreting, and prioritizing sheep genomic variants, supporting research in population genomics, functional genomics, evolutionary biology, and molecular breeding.

1.2 Variant Processing

Raw SRA files were converted to FASTQ format using SRA Toolkit v3.2.0, and adapter sequences and low-quality bases were removed using fastp v0.12.4. Clean reads were aligned to the sheep reference genome using BWA-MEM. The resulting alignments were sorted, BAM files from the same sample were merged, and duplicate reads were removed before variant calling.

SNPs and small InDels were identified using a uniform GATK v4.1.8.1 pipeline. Variants were first called for each sample using HaplotypeCaller in GVCF mode, followed by joint genotyping across samples using GenotypeGVCFs and variant extraction using SelectVariants. Variants with QUAL < 30 were removed, and additional hard-filtering criteria were applied separately to SNPs and InDels. SNPs were filtered using the following expression: QD < 2.0 || FS > 60.0 || MQ < 40.0 || MQRankSum < -12.5 || ReadPosRankSum < -8.0 || SOR > 3.0. InDels were filtered using: QD < 2.0 || FS > 200.0 || SOR > 10.0 || MQRankSum < -12.5 || ReadPosRankSum < -8.0. Only small InDels of 1–30 bp were retained.

For further quality control, only biallelic variants were retained. PLINK v1.90 was used to remove variants and samples with genotyping rates below 90%. After quality control, 5,603 samples were retained, and missing genotypes were imputed using Beagle v5.4. Finally, variants with a minor allele frequency (MAF) ≥ 0.1% were retained, resulting in a high-confidence dataset comprising approximately 62.71 million SNPs and 3.72 million small InDels.

1.3 Contact Us

Yu Jiang

Yu Jiang

Northwest A&F University, Yangling, Shaanxi, China

Email: yu.jiang@nwafu.edu.cn

Ran Li

Ran Li

Northwest A&F University, Yangling, Shaanxi, China

Email: ran.li@nwsuaf.edu.cn

QuanZhong Liu

QuanZhong Liu

Northwest A&F University, Yangling, Shaanxi, China

Email: liuqzhong@nwsuaf.edu.cn

2  Tutorial

2.1 Overview of the Home page

SheepVar home page

1. Main navigation of the database.

2. A simple introduction about the database.

3. Fast retrieval of the database.

4. The statistics of the data in the database.

5. Updates & Related Resources.

2.2 Variation

The General Search module allows users to search genome-wide SNPs and small InDels from WGS and SNP array datasets. Variants can be queried by genomic coordinates, rsIDs, gene symbols, or SNP array marker IDs. Advanced Search additionally supports filtering by criteria such as MAF and predicted consequence type.

General Search

Search results are displayed on the SNPs Found page with key information including genomic position, dbSNP ID, associated gene, MAF, and predicted functional consequence. Results can be further filtered using the panel at the top of the page. Click the arrow next to a variant to open the Variant Details page and view its full annotation.

SNPs Found

2.2.2 Variant Details

The Variant Details page provides integrated annotations for each variant, including basic information, quality metrics, population allele frequencies, evolutionary annotations, regulatory evidence, molecular QTLs, and trait-related QTLs.


Variant Information

This section summarizes the variant identifier (Chr:Pos_Ref_Alt), dbSNP rsID, reference and alternate alleles, minor allele, MAF, ancestral allele, and SnpEff annotations, including gene, transcript, consequence type, and predicted impact. An embedded genome viewer displays the local genomic context and nearby genes.

Variant Information
Quality Metrics

The Quality Metrics section reports variant-level quality information to help users assess the technical reliability of each variant call. The displayed metrics include read depth and standard variant-calling quality metrics, such as QUAL, QD, FS, SOR, MQ, MQRankSum and ReadPosRankSum.

These metrics provide information on read support, call confidence, strand bias, mapping quality and positional bias. Users can use this information to evaluate whether a variant is well supported by sequencing data or may require additional caution in downstream analyses.

Quality Metrics
Frequency Information

This section displays allele frequencies across modern sheep populations and, when available, genotype information from ancient sheep samples. Modern population frequencies are shown for populations with sufficient sample sizes, while wild sheep are retained as comparative references. Ancient records are accompanied by temporal and geographic information, enabling comparison of allele distributions across populations, regions, and historical periods.

Frequency Information
Conservation Scores

This section provides evolutionary conservation annotations based on multi-species alignments, including phyloP and phastCons scores. Positive phyloP scores indicate conservation, whereas negative scores indicate accelerated evolution. phastCons scores range from 0 to 1 and reflect the probability that a site lies within a conserved element. Together, these metrics help assess the evolutionary constraint of a variant site.

Conservation Scores
ATAC-seq

This section shows whether the selected variant overlaps ATAC-seq peaks representing open chromatin regions. When an overlap is detected, the genomic coordinates, tissue or cell type, and related peak information are displayed.

ATAC-seq
ChIP-seq

This section shows whether the selected variant overlaps ChIP-seq peaks. Information on the tissue or experimental context and regulatory marks, including histone modifications such as H3K4me1, H3K27ac, and H3K4me3, is provided when available.

ChIP-seq
Molecular QTL

This section provides cis-molQTL associations from SheepGTEx, including eQTLs, sQTLs, eeQTLs, stQTLs, isoQTLs, 3’ aQTLs, and enQTLs across up to 51 sheep tissues or cell types. Associated tissues, genes, molecular trait types, effect directions, and significance are displayed. Users can select tissue–gene pairs to view interactive violin plots showing genotype-dependent differences in molecular phenotypes.

Molecular QTL
QTL Information

This section shows trait-associated QTLs from Animal QTLdb that overlap the selected variant. Information on associated traits and QTL regions is provided, covering categories such as growth, carcass, wool, reproduction, disease resistance, and other production-related traits.

QTL Information

2.3 Imputation

SheepVar provides SNP array-to-sequence genotype imputation using Beagle v5.4. Five reference panels are currently available, including two global panels and three ancestry-specific regional panels. Detailed information on panel size, population composition, geographic origin, and sequencing depth is available under Panel Composition in the Sample Table.

Step 1: Select a Reference Panel

Global4678 is generally recommended for populations with unknown, mixed, or admixed genetic backgrounds and for lower-density 50K genotype data. For populations with well-defined genetic backgrounds, ancestry-specific regional panels may provide comparable accuracy with lower computational requirements, particularly for dense 600K genotype data. Global3125 is retained mainly for compatibility with previous SheepGTEx datasets.

Reference panel selection

Step 2: Upload Genotype Data

The online service accepts genotype data in VCF format and PLINK-compatible PED or BED formats. VCF format is recommended for genotype imputation.

  • VCF file (file size ≤ 50 MB)
  • PED file (file size ≤ 300 MB)
  • BED file (file size ≤ 50 MB)
  • Number of individuals (≤ 1,000)

Step 3: Submit and Monitor the Job

Click Submit to start genotype imputation. Each user may submit up to 10 concurrent jobs (≤ 10 jobs per user). Additional jobs are queued when this limit is reached or when the server is operating at full capacity.

Users can monitor job status and download completed results through the Job List page. An email notification is sent when the imputation job is completed.

Imputation job list

2.4 Selective Sweep

The Selective sweep module allows users to select populations and genomic regions of interest to examine signals of selection. Results are provided separately for WGS and SNP chip datasets. For WGS data, SheepVar displays Pi, Tajima’s D, CLR, iHS and FST in non-overlapping 30kb windows. For SNP chip data, SheepVar displays Pi, Tajima’s D, iHS and FST in non-overlapping 150kb windows because of the lower marker density of chip-based data; CLR analysis is not provided for SNP chip datasets.

Selective sweep

2.5 Haplotype Block

This module allows users to explore population-specific haplotype blocks by selecting a population and genomic region. To ensure efficient visualization, the queried region is limited to 250 kb.

Haplotype blocks

Identified haplotype blocks are displayed with their genomic locations and SNP composition. Users can click a Block ID to view the corresponding haplotype combinations and their frequencies in the selected population.

Haplotype combinations

Local LD patterns can also be visualized by clicking the “Generate heatmap” option below the results.

LD Heatmap

2.6 Genome Browser

The Genome Browser allows users to visualize genomic features and annotation tracks based on the sheep reference genome assembly ARS-UI_Ramb_v2.0. Currently, 558 tracks are available, including SNPs, small InDels, selection signals, genotype patterns, QTLs, conservation scores, conserved elements and epigenomic regulatory tracks.

Users can search the browser using a gene symbol, transcript name or genomic region to explore genomic information in a genome-wide context. Search results and candidate regions from SheepVar can be directly linked to the Genome Browser, allowing users to examine candidate genes or loci together with multiple annotation tracks. High-quality images can be exported in PDF or PostScript format using the PDF/PS option under the View menu.

Genome Browser

2.7 Other Tools

LiftOver

SheepVar also provides LiftOver for genome coordinate conversion. This tool converts genomic coordinates between different sheep genome assembly versions and can also be used to retrieve putative orthologous regions in other species. Six LiftOver chain files are currently available in SheepVar, allowing users to perform coordinate conversion and cross-assembly comparison conveniently.

BLAT and BLAST

SheepVar provides two sequence alignment tools, webBLAT and ViroBLAST, to support sequence-based searches. webBLAT enables rapid alignment of DNA or mRNA sequences to the sheep reference genome, and matched regions can be directly displayed in the Genome Browser. ViroBLAST is used to identify local sequence similarities, helping users explore potential functional or evolutionary relationships among sequences.