deepvariant - NVIDIA Docs

Run a GPU-accelerated DeepVariant algorithm.

What is DeepVariant?

DeepVariant is a deep learning based variant caller developed by Google for germline variant calling of high-throughput sequencing data. It works by taking aligned sequencing reads in BAM/CRAM format and utilizes a convolutional neural network (CNN) to classify the locus into true underlying genomic variation or sequencing error. DeepVariant can therefore call single nucleotide variants (SNVs) and insertions/deletions (InDels) from sequencing data at high accuracy in germline samples.

Why DeepVariant?

DeepVariant’s approach is able to detect variants that are often missed by traditional (for example Bayesian) variant callers, and is known to reduce false positives. It offers several advantages over similar tools, including its ability to detect a wide range of variants with high accuracy, its scalability for analyzing large datasets, and its open source availability. Additionally, its deep learning-based approach allows it to provide better support for different sequencing platforms, as it can be retrained to provide higher accuracy for specific protocols or research areas.

How should I use DeepVariant?

DeepVariant is designed for use as a germline variant caller that can apply different models trained for specific sample types (such as whole genome and whole exome samples) to yield higher accuracy results. DeepVariant can be deployed within NVIDIA’s Parabricks software suite, which is designed for accelerated secondary analysis in genomics, bringing industry standard tools and workflows from CPU to GPU, and delivering the same results at up to 60x faster runtimes. A 30x whole genome can be run through DeepVariant in as little as 8 minutes on an NVIDIA DGX station, compared to 5 hours on a CPU instance (m5.24xlarge, 96 x vCPU). DeepVariant in Parabricks is used in the same way as other command line tools that users are familiar with: It takes a BAM/CRAM and the reference genome as inputs and produces the variants (a VCF file) as outputs. Currently, DeepVariant is supported for V100 and newer GPUs out of the box.

Note

In version 3.8 the --run-partition option was added, which can lead to a significant speed increase. However, using the --run-partition, --proposed-variants, and --gvcf options at the same time will lead to a substantial slowdown. A warning will be issued and the --run-partition option will be ignored.

Available Operating Modes

Parabricks DeepVariant can run in one of three operating modes:

shortread
PacBio
ONT

See the --mode option below.

Quick Start

Copy
Copied!

            
            # This command assumes all the inputs are in INPUT_DIR and all the outputs go to OUTPUT_DIR.
docker run --rm --gpus all --volume INPUT_DIR:/workdir --volume OUTPUT_DIR:/outputdir \
    --workdir /workdir \
    nvcr.io/nvidia/clara/clara-parabricks:4.2.1-1 \
    pbrun deepvariant \
    --ref /workdir/${REFERENCE_FILE} \
    --in-bam /workdir/${INPUT_BAM} \
    --out-variants /outputdir/${OUTPUT_VCF}

Compatible Google DeepVariant Commands

The commands below are the Google counterpart of the Parabricks command above. The output from these commands will be identical to the output from the above command. See the Output Comparison page for comparing the results.

Copy
Copied!

            
            sudo docker run \
--volume <INPUT_DIR>:/input \
--volume <OUTPUT_DIR>:/output \
google/deepvariant:1.5.0 \
/opt/deepvariant/bin/run_deepvariant \
--model_type WGS \
--ref /input/${REFERENCE_FILE} \
--reads /input/${INPUT_BAM} \
--output_vcf /output/${OUTPUT_VCF} \
--num_shards $(nproc) \
--make_examples_extra_args "ws_use_window_selector_model=true"

Models for additional GPUs

Parabricks DeepVariant supports the following models:

Short-read WGS
Short-read WES
PacBio
ONT

DeepVariant models for T4, V100 and all other GPUs which are Ampere and above architecture ship with the software.

deepvariant Reference

Run DeepVariant to convert BAM/CRAM to VCF.

Input/Output file options

--ref REF
--in-bam IN_BAM
--interval-file INTERVAL_FILE
--out-variants OUT_VARIANTS
--pb-model-file PB_MODEL_FILE
--pb-model-dir PB_MODEL_DIR
--proposed-variants PROPOSED_VARIANTS

Tool Options:

--disable-use-window-selector-model
--gvcf
--norealign-reads
--sort-by-haplotypes
--keep-duplicates
--vsc-min-count-snps VSC_MIN_COUNT_SNPS
--vsc-min-count-indels VSC_MIN_COUNT_INDELS
--vsc-min-fraction-snps VSC_MIN_FRACTION_SNPS
--vsc-min-fraction-indels VSC_MIN_FRACTION_INDELS
--min-mapping-quality MIN_MAPPING_QUALITY
--min-base-quality MIN_BASE_QUALITY
--mode MODE
--alt-aligned-pileup ALT_ALIGNED_PILEUP
--variant-caller VARIANT_CALLER
--add-hp-channel
--parse-sam-aux-fields
--use-wes-model
--include-med-dp
--normalize-reads
--pileup-image-width PILEUP_IMAGE_WIDTH
--channel-insert-size
--no-channel-insert-size
--max-read-size-512
--prealign-helper-thread
--track-ref-reads
--phase-reads
--dbg-min-base-quality DBG_MIN_BASE_QUALITY
--ws-min-windows-distance WS_MIN_WINDOWS_DISTANCE
--channel-gc-content
--channel-hmer-deletion-quality
--channel-hmer-insertion-quality
--channel-non-hmer-insertion-quality
--skip-bq-channel
--aux-fields-to-keep AUX_FIELDS_TO_KEEP
--vsc-min-fraction-hmer-indels VSC_MIN_FRACTION_HMER_INDELS
--vsc-turn-on-non-hmer-ins-proxy-support
--consider-strand-bias
--p-error P_ERROR
--channel-ins-size
--max-ins-size MAX_INS_SIZE
--disable-group-variants
--filter-reads-too-long
-L INTERVAL, --interval INTERVAL

Performance Options:

--num-cpu-threads-per-stream NUM_CPU_THREADS_PER_STREAM
--num-streams-per-gpu NUM_STREAMS_PER_GPU
--run-partition
--gpu-num-per-partition GPU_NUM_PER_PARTITION
--max-reads-per-partition MAX_READS_PER_PARTITION
--partition-size PARTITION_SIZE

Common options:

--logfile LOGFILE
--tmp-dir TMP_DIR
--with-petagene-dir WITH_PETAGENE_DIR
--keep-tmp
--no-seccomp-override
--version

GPU options:

--num-gpus NUM_GPUS