NVIDIA Docs Hub Homepage NVIDIA Clara deepvariant_germline

deepvariant_germline

Given one or more pairs of FASTQ files, you can run the germline variant tool to generate BAM, variants, duplicate metrics and recal.

The deepvariant germline tool includes alignment, sorting, and marking as well as the DeepVariant variant caller.

The inputs are BWA-indexed reference files and pair-ended FASTQ files. The outputs of this tool are the following:

Aligned, co-ordinate sorted, duplicated marked BAM
Variants in vcf/g.vcf/g.vcf.gz format

Quick Start

The following command runs the DeepVariant tool.

Copy
Copied!

            
            # This command assumes all the inputs are in INPUT_DIR and all the outputs go to OUTPUT_DIR.
docker run --rm --gpus all --volume INPUT_DIR:/workdir --volume OUTPUT_DIR:/outputdir \
    --workdir /workdir \
    nvcr.io/nvidia/clara/clara-parabricks:4.2.1-1 \
    pbrun deepvariant_germline \
    --ref /workdir/${REFERENCE_FILE} \
    --in-fq /workdir/${INPUT_FASTQ_1} /workdir/${INPUT_FASTQ_2} \
    --out-variants /outputdir/${OUTPUT_VCF_FILE}

Compatible Google DeepVariant Commands

The commands below are the Google counterpart of the Parabricks command above. The output from these commands will be identical to the output from the above command. See the Output Comparison page for comparing the results.

Copy
Copied!

            
            # Run bwa-mem and pipe output to create sorted BAM
$ bwa mem \
    -t 32 \
    -K 10000000 \
    -R '@RG\tID:sample_rg1\tLB:lib1\tPL:bar\tSM:sample\tPU:sample_rg1' \
    <INPUT_DIR>/${REFERENCE_FILE} \
    <INPUT_DIR>/${INPUT_FASTQ_1} <INPUT_DIR>/${INPUT_FASTQ_2} | \
  gatk SortSam \
      --java-options -Xmx30g \
      --MAX_RECORDS_IN_RAM 5000000 \
      -I /dev/stdin \
      -O cpu.bam \
      --SORT_ORDER coordinate

# Mark Duplicates
$ gatk MarkDuplicates \
    --java-options -Xmx30g \
    -I cpu.bam \
    -O mark_dups_cpu.bam \
    -M metrics.txt

# Run deepvariant
BIN_VERSION="1.5.0"
sudo docker run \
  -v "${PWD}":"/input" \
  -v "${PWD}/output":"/output" \
  -v "${PWD}/Ref":"/reference" \
  google/deepvariant:"${BIN_VERSION}" \
  /opt/deepvariant/bin/run_deepvariant \
  --model_type WGS \
  --ref /reference/Homo_sapiens_assembly38.fasta \
  --reads /output/mark_dups_cpu.bam \
  --output_vcf /output/"${OUTPUT_VCF_FILE}" \
  --num_shards $(nproc) \
  --make_examples_extra_args "ws_use_window_selector_model=true"

--ref REF
--in-fq [IN_FQ [IN_FQ ...]]
--in-se-fq [IN_SE_FQ [IN_SE_FQ ...]]
--knownSites KNOWNSITES
--interval-file INTERVAL_FILE
--pb-model-file PB_MODEL_FILE
--out-recal-file OUT_RECAL_FILE
--out-bam OUT_BAM
--out-variants OUT_VARIANTS
--out-duplicate-metrics OUT_DUPLICATE_METRICS
--proposed-variants PROPOSED_VARIANTS

Tool Options:

-L INTERVAL, --interval INTERVAL
--bwa-options BWA_OPTIONS
--no-warnings
--filter-flag FILTER_FLAG
--skip-multiple-hits
--min-read-length MIN_READ_LENGTH
--align-only
--no-markdups
--fix-mate
--markdups-assume-sortorder-queryname
--markdups-picard-version-2182
--optical-duplicate-pixel-distance OPTICAL_DUPLICATE_PIXEL_DISTANCE
--read-group-sm READ_GROUP_SM
--read-group-lb READ_GROUP_LB
--read-group-pl READ_GROUP_PL
--read-group-id-prefix READ_GROUP_ID_PREFIX
--standalone-bqsr
--max-read-length-fq2bamfast MAX_READ_LENGTH_FQ2BAMFAST
--min-read-length-fq2bamfast MIN_READ_LENGTH_FQ2BAMFAST
--disable-use-window-selector-model
--gvcf
--norealign-reads
--sort-by-haplotypes
--keep-duplicates
--vsc-min-count-snps VSC_MIN_COUNT_SNPS
--vsc-min-count-indels VSC_MIN_COUNT_INDELS
--vsc-min-fraction-snps VSC_MIN_FRACTION_SNPS
--vsc-min-fraction-indels VSC_MIN_FRACTION_INDELS
--min-mapping-quality MIN_MAPPING_QUALITY
--min-base-quality MIN_BASE_QUALITY
--mode MODE
--alt-aligned-pileup ALT_ALIGNED_PILEUP
--variant-caller VARIANT_CALLER
--add-hp-channel
--parse-sam-aux-fields
--use-wes-model
--include-med-dp
--normalize-reads
--pileup-image-width PILEUP_IMAGE_WIDTH
--channel-insert-size
--no-channel-insert-size
--max-read-size-512
--prealign-helper-thread
--track-ref-reads
--phase-reads
--dbg-min-base-quality DBG_MIN_BASE_QUALITY
--ws-min-windows-distance WS_MIN_WINDOWS_DISTANCE
--channel-gc-content
--channel-hmer-deletion-quality
--channel-hmer-insertion-quality
--channel-non-hmer-insertion-quality
--skip-bq-channel
--aux-fields-to-keep AUX_FIELDS_TO_KEEP
--vsc-min-fraction-hmer-indels VSC_MIN_FRACTION_HMER_INDELS
--vsc-turn-on-non-hmer-ins-proxy-support
--consider-strand-bias
--p-error P_ERROR
--channel-ins-size
--max-ins-size MAX_INS_SIZE
--disable-group-variants
--filter-reads-too-long

Performance Options:

--fq2bamfast
--gpuwrite
--gpuwrite-deflate-algo GPUWRITE_DEFLATE_ALGO
--gpusort
--use-gds
--memory-limit MEMORY_LIMIT
--low-memory
--num-cpu-threads-per-stage NUM_CPU_THREADS_PER_STAGE
--bwa-nstreams BWA_NSTREAMS
--bwa-cpu-thread-pool BWA_CPU_THREAD_POOL
--num-cpu-threads-per-stream NUM_CPU_THREADS_PER_STREAM
--num-streams-per-gpu NUM_STREAMS_PER_GPU
--run-partition
--gpu-num-per-partition GPU_NUM_PER_PARTITION
--max-reads-per-partition MAX_READS_PER_PARTITION
--partition-size PARTITION_SIZE
--read-from-tmp-dir

Common options:

--logfile LOGFILE
--tmp-dir TMP_DIR
--with-petagene-dir WITH_PETAGENE_DIR
--keep-tmp
--no-seccomp-override
--version

GPU options:

--num-gpus NUM_GPUS

Note

The --in-fq option takes the names of two FASTQ files, optionally followed by a quoted read group. The FASTQ filenames must not start with a hyphen.

deepvariant_germline

Quick Start

Compatible Google DeepVariant Commands

Models for additional GPUs