-
Minimap2
Minimap2
Minimap2 is a fast general-purpose alignment program to map DNA or long mRNA sequences against a large reference database. It can be used for:
- mapping of accurate short reads (preferably longer than 100 bases)
- mapping 1kb genomic reads at error rate 15% (e.g. PacBio or Oxford Nanopore genomic reads)
- mapping full-length noisy Direct RNA or cDNA reads
- mapping and comparing assembly contigs or closely related full chromosomes of hundreds of megabases in length
License
Free to use and open source under MIT License.
Available
- Roihu: 2.30, via the
bio-appsmodule.
Usage
Minimap2 is part of the bio-apps collection on Roihu. Load the bio-apps module tree and then the Minimap2 module:
Once the module is loaded, Minimap2 starts with the command:
Without any options, minimap2 takes a reference database and a query sequence file as input and produce approximate mapping, without base-level alignment (i.e. no CIGAR), in the PAF format:
If you wish to get the output in SAM format, you can use option -a.
For different data types, Minimap2 needs to be tuned for optimal performance and accuracy.
With option -x you can use case specific parameter sets, pre-defined and recommended by the Minimap2 developers.
Map long noisy genomic reads (map-pb and map-ont)
- PacBio subreads (map-pb):
- Oxford Nanopore reads (map-ont):
Map long mRNA/cDNA reads (splice)
- PacBio Iso-seq/traditional cDNA
- Nanopore 2D cDNA-seq
- Nanopore Direct RNA-seq
- mapping against SIRV control
Find overlaps between long reads (ava-pb and ava-ont)
- PacBio read overlap
- Oxford Nanopore read overlap
Map short accurate genomic reads (sr)
Note, Minimap2 does not work well with short spliced reads.
- single-end alignment
- paired-end alignment
- paired-end alignment
Full genome/assembly alignment (asm5)
- assembly to assembly
Example batch script
Minimap2 jobs should be run as batch jobs. Below is a sample batch job script for running a Minimap2 alignment on Roihu.
#!/bin/bash
#SBATCH --job-name=minimap2
#SBATCH --account=<project>
#SBATCH --output=output_%j.txt
#SBATCH --error=errors_%j.txt
#SBATCH --partition=small
#SBATCH --time=04:00:00
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=8
#SBATCH --mem=16000
module load bio-apps/v202603
module load minimap2/2.30
minimap2 -t $SLURM_CPUS_PER_TASK -ax splice -uf ref.fa iso-seq.fq > aln.sam
In the batch job example above, one task (--ntasks=1) is executed. The Minimap2 job
uses 8 cores (--cpus-per-task=8) with a total of 16 GB of memory (--mem=16000).
The maximum duration of the job is four hours (--time=04:00:00). All the cores
are assigned from one computing node (--nodes=1). In addition to the resource
reservations, you have to define the billing project for your batch job. This
is done by replacing the <project> with the name of your project. You can
use command csc-projects to see what projects you have.
You can submit the batch job file to the batch job system with the command:
See creating a batch job script for Roihu for more information about running batch jobs.
Support
More information
- More information about Minimap2 can be found from the Minimap2 home page.