Skip to content

Docs CSC now features an automatic Finnish translation. Click here for more information.

Warning!

Puhti and Mahti computing services have been decommissioned and no new jobs are accepted or executed on its compute nodes. Puhti and Mahti login nodes and storage services are planned to remain available until 15 October 2026. Clean up unnecessary files and move any data you need to keep by 31 August 2026. See the Roihu data migration guide for instructions on transferring your data to Roihu.

MMseqs2

MMseqs2 (Many-against-Many sequence searching) is a software suite for very fast and sensitive searching and clustering of large protein and nucleotide sequence sets.

License

Free to use and open source under GNU GPLv3.

Available

  • Roihu-CPU: 18-8cc5c, via the bio-apps module.
  • Roihu-GPU: 18-8cc5c (GPU-accelerated), via the bio-apps module.

On GPU nodes, the MMseqs2 build is compiled with CUDA support and can use the GH200 GPUs to accelerate searches.

Usage

MMseqs2 is part of the bio-apps collection on Roihu. Load the bio-apps module tree and then the MMseqs2 module:

module load bio-apps/v202603
module load mmseqs2/18-8cc5c

A typical search converts the query and target FASTA files to MMseqs2 databases and runs mmseqs easy-search:

mmseqs easy-search query.fasta target.fasta results.m8 tmp --threads 8

Example batch script

#!/bin/bash
#SBATCH --job-name=mmseqs2
#SBATCH --account=<project>
#SBATCH --output=output_%j.txt
#SBATCH --error=errors_%j.txt
#SBATCH --partition=small
#SBATCH --time=04:00:00
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=8
#SBATCH --mem-per-cpu=4G

module load bio-apps/v202603
module load mmseqs2/18-8cc5c

mmseqs easy-search query.fasta target.fasta results.m8 tmp --threads $SLURM_CPUS_PER_TASK

Replace <project> with your CSC project (for example project_2001234). To use GPU acceleration, run MMseqs2 in a GPU batch job on a GH200 node; see creating a batch job script for Roihu.

Support

CSC Service Desk

More information