Skip to content

Puhti and Mahti computing services have been decommissioned. Puhti and Mahti login nodes and storage services will remain available until 15 October 2026, but are no longer covered by service contracts. Please clean up and migrate your data to Roihu ASAP. See Roihu data migration guide for instructions.

SeqKit

SeqKit is a cross-platform and ultrafast toolkit for FASTA/Q file manipulation. It provides a wide range of subcommands for common sequence operations such as statistics, searching, filtering, subsampling and format conversion.

License

Free to use and open source under the MIT License.

Available

  • Roihu-CPU: 2.10.0, 2.13.0, via the bio-apps module.

Usage

SeqKit is part of the bio-apps collection on Roihu. Load the bio-apps module tree and then the SeqKit module:

module load bio-apps/v202603
module load seqkit/2.13.0

SeqKit is run through the seqkit command followed by a subcommand. For example, to print summary statistics for a set of FASTA files:

seqkit stats *.fasta

or to filter sequences by minimum length:

seqkit seq -m 1000 input.fasta > long_sequences.fasta

Example batch script

#!/bin/bash
#SBATCH --job-name=seqkit
#SBATCH --account=<project>
#SBATCH --output=output_%j.txt
#SBATCH --error=errors_%j.txt
#SBATCH --partition=small
#SBATCH --time=01:00:00
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=4
#SBATCH --mem-per-cpu=2G

module load bio-apps/v202603
module load seqkit/2.13.0

seqkit stats -j $SLURM_CPUS_PER_TASK *.fasta

Replace <project> with your CSC project (for example project_2001234).

See creating a batch job script for Roihu for more information about running batch jobs.

Support

CSC Service Desk

More information