Skip to content

Docs CSC now features an automatic Finnish translation. Click here for more information.

Warning!

Puhti and Mahti computing services have been decommissioned and no new jobs are accepted or executed on its compute nodes. Puhti and Mahti login nodes and storage services are planned to remain available until 15 October 2026. Clean up unnecessary files and move any data you need to keep by 31 August 2026. See the Roihu data migration guide for instructions on transferring your data to Roihu.

gapseq

gapseq performs informed prediction and analysis of bacterial metabolic pathways and produces gap-filled genome-scale metabolic models from genome sequences.

License

Free to use and open source under GNU GPLv3.

Available

  • Roihu: 2.1.0, via the bio-apps module.

Usage

gapseq is part of the bio-apps collection on Roihu. Load the bio-apps module tree and then the gapseq module:

module load bio-apps/v202603
module load gapseq/2.1.0

Reference sequence database

gapseq requires a reference sequence database for sequence searches. If a shared database is not available on Roihu, download it to a writable location such as your project's /scratch directory.

For bacterial genomes:

DB=/scratch/<project>/gapseq_db

gapseq update-sequences \
    -t Bacteria \
    -D "$DB"

For archaeal genomes, replace Bacteria with Archaea.

Specify the same database directory with -D when running gapseq.

Running gapseq

The full pipeline (pathway prediction, network building and gap-filling) can be run on a genome with the doall subcommand:

gapseq doall genome.fna.gz

Individual steps (find, find-transport, draft, fill) can also be run separately. gapseq analyses can be computationally heavy and should be run as batch jobs:

#!/bin/bash
#SBATCH --job-name=gapseq
#SBATCH --account=<project>
#SBATCH --output=output_%j.txt
#SBATCH --error=errors_%j.txt
#SBATCH --partition=small
#SBATCH --time=12:00:00
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=8
#SBATCH --mem-per-cpu=4G

module load bio-apps/v202603
module load gapseq/2.1.0

DB=/scratch/<project>/gapseq_db

gapseq doall \
    -K "$SLURM_CPUS_PER_TASK" \
    -D "$DB" \
    genome.fna.gz

Replace <project> with your CSC project (for example project_2001234).

See creating a batch job script for Roihu for more information about running batch jobs.

Support

CSC Service Desk

More information