Skip to content

Docs CSC now features an automatic Finnish translation. Click here for more information.

Warning!

Puhti and Mahti computing services have been decommissioned and no new jobs are accepted or executed on its compute nodes. Puhti and Mahti login nodes and storage services are planned to remain available until 15 October 2026. Clean up unnecessary files and move any data you need to keep by 31 August 2026. See the Roihu data migration guide for instructions on transferring your data to Roihu.

BUSCO

BUSCO (Benchmarking Universal Single-Copy Orthologs) assesses the completeness of genome assemblies, gene sets and transcriptomes by searching for a set of orthologs that are expected to be present as single-copy genes in a given lineage.

License

Free to use and open source under the MIT License.

Available

  • Roihu: 5.4.3, 6.1.0, via the bio-apps module.

Usage

BUSCO is part of the bio-apps collection on Roihu. Load the bio-apps module tree and then the BUSCO module:

module load bio-apps/v202603
module load busco/6.1.0

BUSCO is run with the busco command, specifying the input, the mode (genome, proteins or transcriptome) and a lineage dataset:

busco -i genome.fa -m genome -l eukaryota_odb12 -o result -c 8

Lineage datasets

BUSCO downloads the required lineage datasets automatically into a busco_downloads directory in your working directory. Run BUSCO from your project's /scratch directory so there is space for these datasets.

Shared reference databases

CSC plans to provide shared reference databases at a central location on Roihu. This is still being set up. Until it is available, let BUSCO download the lineage datasets to a writable location, or download them yourself with busco --download <lineage>.

Augustus gene predictor

If you run BUSCO with the Augustus gene predictor (--augustus), Augustus needs a writable configuration directory, because it writes trained species parameters there. The module provides a snapshot of this configuration via the $CONFIG_TEMPLATE environment variable. Unpack it to a writable location and point AUGUSTUS_CONFIG_PATH at it before running BUSCO:

tar -xzf $CONFIG_TEMPLATE -C /scratch/<project>/
export AUGUSTUS_CONFIG_PATH=/scratch/<project>/config

The default gene predictor (Metaeuk/Miniprot) does not require this step.

Example batch script

#!/bin/bash
#SBATCH --job-name=busco
#SBATCH --account=<project>
#SBATCH --output=output_%j.txt
#SBATCH --error=errors_%j.txt
#SBATCH --partition=small
#SBATCH --time=08:00:00
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=8
#SBATCH --mem-per-cpu=4G

module load bio-apps/v202603
module load busco/6.1.0

busco -i genome.fa -m genome -l eukaryota_odb12 -o result -c $SLURM_CPUS_PER_TASK

Replace <project> with your CSC project (for example project_2001234).

See creating a batch job script for Roihu for more information about running batch jobs.

Support

CSC Service Desk

More information