Skip to content

Docs CSC now features an automatic Finnish translation. Click here for more information.

Warning!

Puhti and Mahti computing services have been decommissioned and no new jobs are accepted or executed on its compute nodes. Puhti and Mahti login nodes and storage services are planned to remain available until 15 October 2026. Clean up unnecessary files and move any data you need to keep by 31 August 2026. See the Roihu data migration guide for instructions on transferring your data to Roihu.

Applications by discipline

Note

In addition to technical support, CSC also provides expert consulting in questions related to sciences and methods. For more details, see our science-specific support pages at research.csc.fi or directly contact our Service Desk .

Biosciences

  • ABySS — De novo, parallel, paired-end sequence assembler
  • AdmixTools — Inference of population admixture from f-statistics
  • ADMIXTURE — Maximum-likelihood estimation of individual ancestries from SNP data
  • Alphafold — Protein 3D structure prediction
  • ANGSD — Analysis of next-generation sequencing data via genotype likelihoods
  • antiSMASH — Detection of secondary-metabolite biosynthesis gene clusters in microbial genomes
  • ASTRAL — Coalescent-based species-tree estimation from gene trees
  • AUGUSTUS — Gene prediction in eukaryotic genomic sequences
  • BamTools — Tools for working with BAM formatted files
  • Barrnap — Rapid ribosomal RNA (rRNA) prediction
  • BayeScan — Tool for identifying candidate loci under natural selection based on allele frequencies in populations
  • BBMap — BBTools short-read aligner and sequence-processing suite
  • BCFtools — Variant calling and VCF/BCF manipulation
  • BEDOPS — Set operations on genomic intervals (BED)
  • BEDTools — Genome arithmetic on BED/VCF/GFF intervals
  • Bio-apps — Access module to a collection of applications and software often used in biosciences
  • BioPerl — Perl environment with bioperl extension
  • Biopython — Python environment with biopython and other bioinformatics related Python libraries
  • BLAST — Sequence similarity search tool for nucleotides and proteins
  • Bowtie — Ultrafast, memory-efficient short read aligner
  • Bowtie2 — Short read aligner
  • Bracken — Species abundance estimation from Kraken2 output
  • BRAKER — Automatic genome annotation pipeline for eukaryotes
  • BUSCO — Genome/transcriptome completeness assessment via orthologs
  • BWA — Short read aligner
  • bwa-mem2 — Faster successor to bwa mem
  • Canu — Long-read (PacBio/Nanopore) genome assembler
  • CD-HIT — Sequence clustering and redundancy removal tool
  • CheckM2 — Genome quality assessment via machine learning
  • Chipster  — Easy-to-use analysis platform for RNA-seq, single cell RNA-seq and other NGS data
  • Chipster_genomes — Tool to download aligner indexes used by Chipster to Puhti
  • CLUMPP — Alignment of replicate cluster assignments from population-structure analyses
  • Clustal Omega — Multiple sequence alignment
  • ClustalW — Multiple sequence alignment
  • CryoSPARC — Tool to analyse Cryo-EM data on Puhti/Mahti
  • Cufflinks — Transcript assembly and differential expression for RNA-Seq
  • Cutadapt — Trimming high-throughput sequencing reads
  • deepTools — Tools for exploring deep-sequencing coverage data
  • Diamond — Sequence similarity search tool for proteins and nucleotides
  • Dorado — GPU-accelerated Oxford Nanopore basecaller
  • eggNOG-mapper — Functional annotation of sequences via orthology (eggNOG)
  • EMBOSS — Toolkit for classical sequence analysis
  • Entrez Direct — Entrez direct - command line tool to search and retrieve data from NCBI
  • Exonerate — A generic tool for pairwise sequence comparison
  • fastp — Fast all-in-one FASTQ preprocessing and QC
  • FastQC — Quality control tool for high throughput sequence data
  • FASTX-Toolkit — FASTA/FASTQ short-read preprocessing tools
  • Freebayes — Genetic variant detector
  • gapseq — Genome-scale metabolic network reconstruction and analysis
  • GATK — Genome Analysis Toolkit for variant discovery
  • GetOrganelle — Assembly of organelle genomes from whole-genome sequencing data
  • GOLD — Protein Ligand Docking Software
  • Grace — Plotting tool for xvg-files in particular
  • GROMACS — Fast and versatile classical molecular dynamics
  • GTDB-Tk — Taxonomic classification of bacterial and archaeal genomes using the GTDB
  • HADDOCK3 — High Ambiguity Driven biomolecular DOCKing
  • HISAT2 — Spliced aligner for RNA-seq and DNA reads
  • HMMER — Toolkit to create and use sequence profile hidden Markov models
  • HUMAnN — Profiling microbial pathways with metagenomic data
  • HybPiper — Target-capture (Hyb-Seq) locus recovery for phylogenomics
  • HyPhy — Hypothesis testing on phylogenies (selection analysis)
  • IGV — Integrative Genomics Viewer - interactive genome browser
  • Illumina BaseSpace — Command line client for retrieving data from the Illumina BaseSpace environment
  • inStrain — Strain-level population genomics from metagenomic mappings
  • InterProScan — Protein signature/motif search tool
  • iPyrad — toolkit for population genetic and phylogenetic studies of restriction-site associated genomic data sets (e.g., RAD, ddRAD, GBS)
  • IQ-TREE — Maximum-likelihood phylogenetic inference
  • Jellyfish — Fast k-mer counting
  • Kraken — Taxonomic sequence classification system
  • Krona visualization tool — Visualization tool for taxonomic classification and other hierarchical data
  • Lazypipe — A stand-alone pipeline for identifying viruses in host-associated or environmental samples
  • MACS2/3 — ChIP-Seq analysis tool
  • Maestro — Versatile drug discovery and materials modeling suite
  • MAFFT — Multiple sequence alignment
  • Mash — Fast genome/metagenome distance estimation (MinHash)
  • MaxQuant  — A proteomics software for processing of Mass-spectromtery data
  • medaka — Nanopore consensus and variant calling
  • Megahit — Metagenomics assembly
  • MEME Suite — Motif discovery and analysis (MEME Suite)
  • MetaBAT — Metagenome binning (MetaBAT2)
  • MetaPhlAn — Profiling the composition of microbial communities with metagenomic data
  • MICOM — Metabolic modelling of microbial communities
  • Minimap2 — Short read aligner
  • MIRA — Whole genome shotgun and EST sequence Assembler
  • MMseqs2 — Very fast protein search and clustering
  • Mothur — Package for microbial community analysis of amplicon sequencing data
  • MrBayes — Program for inferring phylogenies using Bayesian methods
  • MUMmer — Genome alignment (MUMmer)
  • MUSCLE — Multiple sequence alignment (v3)
  • MUSCLE5 — Multiple sequence alignment (v5)
  • NCBI C++ Toolkit — NCBI C++ Toolkit libraries and command-line applications
  • Nextflow — Nextflow is a scientific workflow management system for creating scalable, portable, and reproducible workflows
  • PANNZER2/SANSPANZ — Automatic protein annotation tool
  • PHYLIP — PHYLIP phylogeny inference package
  • Picard Tools — Tools for working with SAM,BAM,CRAM and VCF files
  • Prodigal — Prokaryotic gene prediction
  • Prokka — Rapid prokaryotic genome annotation
  • QIIME — Package for microbial community analysis of amplicon sequencing data
  • RAxML — Program for inferring phylogenies with likelihood
  • RAxML-NG — Maximum-likelihood phylogenetic inference (RAxML-NG)
  • Roary — Pan genome pipeline
  • run_dbcan — Automated carbohydrate-active enzyme (CAZyme) annotation
  • SALMON — Program to produce transcript-level quantification estimates from RNA-seq data
  • SameStr — Strain-level sharing analysis from metagenomic SNV profiles
  • SAMtools — Utilities for managing SAM/BAM formatted alignment files
  • SeqKit — Cross-platform FASTA/FASTQ toolkit
  • Seqtk — Tool for processing sequences in the FASTA or FASTQ format
  • Snakemake — Snakemake is a scientific workflow management system for creating scalable, portable, and reproducible workflows
  • SortMeRNA — Filtering and sorting of rRNA reads from (meta)transcriptomic data
  • SPAdes — Genome assembly
  • SRA Toolkit — NCBI SRA Toolkit for accessing and converting SRA data
  • Stacks — Pipeline for building loci from short-read sequences (e.g. RAD-seq data)
  • STAR — Short read aligner
  • SteadierCom — Steady-state metabolic simulation of microbial communities
  • StrAuto — Automation and parallelization of STRUCTURE analysis
  • StringTie — Transcript assembly and quantification for RNA-Seq
  • Structure — Inference of population structure in genetics
  • Structure Harvester — Post-processing of STRUCTURE results (Evanno method)
  • TopHat — Splice junction mapper for RNA-Seq reads
  • Trimmomatic — Trim Illumina paired-end and single-read data
  • Trinity — Transcriptome assembly tool
  • VCFtools — VCF manipulation and statistics
  • Velvet — Genome assembler
  • VirusDetect — Virus identification with sRNA data
  • VMD — Molecular visualization program
  • VSEARCH — Versatile sequence search and clustering
  • wtdbg2 — Fast assembler for long-read data
  • XHMM (eXome-Hidden Markov Model) — Copy number variation calling from targeted sequencing data

Chemistry

  • Amber — Molecular dynamics suite
  • AMS — Modelling suite providing the ADF engine
  • AMS-GUI — AMS integrated GUI
  • COSMO-RS — Toolbox for the prediction of fluid phase thermodynamic properties using the COSMO-RS model
  • CP2K — DFT, quantum chemistry, QM/MM, AIMD etc. in particular for periodic systems
  • CSD — Cambridge Crystallographic Database - organic and metallo-organic crystal structures and tools
  • Gaussian — Versatile computational chemistry package
  • GOLD — Protein Ligand Docking Software
  • GPAW — Versatile DFT package
  • GROMACS — Fast and versatile classical molecular dynamics
  • HADDOCK3 — High Ambiguity Driven biomolecular DOCKing
  • LAMMPS — Fast molecular dynamics engine with large force field selection
  • Maestro — Versatile drug discovery and materials modeling suite
  • Molden — Processing program for molecular and electronic structure calculations
  • MOLPRO — Package for accurate ab initio quantum chemistry calculations
  • NAMD — Highly scalable classical molecular dynamics
  • NMRLipids — NMRLipids databank containing MD simulations
  • NWChem — A computational chemistry software package designed to perform well on parallel HPC systems
  • Open Babel — Program to interconvert file formats currently used in molecular modeling
  • ORCA — General-purpose quantum chemistry package
  • PLUMED — Library and tools for enhanced sampling methods
  • Quantum ESPRESSO — Electronic-structure calculations and materials modeling at the nanoscale
  • TmoleX — GUI for setting up and analyzing TURBOMOLE jobs
  • TURBOMOLE — Fast and robust quantum chemistry program package
  • VASP — Ab initio DFT electronic structures
  • VMD — Molecular visualization program

Computational Engineering

  • Abaqus — Dassault Systemes' SIMULIA academic research suite
  • Ansys — Ansys Academic research CFD
  • COMSOL Multiphysics — General-purpose simulation software
  • Elmer — Open source multi-physics FEM package
  • OpenFOAM — OpenFOAM® is the leading free, open source software for computational fluid dynamics (CFD)
  • PALM — Meteorological model system for atmospheric and oceanic boundary-layer flows
  • Star-CCM+ — Computational Fluid Dynamics software by Siemens Digital Industries Software

Data Analytics and Machine Learning

  • JAX — Autograd and XLA, brought together for high-performance machine learning
  • Python Data — Collection of Python libraries for data analytics and machine learning
  • PyTorch — Machine learning framework for Python
  • Spark — High-performance distributed computing framework
  • TensorFlow — Deep learning library for Python
  • vLLM — A fast and easy-to-use library for LLM inference and serving
  • Whisper — General-purpose speech recognition model

Geosciences

  • ArcGIS Python API — Spatial analysis and data science
  • CloudCompare — for visualizing, editing and processing point clouds
  • GDAL — for geospatial data formats
  • Geoconda — Python libraries for spatial analysis
  • GRASS GIS — General purpose GIS software family for viewing, editing and analysing geospatial data
  • LAStools — for LiDAR datasets
  • OpenDroneMap (ODM) — for processing aerial drone imagery
  • Orfeo ToolBox — for remote sensing applications
  • PDAL — for point cloud translations and processing
  • Python-geo — Python libraries for spatial analysis
  • QGIS — General purpose GIS software family for viewing, editing and analysing geospatial data
  • R for GIS — R spatial analysis libraries
  • SAGA GIS — General purpose GIS software family for viewing, editing and analysing geospatial data
  • SNAP — for remote sensing applications
  • WhiteboxTools — an advanced geospatial data analysis platform
  • Zonation — Spatial conservation prioritization framework

Language Research and Other Digital Humanities and Social Sciences

  • eBay's tsv-utils — Utilities for manipulating large tabular data files
  • Finnish Tagtools  — Finnish Tagtools
  • HeLI-OTS — Off-the-shelf language identifier with language models for 220 languages
  • HFST  — Helsinki Finite-State Transducer Technology
  • HFST-fi  — Helsinki Finite-State Technology for Finnish
  • HFST-sv  — Helsinki Finite-State Technology for Swedish
  • kp-spell (enchant) — Finnish and Swedish spell-checking via the enchant interface
  • openSMILE — Toolkit for extracting audio features for speech and music analysis
  • Praat — Toolkit for annotating, processing and analysing speech and other audio samples
  • trankit — Transformer-based Python toolkit for multilingual Natural Language Processing (NLP)
  • UDPipe — Trainable pipeline for tokenization, tagging, lemmatization and dependency parsing
  • vrt-tools — Tools for converting VRT (Vertical Text) corpus files

Mathematics and Statistics

  • IDL — Programming Language, Numeric Analysis, Manipulation and Visualization of Scientific Data
  • Julia Language — High-level, high-performance dynamic programming language for numerical computing
  • MATLAB — High-level technical computing language
  • Octave — High-level interpreted language for numerical computations
  • Python — The programming language and its modules at CSC
  • r-env — R and RStudio Server
  • RStudio IDE — Integrated development environment for R
  • SageMath — Free open-source mathematics software system

Physics

  • VASP — Ab initio DFT electronic structures

Quantum

  • Cirq-on-iqm — open-source cirq adapter for quantum computing
  • Pennylane — Free open-source software framework for quantum machine learning and quantum computing
  • Qiskit — open-source toolkit for useful quantum computing
  • Qiskit-on-iqm — open-source qiskit adapter for quantum computing

Atmosphere research

  • CDO — Command line tools to manipulate and analyse Climate and NWP model Data

Miscellaneous

  • Accelerated visualization — A selection of GPU accelerated visualization applications
  • Blender — 3D modeling, visualization and rendering software
  • compute-sanitizer — Functional correctness checking suite included in the CUDA toolkit
  • cProfile — Built-in profiler for Python programs
  • cuda-gdb — Nvidia extension of the GNU debugger GDB
  • DDT — Parallel debugger
  • Desktop — Remote desktop environment
  • FFmpeg — Tools and libraries for recording, converting and streaming audio and video
  • FireWorks — FireWorks is a free, open-source tool for defining, managing and executing workflows with multiple steps and complex dependencies
  • gdb — GNU debugger for compiled programs
  • HyperQueue — Scheduler for sub-node tasks
  • Intel Trace Analyzer and Collector (ITAC) — MPI profiling and tracing tool
  • Intel VTune Profiler — Performance analysis tool for single core and threading performance
  • Julia-Jupyter — Interactive computational environment for Julia
  • Jupyter — Interactive computational environment for Python
  • Jupyter for courses — A version of the Jupyter app for course environments
  • ncu — Nvidia CUDA kernel profiler
  • nsys — Nvidia GPU and CPU profiler
  • nvprof — Nvidia profiling tool that collects and views profiling data
  • ParaView — Free open-source visualization application
  • pdb — Built-in Python debugger
  • perf — Command line tool for performance analysis
  • Scalasca — Performance profiler for parallel programs
  • TensorBoard — The visualization toolkit for TensorFlow
  • VisIt — Free open-source visualization application
  • Visual Studio Code — Source code editor