niekwit/crispr-screens

Snakemake workflow for CRISPR-Cas9 screen analysis

Overview

Latest release: v1.2.0, Last update: 2026-10-01

Share link: https://snakemake.github.io/snakemake-workflow-catalog?wf=niekwit/crispr-screens

Quality control: linting: failed formatting: passed

Topics: bioinformatics-pipeline crispr-screen-analysis snakemake-workflow

Wrappers: bio/cutadapt/se bio/fastqc bio/multiqc

Deployment

Step 1: Install Snakemake and Snakedeploy

Snakemake and Snakedeploy are best installed via the Conda package manager. It is recommended to install conda via Miniforge. Run

conda create -c conda-forge -c bioconda -c nodefaults --name snakemake snakemake snakedeploy

to install both Snakemake and Snakedeploy in an isolated environment. For all following commands ensure that this environment is activated via

conda activate snakemake

For other installation methods, refer to the Snakemake and Snakedeploy documentation.

Step 2: Deploy workflow

With Snakemake and Snakedeploy installed, the workflow can be deployed as follows. First, create an appropriate project working directory on your system and enter it:

mkdir -p path/to/project-workdir
cd path/to/project-workdir

In all following steps, we will assume that you are inside of that directory. Then run

snakedeploy deploy-workflow https://github.com/niekwit/crispr-screens . --tag v1.2.0

Snakedeploy will create two folders, workflow and config. The former contains the deployment of the chosen workflow as a Snakemake module, the latter contains configuration files which will be modified in the next step in order to configure the workflow to your needs.

Step 3: Configure workflow

To configure the workflow, adapt config/config.yml to your needs following the instructions below.

Step 4: Run workflow

The deployment method is controlled using the --software-deployment-method (short --sdm) argument.

To run the workflow with automatic deployment of all required software via conda/mamba, use

snakemake --cores all --sdm conda

To run the workflow using apptainer/singularity, use

snakemake --cores all --sdm apptainer

To run the workflow using a combination of conda and apptainer/singularity for software deployment, use

snakemake --cores all --sdm conda apptainer

Snakemake will automatically detect the main Snakefile in the workflow subfolder and execute the workflow module that has been defined by the deployment in step 2.

For further options such as cluster and cloud execution, see the docs.

Step 5: Generate report

After finalizing your data analysis, you can automatically generate an interactive visual HTML report for inspection of results together with parameters and code inside of the browser using

snakemake --report report.zip

Configuration

The following section is imported from the workflow’s config/README.md.

CONFIGURATION

config.yml

Use config.yml to provide information about your experiment. See workflow/schemas/config.schema.yaml for the full, authoritative list of keys.

lib_info

Under lib_info, set library_file to the path of the sgRNA library CSV file (in the resources folder) and species to the species the library targets (e.g. human).

cutadapt_args

Extra arguments passed to cutadapt when trimming the raw reads, as a single string (e.g. "-q 20 -l 20"). --minimum-length 10 is always added by the workflow and does not need to be included. See the cutadapt manual for all options.

csv

A CSV file (in the resources folder) must be provided that contains the sgRNA names, sequences, and gene names, each in its own column. Under name_column, sequence_column, and gene_column (0-based), set the column numbers of these columns. If no fasta file is available for building the Bowtie index, one is generated from this CSV file.

bowtie_args

Extra arguments passed to bowtie (v1) when aligning the trimmed reads to the sgRNA index. The read file (-q), index (-x) and number of threads (-p) are set by the workflow and should not be included here.

The default (-v 1 -m 1) allows one mismatch (-v) and discards reads that align to more than one sgRNA (-m 1). Note that Bowtie aligns end-to-end, so reads are expected to be trimmed to the length of the sgRNAs (see cutadapt_args). Reads that are longer than the sgRNA they originate from (e.g. in libraries with variable sgRNA lengths) will not align unless they are trimmed accordingly. Use -v 0 to only allow perfect matches. See the Bowtie manual for all options.

stats

Each statistical tool (bagel2, mageck, drugz) has its own run: True/False switch to enable or disable it.

  • crisprcleanr: settings for the CRISPRcleanR normalisation step. It always runs ahead of BAGEL2, and optionally ahead of MAGeCK/DrugZ (see apply_crisprcleanr under those sections). library_name is either the name of one of the libraries bundled with CRISPRcleanR (see the comments in config.yml for the full list) or any other name, in which case a CRISPRcleanR-formatted library file must be provided (see the main README.md, e.g. via annotate_sgrna_coordinates.py, for how to generate one). min_reads sets the minimum read count in the control sample for an sgRNA to be kept.

  • bagel2: custom_gene_lists.essential_genes/non_essential_genes can point to custom essential/non-essential gene list files (none uses BAGEL2’s own default lists). extra_args.bf/pr pass extra arguments to the BAGEL2 bf and pr subcommands respectively.

  • mageck: command is test (pairwise, needs config/stats.csv) or mle (needs one or more design matrices, see mle.design_matrix; each file must be placed in the config directory). extra_mageck_arguments passes extra arguments to the MAGeCK test/mle command. mageck_control_genes is all or a path to a file with control gene names, one per line, used to build the null distribution/normalisation instead of the whole library. apply_CNV_correction and cell_line enable copy-number correction of MAGeCK results.

  • drugz: extra passes extra arguments to the drugz command.

  • string_db: run STRING-db enrichment analysis on the MAGeCK, DrugZ, and BAGEL2 results. For MAGeCK/DrugZ, data selects enriched, depleted, or both gene sets and fdr sets the significance threshold. BAGEL2 has no enriched set, so it always runs depleted-only, selecting genes by bf_cutoff (genes with a Bayes Factor above this, from the .bf output) instead of fdr. top_genes, if not 0, overrides fdr/bf_cutoff and takes the top N genes by significance/BF instead.

stats.csv

Pairwise comparisons for MAGeCK (command: test), DrugZ, and BAGEL2 are defined in config/stats.csv, with test and control columns naming samples exactly as they appear in the reads directory (without the .fastq.gz/.cram extension). Multiple replicate samples can be combined in one comparison by separating their names with a semicolon. An optional bagel2_only column (y/n per row) can be added to run some comparisons only through BAGEL2 (and CRISPRcleanR) and the rest only through MAGeCK/DrugZ; if the column is absent, every row is available to the enabled tools. See the main documentation for details and examples.

ngs_tracker

Optional integration with NGS Tracker to register the workflow run and attach output files. Set enabled: false to skip this entirely. When enabled, base_url and project_id identify the NGS Tracker instance/project, and files lists the paths (glob patterns allowed) and types of files to attach after a successful run.

resources

Computational resources (threads, runtime, memory) are set per rule in the workflow itself, not in config.yml. To override them, use a Snakemake profile (--set-resources/--set-threads) or edit the relevant rule.

Workflow parameters

The following table is automatically parsed from the workflow’s config.schema.y(a)ml file.

Parameter

Type

Description

Required

Default

lib_info

yes

. library_file

string

Path to the library file with sgRNA annotations

yes

. species

string

Species of the library (e.g., human, mouse)

yes

cutadapt_args

string

Extra arguments for cutadapt

csv

yes

. name_column

integer

Column number with sgRNA names

yes

. sequence_column

integer

Column number with sgRNA sequences

yes

. gene_column

integer

Column number with gene names

bowtie_args

string

Extra arguments for Bowtie

yes

stats

yes

. crisprcleanr

. . library_name

string

sgRNA library name (e.g., AVANA_Library, Brunello_Library, etc.)

yes

. . min_reads

integer

Keep sgRNAs with at least this many reads in control sample

yes

. bagel2

. . run

boolean

Perform bagel2 analysis

yes

. . custom_gene_lists

yes

. . . essential_genes

string

Path to custom gene list for bagel2 analysis

. . . non_essential_genes

string

Path to custom gene list for bagel2 analysis

. . extra_args

yes

. . . bf

string

Extra arguments for bagel2 bf subcommand

. . . pr

string

Extra arguments for bagel2 pr subcommand

. mageck

. . run

boolean

Perform mageck analysis

yes

. . command

string

mageck command to run (test or mle)

. . mle

. . . design_matrix

array

Paths to design matrices for mageck mle

. . apply_crisprcleanr

boolean

Apply crisprcleanr to mageck results (recommended for genome-wide libraries)

. . extra_mageck_arguments

string

Extra arguments for mageck

yes

. . mageck_control_genes

string

All or path to file with control genes

yes

. . apply_CNV_correction

boolean

Apply CNV correction to mageck results

yes

. . cell_line

string

Cell line for CNV correction

yes

. drugz

. . run

boolean

Perform drugZ analysis

yes

. . apply_crisprcleanr

boolean

Apply crisprcleanr before drugZ analysis

. . extra

string

Extra arguments for drugZ

yes

. string_db

. . run

boolean

Perform STRING-db analysis on mageck, drugz, and bagel2 results

yes

. . data

string

enriched, depleted, or both (mageck/drugz only; bagel2 is always depleted-only)

yes

. . fdr

number

FDR threshold for significant genes (mageck/drugz)

yes

. . bf_cutoff

number

BAGEL2 Bayes Factor cutoff; genes with BF above this are considered depleted/essential (bagel2 only)

yes

. . top_genes

integer

Number of top genes to consider for STRING-db analysis (overrides fdr/bf_cutoff, use 0 to disable)

yes

ngs_tracker

. enabled

boolean

Set to false to skip NGS Tracker registration

yes

. base_url

string

NGS Tracker API base URL

yes

. project_id

integer

Project ID from the NGS Tracker project detail page

yes

. workflow_name

string

. workflow_tag

string

. workflow_system

string

. description

string

. tags

array

. files

array

Linting and formatting

Linting results
 1/home/runner/work/snakemake-workflow-catalog/snakemake-workflow-catalog/.pixi/envs/default/lib/python3.13/site-packages/google/auth/transport/grpc.py:43: FutureWarning: grpcio < 1.83.0 does not support Post-Quantum Cryptography (PQC). Support for non-PQC environments is deprecated. In April 2027, google-auth will raise its minimum requirements to enforce grpcio >= 1.83.0. For more details on Google Cloud's post-quantum security migration, visit: https://cloud.google.com/security/resources/post-quantum-cryptography
 2  warnings.warn(
 3Workflow version: v1.2.0
 4No validator found for JSON Schema version identifier 'http://json-schema.org/draft-06/schema#'
 5Defaulting to validator for JSON Schema version 'https://json-schema.org/draft/2020-12/schema'
 6Note that schema file may not be validated correctly.
 7No validator found for JSON Schema version identifier 'http://json-schema.org/draft-06/schema#'
 8Defaulting to validator for JSON Schema version 'https://json-schema.org/draft/2020-12/schema'
 9Note that schema file may not be validated correctly.
10ValueError in file "/tmp/tmpl31le160/niekwit-crispr-screens-cb5f0f8/workflow/scripts/general_functions.smk", line 190:
11No fastq or cram files found in reads directory
12  File "/tmp/tmpl31le160/niekwit-crispr-screens-cb5f0f8/workflow/Snakefile", line 33, in <module>
13  File "/tmp/tmpl31le160/niekwit-crispr-screens-cb5f0f8/workflow/scripts/general_functions.smk", line 190, in sample_names
Formatting results
All tests passed!