niekwit/crispr-screens
Snakemake workflow for CRISPR-Cas9 screen analysis
Overview
Latest release: v1.2.0, Last update: 2026-10-01
Share link: https://snakemake.github.io/snakemake-workflow-catalog?wf=niekwit/crispr-screens
Quality control: linting: failed formatting: passed
Topics: bioinformatics-pipeline crispr-screen-analysis snakemake-workflow
Wrappers: bio/cutadapt/se bio/fastqc bio/multiqc
Deployment
Step 1: Install Snakemake and Snakedeploy
Snakemake and Snakedeploy are best installed via the Conda package manager. It is recommended to install conda via Miniforge. Run
conda create -c conda-forge -c bioconda -c nodefaults --name snakemake snakemake snakedeploy
to install both Snakemake and Snakedeploy in an isolated environment. For all following commands ensure that this environment is activated via
conda activate snakemake
For other installation methods, refer to the Snakemake and Snakedeploy documentation.
Step 2: Deploy workflow
With Snakemake and Snakedeploy installed, the workflow can be deployed as follows. First, create an appropriate project working directory on your system and enter it:
mkdir -p path/to/project-workdir
cd path/to/project-workdir
In all following steps, we will assume that you are inside of that directory. Then run
snakedeploy deploy-workflow https://github.com/niekwit/crispr-screens . --tag v1.2.0
Snakedeploy will create two folders, workflow and config. The former contains the deployment of the chosen workflow as a Snakemake module, the latter contains configuration files which will be modified in the next step in order to configure the workflow to your needs.
Step 3: Configure workflow
To configure the workflow, adapt config/config.yml to your needs following the instructions below.
Step 4: Run workflow
The deployment method is controlled using the --software-deployment-method (short --sdm) argument.
To run the workflow with automatic deployment of all required software via conda/mamba, use
snakemake --cores all --sdm conda
To run the workflow using apptainer/singularity, use
snakemake --cores all --sdm apptainer
To run the workflow using a combination of conda and apptainer/singularity for software deployment, use
snakemake --cores all --sdm conda apptainer
Snakemake will automatically detect the main Snakefile in the workflow subfolder and execute the workflow module that has been defined by the deployment in step 2.
For further options such as cluster and cloud execution, see the docs.
Step 5: Generate report
After finalizing your data analysis, you can automatically generate an interactive visual HTML report for inspection of results together with parameters and code inside of the browser using
snakemake --report report.zip
Configuration
The following section is imported from the workflow’s config/README.md.
CONFIGURATION
config.yml
Use config.yml to provide information about your experiment. See workflow/schemas/config.schema.yaml for the full, authoritative list of keys.
lib_info
Under lib_info, set library_file to the path of the sgRNA library CSV file (in the resources folder) and species to the species the library targets (e.g. human).
cutadapt_args
Extra arguments passed to cutadapt when trimming the raw reads, as a single string (e.g. "-q 20 -l 20"). --minimum-length 10 is always added by the workflow and does not need to be included. See the cutadapt manual for all options.
csv
A CSV file (in the resources folder) must be provided that contains the sgRNA names, sequences, and gene names, each in its own column. Under name_column, sequence_column, and gene_column (0-based), set the column numbers of these columns. If no fasta file is available for building the Bowtie index, one is generated from this CSV file.
bowtie_args
Extra arguments passed to bowtie (v1) when aligning the trimmed reads to the sgRNA index. The read file (-q), index (-x) and number of threads (-p) are set by the workflow and should not be included here.
The default (-v 1 -m 1) allows one mismatch (-v) and discards reads that align to more than one sgRNA (-m 1). Note that Bowtie aligns end-to-end, so reads are expected to be trimmed to the length of the sgRNAs (see cutadapt_args). Reads that are longer than the sgRNA they originate from (e.g. in libraries with variable sgRNA lengths) will not align unless they are trimmed accordingly. Use -v 0 to only allow perfect matches. See the Bowtie manual for all options.
stats
Each statistical tool (bagel2, mageck, drugz) has its own run: True/False switch to enable or disable it.
crisprcleanr: settings for the CRISPRcleanR normalisation step. It always runs ahead of BAGEL2, and optionally ahead of MAGeCK/DrugZ (seeapply_crisprcleanrunder those sections).library_nameis either the name of one of the libraries bundled with CRISPRcleanR (see the comments inconfig.ymlfor the full list) or any other name, in which case a CRISPRcleanR-formatted library file must be provided (see the mainREADME.md, e.g. viaannotate_sgrna_coordinates.py, for how to generate one).min_readssets the minimum read count in the control sample for an sgRNA to be kept.bagel2:custom_gene_lists.essential_genes/non_essential_genescan point to custom essential/non-essential gene list files (noneuses BAGEL2’s own default lists).extra_args.bf/prpass extra arguments to the BAGEL2bfandprsubcommands respectively.mageck:commandistest(pairwise, needsconfig/stats.csv) ormle(needs one or more design matrices, seemle.design_matrix; each file must be placed in theconfigdirectory).extra_mageck_argumentspasses extra arguments to the MAGeCKtest/mlecommand.mageck_control_genesisallor a path to a file with control gene names, one per line, used to build the null distribution/normalisation instead of the whole library.apply_CNV_correctionandcell_lineenable copy-number correction of MAGeCK results.drugz:extrapasses extra arguments to thedrugzcommand.string_db: run STRING-db enrichment analysis on the MAGeCK, DrugZ, and BAGEL2 results. For MAGeCK/DrugZ,dataselectsenriched,depleted, orbothgene sets andfdrsets the significance threshold. BAGEL2 has noenrichedset, so it always runs depleted-only, selecting genes bybf_cutoff(genes with a Bayes Factor above this, from the.bfoutput) instead offdr.top_genes, if not 0, overridesfdr/bf_cutoffand takes the top N genes by significance/BF instead.
stats.csv
Pairwise comparisons for MAGeCK (command: test), DrugZ, and BAGEL2 are defined in config/stats.csv, with test and control columns naming samples exactly as they appear in the reads directory (without the .fastq.gz/.cram extension). Multiple replicate samples can be combined in one comparison by separating their names with a semicolon. An optional bagel2_only column (y/n per row) can be added to run some comparisons only through BAGEL2 (and CRISPRcleanR) and the rest only through MAGeCK/DrugZ; if the column is absent, every row is available to the enabled tools. See the main documentation for details and examples.
ngs_tracker
Optional integration with NGS Tracker to register the workflow run and attach output files. Set enabled: false to skip this entirely. When enabled, base_url and project_id identify the NGS Tracker instance/project, and files lists the paths (glob patterns allowed) and types of files to attach after a successful run.
resources
Computational resources (threads, runtime, memory) are set per rule in the workflow itself, not in config.yml. To override them, use a Snakemake profile (--set-resources/--set-threads) or edit the relevant rule.
Workflow parameters
The following table is automatically parsed from the workflow’s config.schema.y(a)ml file.
Parameter |
Type |
Description |
Required |
Default |
|---|---|---|---|---|
lib_info |
yes |
|||
. library_file |
string |
Path to the library file with sgRNA annotations |
yes |
|
. species |
string |
Species of the library (e.g., human, mouse) |
yes |
|
cutadapt_args |
string |
Extra arguments for cutadapt |
||
csv |
yes |
|||
. name_column |
integer |
Column number with sgRNA names |
yes |
|
. sequence_column |
integer |
Column number with sgRNA sequences |
yes |
|
. gene_column |
integer |
Column number with gene names |
||
bowtie_args |
string |
Extra arguments for Bowtie |
yes |
|
stats |
yes |
|||
. crisprcleanr |
||||
. . library_name |
string |
sgRNA library name (e.g., AVANA_Library, Brunello_Library, etc.) |
yes |
|
. . min_reads |
integer |
Keep sgRNAs with at least this many reads in control sample |
yes |
|
. bagel2 |
||||
. . run |
boolean |
Perform bagel2 analysis |
yes |
|
. . custom_gene_lists |
yes |
|||
. . . essential_genes |
string |
Path to custom gene list for bagel2 analysis |
||
. . . non_essential_genes |
string |
Path to custom gene list for bagel2 analysis |
||
. . extra_args |
yes |
|||
. . . bf |
string |
Extra arguments for bagel2 bf subcommand |
||
. . . pr |
string |
Extra arguments for bagel2 pr subcommand |
||
. mageck |
||||
. . run |
boolean |
Perform mageck analysis |
yes |
|
. . command |
string |
mageck command to run (test or mle) |
||
. . mle |
||||
. . . design_matrix |
array |
Paths to design matrices for mageck mle |
||
. . apply_crisprcleanr |
boolean |
Apply crisprcleanr to mageck results (recommended for genome-wide libraries) |
||
. . extra_mageck_arguments |
string |
Extra arguments for mageck |
yes |
|
. . mageck_control_genes |
string |
All or path to file with control genes |
yes |
|
. . apply_CNV_correction |
boolean |
Apply CNV correction to mageck results |
yes |
|
. . cell_line |
string |
Cell line for CNV correction |
yes |
|
. drugz |
||||
. . run |
boolean |
Perform drugZ analysis |
yes |
|
. . apply_crisprcleanr |
boolean |
Apply crisprcleanr before drugZ analysis |
||
. . extra |
string |
Extra arguments for drugZ |
yes |
|
. string_db |
||||
. . run |
boolean |
Perform STRING-db analysis on mageck, drugz, and bagel2 results |
yes |
|
. . data |
string |
enriched, depleted, or both (mageck/drugz only; bagel2 is always depleted-only) |
yes |
|
. . fdr |
number |
FDR threshold for significant genes (mageck/drugz) |
yes |
|
. . bf_cutoff |
number |
BAGEL2 Bayes Factor cutoff; genes with BF above this are considered depleted/essential (bagel2 only) |
yes |
|
. . top_genes |
integer |
Number of top genes to consider for STRING-db analysis (overrides fdr/bf_cutoff, use 0 to disable) |
yes |
|
ngs_tracker |
||||
. enabled |
boolean |
Set to false to skip NGS Tracker registration |
yes |
|
. base_url |
string |
NGS Tracker API base URL |
yes |
|
. project_id |
integer |
Project ID from the NGS Tracker project detail page |
yes |
|
. workflow_name |
string |
|||
. workflow_tag |
string |
|||
. workflow_system |
string |
|||
. description |
string |
|||
. tags |
array |
|||
. files |
array |
Linting and formatting
Linting results
1/home/runner/work/snakemake-workflow-catalog/snakemake-workflow-catalog/.pixi/envs/default/lib/python3.13/site-packages/google/auth/transport/grpc.py:43: FutureWarning: grpcio < 1.83.0 does not support Post-Quantum Cryptography (PQC). Support for non-PQC environments is deprecated. In April 2027, google-auth will raise its minimum requirements to enforce grpcio >= 1.83.0. For more details on Google Cloud's post-quantum security migration, visit: https://cloud.google.com/security/resources/post-quantum-cryptography
2 warnings.warn(
3Workflow version: v1.2.0
4No validator found for JSON Schema version identifier 'http://json-schema.org/draft-06/schema#'
5Defaulting to validator for JSON Schema version 'https://json-schema.org/draft/2020-12/schema'
6Note that schema file may not be validated correctly.
7No validator found for JSON Schema version identifier 'http://json-schema.org/draft-06/schema#'
8Defaulting to validator for JSON Schema version 'https://json-schema.org/draft/2020-12/schema'
9Note that schema file may not be validated correctly.
10ValueError in file "/tmp/tmpl31le160/niekwit-crispr-screens-cb5f0f8/workflow/scripts/general_functions.smk", line 190:
11No fastq or cram files found in reads directory
12 File "/tmp/tmpl31le160/niekwit-crispr-screens-cb5f0f8/workflow/Snakefile", line 33, in <module>
13 File "/tmp/tmpl31le160/niekwit-crispr-screens-cb5f0f8/workflow/scripts/general_functions.smk", line 190, in sample_names
Formatting results
All tests passed!