andyngh/noHiC-Snakemake
A Snakemake reimplementation of noHiC - the personalized reference-guided contig scaffolding pipeline
Overview
Latest release: v1.1.0, Last update: 2026-09-27
Share link: https://snakemake.github.io/snakemake-workflow-catalog?wf=andyngh/noHiC-Snakemake
Quality control: linting: failed formatting: failed
Deployment
Step 1: Install Snakemake and Snakedeploy
Snakemake and Snakedeploy are best installed via the Conda package manager. It is recommended to install conda via Miniforge. Run
conda create -c conda-forge -c bioconda -c nodefaults --name snakemake snakemake snakedeploy
to install both Snakemake and Snakedeploy in an isolated environment. For all following commands ensure that this environment is activated via
conda activate snakemake
For other installation methods, refer to the Snakemake and Snakedeploy documentation.
Step 2: Deploy workflow
With Snakemake and Snakedeploy installed, the workflow can be deployed as follows. First, create an appropriate project working directory on your system and enter it:
mkdir -p path/to/project-workdir
cd path/to/project-workdir
In all following steps, we will assume that you are inside of that directory. Then run
snakedeploy deploy-workflow https://github.com/andyngh/noHiC-Snakemake . --tag v1.1.0
Snakedeploy will create two folders, workflow and config. The former contains the deployment of the chosen workflow as a Snakemake module, the latter contains configuration files which will be modified in the next step in order to configure the workflow to your needs.
Step 3: Configure workflow
To configure the workflow, adapt config/config.yml to your needs following the instructions below.
Step 4: Run workflow
The deployment method is controlled using the --software-deployment-method (short --sdm) argument.
Snakemake will automatically detect the main Snakefile in the workflow subfolder and execute the workflow module that has been defined by the deployment in step 2.
For further options such as cluster and cloud execution, see the docs.
Step 5: Generate report
After finalizing your data analysis, you can automatically generate an interactive visual HTML report for inspection of results together with parameters and code inside of the browser using
snakemake --report report.zip
Configuration
The following section is imported from the workflow’s config/README.md.
Configuration
Conventions for parsing command-line (CI) arguments
Follow these conventions when parsing the CI arguments:
Switch arguments take
"yes"/"no"to turn an option or a computational task on/off.[chained]arguments in an assembly stage are filled in automatically from the preceding stage or from the[global]section. Leave them as"". Only fill in one of these arguments if you want to use a different input file; your manually provided inputs always take precedence over the defaults.[Required]arguments must be filled in manually.We highly recommend using absolute paths when filling in your inputs.
You must use distinct output directory names for the five assembly stages of the workflow (
refpick,refpolish,clean,asm, andeval).
Section 1. [stages] - Choose which sub-workflows to run
Config keys |
CI arguments |
Main tasks |
|---|---|---|
|
|
Build (and patch) the synref using the provided pangenome graph (default: yes). |
|
|
Polish the synref or a real reference genome (default: yes). |
|
|
Decontaminate the target contig assembly (default: yes). |
|
|
Correct and scaffold the target contig assembly (default: yes). |
|
|
Check the quality of the final assembly (default: yes). |
A stage set to "yes" must have its own [Required] arguments filled in.
Section 3. [refpick] - Building the synref for your target genome
Config keys |
CI arguments |
Description |
|---|---|---|
|
|
[chained] Reads for k-mer counting; defaults to |
|
|
K-mer length (bp) for KMC-based k-mer counting (default: 29). |
|
|
Memory (in GB) for KMC-based k-mer counting (default: 32). |
|
|
Depends on |
|
|
[chained] Set outputs’ prefix. Take “global: prefix” by default. |
|
|
Output directory of the |
|
|
Number of threads for KMC (default: 1). This argument is filled automatically via the |
|
|
[Required] Path to a pangenome graph in GBZ format. |
|
|
[Required] Path to the |
|
|
Number of threads for the haplotype sampling and synref extraction steps (default: 1). This argument is filled automatically via the |
|
|
Fill in |
|
|
Required when |
|
|
Number of threads for synref patching (default: 1). This argument is filled automatically via the |
Section 4. [refpolish] - Polishing your reference genome
Config keys |
CI arguments |
Description |
|---|---|---|
|
|
[chained] Takes the synref from |
|
|
minimap2 preset used to map the reads of your target genome to the reference. Values can be |
|
|
Number of threads for |
|
|
The polishing tool to use. Fill in either |
|
|
[chained] The estimated coverage of the reads on the reference genome. Used when |
|
|
[Required] The estimated reference genome size (e.g., |
|
|
Output directory of the |
|
|
[chained] Basename of the polished reference. Take “global: prefix” by default. |
Section 5. [clean] - Decontaminating your target contig assembly
Config keys |
CI arguments |
Description |
|---|---|---|
|
|
[Required] Path to the target contig assembly to clean. Its basename (i.e., without |
|
|
Output directory of the |
|
|
[Required] Path to a FASTA file containing the adapter sequences to screen for. |
|
|
Number of threads for the adapter detection step (default: 1). This argument is filled automatically via the |
|
|
[Required] Path to a downloaded Kraken2 database directory. Fill in |
|
|
Number of threads for Kraken2 (default: 1). This argument is filled automatically via the |
|
|
Fill in |
|
|
The clade to keep. Contigs whose Kraken2 lineage does not contain this string are treated as contaminants. (default: Viridiplantae). |
|
|
Fill in |
|
|
Required when |
|
|
Number of threads for BLASTn (default: 1). This argument is filled automatically via the |
[!NOTE] The adapter check is a hard gate: if any adapter is found in your contigs, the workflow stops and points you to
1_adapter_check/{sample}.adapter_positions.bed.
Section 6. [asm] - Target contig correction, scaffolding, and gap closing
Config keys |
CI arguments |
Description |
|---|---|---|
|
|
[chained] Takes the decontaminated contigs from |
|
|
[chained] Takes the polished synref from |
|
|
Output directory of the |
|
|
[chained] Basename of the outputs. Take “global: prefix” by default. |
|
|
Fill in |
|
|
Number of threads for CRAQ; should be 5–6 lower than for the other steps (default: 1). This argument is filled automatically via the |
|
|
Filling in |
|
|
Fill in |
|
|
Number of threads for Inspector (default: 1). This argument is filled automatically via the |
|
|
Fill in |
|
|
Number of threads for |
|
|
Aggressiveness of |
|
|
Fill in |
|
|
Number of threads for gap closing (default: 1). This argument is filled automatically via the |
RagTag correct presets
See the detailed descriptions of the presets in our preprint.
|
Behavior |
Use case |
|---|---|---|
|
Uses the default settings of |
When you want an initial assembly without much correction. |
|
Sets the window size for read-based misassembly validation to 45000 and removes short alignments (< 1000 bp) between the target contigs and the reference genome (default preset). |
A balanced preset for when you want more aggressive contig correction but don’t want to force the target genome’s sequence structure to follow the reference genome too closely. |
|
Similar to |
Use this preset when you have a high-quality (possibly gapless/T2T) reference genome that is genetically close to your target genome, because this preset strictly forces the sequence structure of the target genome to follow the reference. |
|
Similar to |
Even stricter than |
|
|
Use this preset when you want to break the contigs at every point where they disagree with the reference genome (not recommended). |
[!NOTE]
RagTag correctonly acceptshifi,ont,corrected_clr, orcorrected_ontas the sequencing platform; rawclris not supported by this step.
Section 7. [eval] - Final assembly QC
Config keys |
CI arguments |
Description |
|---|---|---|
|
|
[chained] Takes the final assembly from |
|
|
[chained] Takes the synref from |
|
|
Output directory of the |
|
|
The tool used to calculate contiguity metrics. Fill in |
|
|
Number of threads for the contiguity metric calculations (default: 1). This argument is filled automatically via the |
|
|
The tool used to evaluate gene-space completeness. Fill in |
|
|
Number of threads for the gene-space completeness evaluation (default: 1). This argument is filled automatically via the |
|
|
The BUSCO/compleasm lineage, e.g., |
|
|
The OrthoDB release (default: |
|
|
[chained] Output prefix for BUSCO. Used when |
|
|
Fill in |
|
|
Number of threads for CRAQ; should be 5–6 lower than for the other steps (default: 1). This argument is filled automatically via the |
|
|
Fill in |
|
|
Number of threads for Inspector (default: 1). This argument is filled automatically via the |
|
|
Fill in |
|
|
minimap2 preset for the dot plot mapping (default: |
|
|
Number of threads for the dot plot visualization (default: 1). This argument is filled automatically via the |
[!NOTE]
compleasmis run from a sibling environment ofnohic_env_path: if the environment is/path/to/noHiC-Snakemake/workflow/envs/noHiC, compleasm is expected at/path/to/noHiC-Snakemake/workflow/envs/compleasm. BUSCO runs from the main noHiC environment.
Using SLURM in nohic-eval
Only the evaluation stage submits jobs. use_slurm: "yes" attaches SLURM resources to its
heavy rules. Use the following CI arguments for SLURM setting.
CI arguments |
Defaults |
Descriptions |
|---|---|---|
|
no |
Fill in “yes” or “no” to use SLURM. |
|
250G |
Set the memory requirement for SLURM jobs (e.g., “32000M” or “32G”). |
|
‘’ |
Set the SLURM partition. Leave it empty (“”) to let the SLURM site default decide. |
|
‘’ |
Set the SLURM account. “” if your cluster does not use accounts. |
|
24h |
Set the wall time for each evaluation step (e.g., “4h”, “2d”; “” = partition default). |
The above settings are for all steps of nohic-eval. You can also set the resource requirements specifically for a step. Check the help message for details.
snakemake -s /path/to/noHiC-Snakemake/workflow/Snakefile --config help=yes
Linting and formatting
Linting results
1/home/runner/work/snakemake-workflow-catalog/snakemake-workflow-catalog/.pixi/envs/default/lib/python3.13/site-packages/google/auth/transport/grpc.py:43: FutureWarning: grpcio < 1.83.0 does not support Post-Quantum Cryptography (PQC). Support for non-PQC environments is deprecated. In April 2027, google-auth will raise its minimum requirements to enforce grpcio >= 1.83.0. For more details on Google Cloud's post-quantum security migration, visit: https://cloud.google.com/security/resources/post-quantum-cryptography
2 warnings.warn(
3Using workflow specific profile profiles/default for setting default command line arguments.
4noHiC: threads from --cores: 1 (CRAQ: 1)
5WorkflowError in file "/tmp/tmpn7s3sbc9/andyngh-noHiC-Snakemake-c487035/workflow/Snakefile", line 568:
6ERROR: the following keys are missing from the config file.
7Keys that are normally filled in by an earlier stage have to be given
8by hand when that stage is switched off in 'stages:'.
9 refpick: seq_file
10 refpick: prefix
11 refpick: gbz
12 refpick: hapl
13 refpolish: reads
14 refpolish: prefix
15 clean: contig_assembly
16 clean: adapters
17 clean: kraken2_db
18 asm: out_prefix
19 asm: reads
20 eval: reads
21 File "/tmp/tmpn7s3sbc9/andyngh-noHiC-Snakemake-c487035/workflow/Snakefile", line 568, in <module>
Formatting results
1[DEBUG]
2[DEBUG] In file "/tmp/tmpn7s3sbc9/andyngh-noHiC-Snakemake-c487035/workflow/rules/nohic-eval.slurm.c.smk": Formatted content is different from original
3[DEBUG]
4[DEBUG] In file "/tmp/tmpn7s3sbc9/andyngh-noHiC-Snakemake-c487035/workflow/rules/nohic-clean.c.smk": Formatted content is different from original
5[DEBUG]
6[DEBUG] In file "/tmp/tmpn7s3sbc9/andyngh-noHiC-Snakemake-c487035/workflow/Snakefile": Formatted content is different from original
7[DEBUG]
8[DEBUG] In file "/tmp/tmpn7s3sbc9/andyngh-noHiC-Snakemake-c487035/workflow/rules/nohic-refpick.c.smk": Formatted content is different from original
9[DEBUG]
10[DEBUG] In file "/tmp/tmpn7s3sbc9/andyngh-noHiC-Snakemake-c487035/workflow/rules/nohic-refpolish.c.smk": Formatted content is different from original
11[DEBUG]
12[DEBUG] In file "/tmp/tmpn7s3sbc9/andyngh-noHiC-Snakemake-c487035/workflow/rules/nohic-asm.c.smk": Formatted content is different from original
13[INFO] 6 file(s) would be changed 😬
14
15snakefmt version: 0.11.5