Hi-C, PacBio HiFi, and ONT for Hybrid Genome Assembly: Planning Chromosome-Scale Animal Genome Projects

Workflow overview of hybrid genome assembly using PacBio HiFi, Oxford Nanopore ultra-long reads, and Hi-C scaffolding.

Genome assemblies usually fail for boring reasons.

Not because you picked the wrong assembler brand name, or because you missed the newest chemistry. They fail because the project was planned as if "more sequencing" automatically means "more truth." In chromosome-scale assembly, different datasets reduce different kinds of uncertainty. If you mix them without a model, you can spend a lot of money and still ship an assembly that is structurally wrong in the places that matter.

This guide is written for PI-led teams and senior bioinformaticians planning animal or primate genomes in a consideration-stage mindset: you already know what long reads and Hi-C are, but you want a defensible way to decide which combination you need, and what a quote should commit to.

Key takeaways

A few planning statements that hold up in practice:

  • PacBio HiFi and ONT ultra-long reads solve different contig-level problems. HiFi primarily buys consensus accuracy and cleaner assembly graphs; ONT ultra-long reads primarily buy repeat spanning and gap closure.
  • Hi-C solves a different class of problem than long reads. Hi-C constrains chromosome-scale order/orientation and is a powerful misjoin detector. It does not "fix" repeat collapses or base errors.
  • For multi-genome cohorts, you need a consistency strategy, not just a deeper reference. Cohort scaffolding can be efficient, but it breaks when real structural variation is present.
  • Your quote should specify acceptance criteria and curation deliverables. Ask for FASTQ + phased assemblies (when required) + contact maps + BUSCO/QV/N50 + an annotation-ready package, plus a documented misjoin-correction loop.

Why long reads and Hi-C solve different assembly problems

A chromosome-scale assembly has at least three separate objectives. If you don't separate them, procurement conversations get fuzzy and you end up paying twice.

Objective 1: base-level correctness

Base correctness is about the consensus sequence: substitutions and small indels. High base accuracy reduces downstream polishing effort, lowers false positive variant calls, and makes gene model prediction less brittle.

HiFi is built to attack this problem. That's why HiFi-first assembly has become a common baseline for diploid animal genomes (for example, the hifiasm assembly graph approach described by Cheng and colleagues in Nature Methods, 2021).

Objective 2: repeat resolution and long-range spanning

Repeat resolution is about whether the assembly graph can be simplified into correct contigs without collapsing or duplicating repeat structures.

A practical mental model: repeats aren't "hard" because they are repetitive. They're hard because the ambiguous region is longer than your reads, so the assembler cannot uniquely choose a path.

That is where ONT ultra-long reads can change the outcome. Ultra-long molecules can bridge repeat structures that exceed typical HiFi read lengths, and that bridging power is why hybrid pipelines exist for telomere-to-telomere-oriented work. Verkko is a canonical example: it is designed to use HiFi plus ONT ultra-long reads for diploid telomere-to-telomere assembly (see Verkko, Nature Methods 2023).

Objective 3: chromosome-scale structure (ordering, orientation, and chimeric joins)

Even with excellent contigs, you can still be wrong at the chromosome scale.

Hi-C contact data provides long-range linkage information that helps you:

  • order and orient contigs into chromosome-scale scaffolds, and
  • detect misjoins when contact patterns break expectations.

A key point for planning: Hi-C reduces uncertainty about structure, not about bases.

The 3D-DNA pipeline demonstrated Hi-C-based chromosome-length scaffolding in a Science paper on Aedes aegypti (Dudchenko et al., Science 2017). For manual curation, Juicebox remains a standard interactive viewer for inspecting contact maps (Durand et al., Cell Systems 2016).

When PacBio HiFi, ONT ultra-long reads, and Hi-C should be combined

Most teams overbuy data in one of two ways: they buy a little bit of everything without enough of any one thing to fix the actual bottleneck, or they buy huge depth of one dataset when the real failure mode is orthogonal.

To decide, start with a hard question: what do you mean by "reference quality" for this study?

Define "done" before you define coverage

In genome assembly, "reference" can mean any of the following, and they are not interchangeable:

  • Chromosome-scale draft: ordered/oriented scaffolds with high continuity, but gaps remain in extreme repeats.
  • Haplotype-resolved assembly: two haplotypes (or primary + alternate) rather than a collapsed mosaic.
  • Near-T2T / T2T-like: explicit ambition to close telomeres/centromeres and long satellites.
  • Cohort-ready set: consistent deliverables and QC across many individuals, even if none is T2T.

Once "done" is explicit, dataset choices become defensible.

A decision matrix you can actually use

Below is a compact comparison based on what each data type constrains in the inference problem.

Decision axis PacBio HiFi ONT ultra-long reads Hi-C
Primary value base accuracy + clean graphs spanning across long repeats long-range ordering/orientation + misjoin detection
Typical "best use" contig assembly backbone bridge repeats; improve contiguity; gap closure scaffolding + curation; chromosome-scale sanity checks
Where it does not help doesn't guarantee chromosome placement does not automatically give high consensus accuracy does not fix base errors or collapsed repeats
Risk if underplanned overpolishing and hidden errors expensive prep failures; insufficient UL length "chromosome-scale" but structurally wrong scaffolds

A decision tree (quick triage)

If you want an answer in five minutes, use this.

  1. Is your primary goal simply a high-accuracy, contig-level assembly for annotation and variant calling?
  • If yes, start with HiFi-first assembly. You will often get long contigs and strong base accuracy with relatively little polishing.
  1. Do you expect long repeats that exceed typical HiFi read lengths (satellite arrays, long SDs, rDNA blocks), and do you care about closing them?
  • If yes, add ONT ultra-long reads and plan the extraction as a milestone. Ultra-long data is most valuable when it spans the ambiguity the graph cannot resolve.
  1. Do you need chromosome-scale ordering and an explicit misjoin correction step?
  • If yes (most publication-facing projects), add Hi-C and require contact-map review deliverables.
  1. Do you need a phased output (two haplotypes) or is a collapsed reference acceptable?
  • If you need phasing, specify it explicitly. A chromosome-scale scaffold can still be a haplotype mosaic if phasing is not treated as a deliverable.

The reason this matters: it prevents teams from buying "chromosome-scale genome assembly" as a slogan, rather than planning for the specific uncertainty that breaks their use case.

When "all three" is justified

Combining HiFi + ONT ultra-long + Hi-C is usually justified in three scenarios.

1) You expect extreme repeats and you care about closing them

If the goal is near-T2T, the project bottleneck is often repeat spanning, not consensus.

Verkko shows why hybrid inputs matter: in its diploid example, combining HiFi with ultra-long reads enabled telomere-to-telomere assemblies for a subset of chromosomes and achieved very high consensus accuracy (as reported in the Nature Methods paper; see Verkko, Nature Methods 2023).

Hi-C is then used as a structural constraint and confirmation signal. It won't close centromeres by itself, but it can tell you whether chromosome-scale joins are plausible.

2) You need haplotype separation, not just a long scaffold

Phasing is a deliverable. Don't outsource it to luck.

Hifiasm is explicitly designed for haplotype-resolved assembly using phased assembly graphs (Cheng et al., Nature Methods 2021). If haplotypes matter for your downstream biology (allele-specific regulation, population comparisons, avoiding reference collapse), plan for phased outputs.

Hi-C can help with chromosome-scale phasing signals in some workflows, but contig-level haplotype separation still depends heavily on long-read information.

3) Misjoins would poison downstream interpretation

If you intend to publish a "reference," interpret structural variation, or use the assembly to anchor functional genomics, structural correctness is not optional.

⚠️ Warning: A chimeric join can look fine by N50. It still breaks synteny, fragments genes, and creates false SV calls. Hi-C contact maps are one of the fastest ways to catch these problems before annotation.

Planning 20 animal or primate genomes: one reference, multiple assemblies, or cohort scaffolding

Scaling from one assembly to 20 changes what "good planning" looks like. You now care about repeatability, comparability, and a failure policy.

Decide what must be consistent across the cohort

Before you choose Strategy A/B/C, define what "consistent" means for your cohort. In most primate/animal projects, it is at least four things:

  • Consistent naming and coordinate conventions (chromosome names, haplotype labels, mitochondrial handling).
  • Consistent QC reporting (the same metrics and plots for every sample).
  • Consistent filtering policy (what is labeled contamination vs retained as unplaced).
  • Consistent failure policy (when you stop, when you resequence, when you accept a reduced scope).

Without these, even technically good assemblies become hard to compare and easy to misinterpret.

Common cohort pitfalls (that aren't fixed by more coverage)

A few issues that show up repeatedly when projects scale:

  • Sex chromosome imbalance: mixed sexes in the cohort changes coverage expectations and complicates "one pipeline for all."
  • Mitochondrial overrepresentation: high mtDNA fraction can distort coverage assumptions and polishing behavior.
  • Sample-to-sample DNA quality variance: the cohort's bottleneck is often the worst samples, not the average.
  • Real structural heterogeneity: inversions and translocations can look like assembly errors if you force every sample into the same scaffold structure.

These are planning problems. They need explicit rules, not hero debugging at the end.

Option 1: one deep reference, lighter assemblies for the rest

This is common when you have one high-quality individual (good tissue, strong HMW DNA yield, clean Hi-C) and many cohort samples where a full T2T ambition is unrealistic.

It works well when the cohort is genetically close and the primary deliverable is "a coordinate system plus per-sample variants."

Where it can fail:

  • the cohort contains large inversions/translocations relative to the reference,
  • sex chromosomes vary (particularly in primates where sex chromosome assembly is already difficult),
  • you have substantial repeat architecture differences,
  • or contamination and mixed tissue sources vary across samples.

A safeguard that helps: define a "graduation rule." If a sample's assembly QC deviates beyond a threshold (e.g., unusually low Merqury QV or suspicious contact-map patterns), that sample gets a deeper de novo assembly rather than being forced into reference-guided scaffolding.

Option 2: full assemblies for each individual (highest cost, lowest assumption)

If you expect real structural heterogeneity, this is the conservative option.

It's also operationally clean: each genome is judged against the same deliverable list and QC gates, instead of being judged by similarity to a reference. That matters when multiple people will use the dataset later.

Option 3: cohort scaffolding (efficient, but assumption-heavy)

Cohort scaffolding or reference-guided placement can reduce costs, but it requires discipline. It tends to break in exactly the cases where people care most about cohorts: structural variation, lineage divergence, and repeats.

If you choose this option, treat it as a hypothesis: you must still validate per-sample structure, ideally with either contact map evidence or other long-range signals. Otherwise you risk producing a set of "consistent" assemblies that are consistently wrong.

Sample requirements for HMW DNA and Hi-C libraries

Hybrid assembly success is often determined before sequencing begins. Sample realism should be in the quote, not in the postmortem.

A compact acceptance checklist (what to collect and how to decide)

When teams send "good DNA" without a shared definition, quotes become guesswork. What you want is a minimal acceptance checklist that is easy to satisfy and hard to misinterpret.

  • For HiFi DNA inputs: purity ratios, dsDNA mass by fluorometry, and an integrity readout that shows whether the sample is dominated by long fragments.
  • For ONT ultra-long inputs: the same purity/quantity checks, plus a handling note that documents how shear was minimized (pipetting, mixing, transport, freeze-thaw).
  • For Hi-C: tissue type, preservation method, time-to-fixation or stabilization, and any expected contamination sources.

If any of these items are missing, the risk is not "slightly lower N50." The risk is that you can't interpret why the project failed, so you can't fix it efficiently.

HMW DNA for PacBio HiFi: controlled length, high purity, predictable behavior

HiFi assembly typically performs best when the DNA behaves predictably during library construction. You don't need the absolute longest molecules for HiFi; you need intact DNA that can be size-selected or sheared into the target insert range without chemical inhibition.

In practice, teams should treat these as acceptance categories:

  • Purity: A260/280 around 1.8 and A260/230 near ~2.0 are common indicators that inhibitors are low.
  • Integrity: a size distribution dominated by high molecular weight DNA rather than a broad smear.
  • Quantification: fluorometric dsDNA quantification (UV alone is not sufficient for viscous HMW DNA).
  • Handling log: how the DNA was mixed, aliquoted, shipped, and stored (freeze-thaw history).

A small planning improvement that pays off: ask for a one-page sample QC report before the sequencing run is committed. If the vendor can't show you how they screen inhibitors and degradation, you can't estimate the redo risk.

UHMW DNA for ONT ultra-long: minimize shear and plan for variability

Ultra-long sequencing is sensitive to mechanical handling. In quotes, this often shows up as vague promises ("ultra-long reads") with no operational definition.

For planning, you want the vendor to define:

  • how they define "ultra-long" (read N50? maximum read length? a target fraction of reads >100 kb?),
  • how they will assess success (run metrics reported back to you),
  • and what happens if the extracted DNA cannot support ultra-long libraries.

Hi-C libraries: tissue handling is the real constraint

Hi-C is fundamentally different because it relies on capturing native chromatin contacts. It is not a DNA-only problem.

For animal tissues, the main determinants are:

  • how fast tissue is stabilized relative to collection,
  • whether there are repeated freeze-thaw cycles,
  • and whether the tissue is contamination-prone (which affects mapping).

If your group is using Hi-C primarily as an assembly tool, it still helps to align expectations about contact-map QC and deliverables. CD Genomics summarizes workflow-level QC expectations in its "QC metrics for 3D genomics workflows" resource page.

Assembly workflow: contig assembly, polishing, Hi-C scaffolding, phasing, and misjoin correction

The most common planning mistake is treating assembly as a linear pipeline. For chromosome-scale work, it's a loop with checkpoints.

Step 1: contig assembly (choose your backbone)

For many diploid animal genomes, a HiFi-first contig assembly is the base case. Hifiasm is a common choice for haplotype-aware assembly with HiFi reads (see Cheng et al., Nature Methods 2021).

If you are explicitly targeting repeat closure and extreme contiguity, a hybrid approach designed for HiFi plus ultra-long reads becomes more attractive; Verkko is designed for this input regime (see Verkko, Nature Methods 2023).

Step 2: polishing (be explicit about what is allowed to change bases)

Polishing is where some projects unknowingly trade one error type for another. If you use short reads, you need safeguards against mis-mapping in repeats.

Your quote should specify:

  • which data types are used for polishing (HiFi-only polishing differs from ONT-based polishing),
  • the mapper and variant caller, and
  • how changes are validated (e.g., k-mer based evaluation).

Merqury is widely used for reference-free consensus quality (QV) and completeness using k-mers (Rhie et al., Genome Biology 2020).

Step 3: Hi-C scaffolding (ordering and orientation)

Hi-C scaffolding tools automate ordering/orientation with different assumptions.

Representative peer-reviewed options include 3D-DNA (Dudchenko et al., Science 2017), SALSA2 (Ghurye et al., PLOS Comput Biol 2019), and YaHS (Zhou et al., Bioinformatics 2023).

Planning point: scaffolding parameters matter less than whether your pipeline includes a deliberate correction step.

Step 4: misjoin correction (manual curation is not a luxury)

Misjoin correction is where chromosome-scale assemblies become trustworthy.

The usual workflow is: generate contact maps, inspect for discontinuities and off-diagonal patterns, split suspicious joins, then rerun scaffolding if needed. Juicebox was introduced as a viewer for interactive inspection of Hi-C contact maps (Durand et al., Cell Systems 2016).

Pro Tip: Require an auditable curation deliverable: a list of manual breaks (coordinates) and at least one pre/post contact map snapshot.

Step 5: phasing (define what you will receive)

"Phased" can mean several different deliverables:

  • primary + alternate assemblies,
  • partially phased contigs,
  • or chromosome-scale phasing (often requiring additional information).

If your downstream analysis depends on haplotypes, define the phasing target and the expected outputs in writing.

Step 6: QC, release, and annotation readiness

At minimum, release criteria should include:

  • contiguity metrics (e.g., N50, NG50 with stated genome size assumption),
  • gene completeness metrics such as BUSCO (Simão et al., Bioinformatics 2015),
  • base-level consensus quality via QV and completeness via k-mers (e.g., Merqury QV),
  • and structural validation artifacts (Hi-C contact maps).

Deliverables: FASTQ, assembly FASTA, Hi-C contact maps, BUSCO/QV/N50, annotation-ready files

Deliverables are where misunderstandings become expensive. You want a package that a different lab member can pick up six months later and still reproduce what happened.

A simple way to sanity-check a deliverable set

If a vendor delivers "chromosome-scale" results but cannot give you:

  • raw FASTQs,
  • an assembly FASTA that matches the reported scaffolds,
  • a contact map you can inspect,
  • and a QC report with clearly defined metrics,

then you do not actually have a reviewable assembly. You have an interpretation.

In MOFU procurement terms: you want artifacts that allow an independent re-check, not a slide deck.

Raw reads and run metadata

Ask for raw FASTQ files for each dataset (HiFi, ONT if included, Hi-C), plus minimal run metadata (platform and chemistry; basecaller version for ONT; library prep identifiers).

Assembly outputs

Ask for:

  • primary assembly FASTA,
  • alternate/haplotype FASTA (if phased outputs are in scope),
  • assembly graph artifacts when available (useful for troubleshooting and interpretation),
  • and a clear naming convention (contigs, scaffolds, haplotypes).

Hi-C scaffolding and curation artifacts

If Hi-C is used for scaffolding, require:

  • at least one standard contact map format suitable for review,
  • a brief parameter summary (scaffolder + key settings),
  • and a curation log if manual edits were performed.

If you want a 3D-genome deliverables reference point, the CD Genomics QC metrics for 3D genomics workflows page summarizes common QC expectations and deliverables.

QC report: completeness, correctness, and structural sanity

A QC report should include, at minimum:

  • N50/NG50 (with assumptions),
  • BUSCO completeness (Simão et al., Bioinformatics 2015),
  • a reference-free QV and completeness assessment using k-mers (for example, Rhie et al., Genome Biology 2020),
  • contamination screening summary,
  • contact-map diagnostics and any identified misjoins.

Annotation-ready package (make the definition explicit)

"Annotation-ready" can be vague. In quotes, define it as a concrete set of files and conventions, for example:

  • a cleaned primary assembly FASTA with consistent scaffold naming,
  • separate mitochondrial/organellar sequences or clearly labeled contigs,
  • a file describing contaminant filtering (what was removed and why),
  • repeat masking plan or deliverables (depending on your annotation strategy).

Quote checklist for hybrid genome assembly projects

A quote should behave like a spec sheet. If it reads like marketing, it is not protecting you.

Below is a checklist you can use as an email template. Keep it short enough that vendors will answer it, but specific enough that you can hold them to it.

1) Project definition

State species, estimated genome size, expected ploidy, and whether you require haplotype-resolved outputs. Define "done" as one of: chromosome-scale draft, haplotype-resolved reference, or near-T2T.

2) Data plan and purpose

Require the vendor to state what each dataset is doing in the workflow. "HiFi for consensus; ONT UL for repeat spanning; Hi-C for scaffolding and misjoin detection" is a reasonable default, but the point is that they must commit to a rationale.

3) Sample acceptance criteria and failure policy

Ask for explicit QC gates for:

  • HiFi DNA input (mass, purity ratios, integrity expectations),
  • ONT ultra-long DNA input (what counts as ultra-long, and what happens if it fails),
  • Hi-C tissue handling and fixation constraints.

Also ask what the project does when samples fail QC: stop-and-consult, reduced scope, re-extraction, or additional sequencing.

4) Assembly and curation workflow

Require:

  • assembler name/version,
  • scaffolder name/version,
  • and a misjoin-correction loop with contact-map review.

If they say "chromosome-scale," they should be able to name a viewer and show how manual corrections are documented.

5) QC metrics and release criteria

Require a written list of metrics, including BUSCO and a k-mer based QV (e.g., Merqury QV). Define minimum acceptable ranges if your project has them, or at least require that the vendor propose them.

6) Deliverables, reproducibility, and ownership

Ask for:

  • all raw FASTQs,
  • all final FASTAs (primary and alternate where in scope),
  • key intermediate artifacts (contact maps; scaffolding logs; curation notes),
  • and enough parameter context to reproduce the pipeline.

Also clarify data ownership, IP, and retention.

Key Takeaway: If a quote doesn't specify deliverables and a misjoin-correction loop, you aren't buying a chromosome-scale assembly. You're buying a hope.

Next steps (RUO)

All services and outputs discussed here are for research use only.

If you want to align your sequencing plan with a 3D-genome-informed scaffolding and QC deliverable set, start by writing down three items: estimated genome size, sample constraints (tissue type, shipping, DNA yield), and your definition of "done." Then map those to a data plan.

For teams that want to integrate long-range contact data into assembly planning, the CD Genomics pages on Hi-C sequencing and Pore-C sequencing describe contact-map deliverables and workflow options.


Author

Dr. Yang H. is a Senior Scientist at CD Genomics, working on 3D genome methods (Hi-C and long-read contact mapping) and long-read sequencing strategy for genome assembly planning.

Connect on LinkedIn: Dr. Yang H. on LinkedIn

For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.
Leading Your Research Forward

Enhancing your vision research capabilities.

High-confidence 3D genomics services for chromatin interaction analysis and regulatory insight.

Contact Us
Copyright © CD Genomics. All Rights Reserved.
Top