How to Plan Hi-C Sequencing for 20+ Samples: Batch Design, QC Gates, and Quote Inputs

Planning Hi-C for a large cohort isn't hard because the method is mysterious. It's hard because small upstream differences get amplified when you have 20, 40, or 60 samples moving through fixation, enzymatic steps, PCR, and multiple sequencing runs. If you don't make the key decisions early, you often discover the real constraints (sample quality, complexity, batch confounding, depth limits) only after you've spent the budget.
This guide is written as practical Hi-C project planning advice for PI-led teams. The goal is to help you design a multi-sample Hi-C study that is interpretable, reproducible, and quote-ready, without promising loop-level resolution from a dataset that can't support it.
Key Takeaway: For 20+ sample Hi-C projects, your biggest cost risks are (1) confounding biology with batch, and (2) sequencing deeper than your library quality and complexity can support. Design gates to catch both early.
Define the project goal before pricing: assembly scaffolding, 3D regulation, SV screening, or comparison
Teams often ask for a quote before they've decided what "success" means. That usually leads to mis-scoping. Hi-C can support very different outcomes, and each outcome changes what matters most: sample handling, depth, replicate strategy, and deliverables.
Below are four common goal archetypes. Pick the one that matches your primary decision.
1) Assembly scaffolding (chromosome-scale genome build)
If your primary goal is scaffolding a draft genome assembly to chromosome scale, your focus is:
- Signal content and long-range contacts (you need contact information that spans large genomic distances).
- Uniformity across the genome, rather than detecting subtle condition-specific differences.
- Deliverables that integrate into an assembly workflow (scaffolded FASTA, assembly QC summaries, contact map for validation).
This use case is often compatible with fewer samples, but it can still show up in "20+ sample" settings when teams scaffold multiple lines, strains, or accessions.
2) 3D regulation (compartments, domains, enhancer-promoter contacts, loops)
If your goal is to interpret regulatory wiring, decide which feature class you need to claim:
- Compartments (A/B): coarse, megabase-scale segregation.
- Domains / TADs: tens-of-kilobases to hundreds-of-kilobases scale.
- Loops / specific contacts: typically the most depth-hungry and the most fragile to low complexity.
A common planning mistake is to write "loops" into the goal statement by default. Many projects only need compartment/domain-level changes to support the biology. If you truly need loop calling, you should expect a steeper depth requirement and stricter QC expectations.
The reproducibility benchmarks in Yardımcı et al. (Genome Biology, 2019) show why: as coverage drops, the number of detectable mid-range significant interactions (loop-like signals) falls rapidly, while some higher-level structures (like domain boundaries) are more robust to noise and lower depth.
3) SV screening (large rearrangements, translocations, misjoins)
Hi-C can act as a genome-wide screen for major structural disruptions, especially when you care about long-range adjacency patterns. But it's not a breakpoint-precise assay by default.
For SV-focused projects, define up front:
- What SV class matters most (large translocations vs inversions vs complex rearrangements).
- Whether you need screening only, or screening plus orthogonal validation.
- Whether your sample type (e.g., tumor tissue) is expected to carry copy-number changes or karyotype abnormalities that complicate normalization.
This matters because between-sample normalization can behave differently in the presence of structural variation. Batch-correction methods like BNBC explicitly note limitations in settings with structural variation and copy-number changes (Fletez-Brant et al., Briefings in Bioinformatics, 2024).
4) Comparison across conditions (treatment vs control, time course, genotype, cohort)
This is the common "20+ sample" scenario: you want to compare contact maps across groups.
Your biggest risks here are:
- Batch confounding (one condition mostly processed in one batch).
- Unequal data quality across groups (one group has lower complexity or worse fixation).
- Overinterpreting differences that are really depth/QC artifacts.
If your primary goal is comparison, your design should prioritize:
- Balanced batching.
- Replicate logic (biological replication, not just technical repeats).
- A pre-registered QC acceptance framework (what gets resequenced, what gets rebuilt).
What information a service team needs before quoting
When you request a Hi-C quote for 20+ samples, most back-and-forth happens because key inputs are missing. The most quote-sensitive inputs are not "what organism" and "what platform." They are the variables that drive feasibility, expected success rate, and depth requirements.
Here is what a service team typically needs to scope accurately.
A. Sample metadata that changes feasibility
Provide these in a spreadsheet on day one:
- Sample type: cell line, primary tissue, organoid, sorted population, nuclei prep.
- Organism and genome size (or reference availability).
- Per-sample cell number (or nuclei count) and expected viability.
- Expected heterogeneity: mixed cell types, tumor purity issues, stressed cells.
- Biosafety and logistics constraints: shipping conditions, fix-first vs ship-live.
Why it matters: Hi-C is unforgiving about upstream material. Over-fixation, under-fixation, low viability, and degraded nuclei can all produce libraries that look "fine" on a Bioanalyzer but carry low proximity-ligation signal.
B. Project goal translated into analysis deliverables
"Hi-C analysis" is not a deliverable. You need to specify what outputs you plan to interpret.
Examples:
- Compartment calls and compartment switch analysis.
- Domain/TAD calling and boundary strength metrics.
- Loop calling (and what loop caller / resolution you expect).
- Differential contact analysis (what statistical framework and what covariates).
- For scaffolding: contact-guided scaffolding and assembly QC.
If you don't define deliverables, you can't define depth, and you can't set QC gates that match the downstream promise.
C. The batch plan you're willing to accept
For 20+ samples, it is normal to need multiple wet-lab batches and multiple sequencing runs. The quote should reflect how you want those batches handled.
Clarify:
- Will samples be processed over multiple days/operators?
- Are there constraints that force one condition into one batch?
- Are you willing to include a reference control sample repeated across batches?
If you don't specify this, you may get an efficient but confounded schedule.
D. Your decision rule for pilot vs full commitment
A quote can be structured as:
- Pilot-only.
- Pilot + optional scale-up.
- Full cohort with internal pilot QC.
If you want cost control, ask for a pilot-first structure explicitly.
Batch design for multi-sample Hi-C projects
This is the core of Hi-C batch design for large cohorts: don't let technical scheduling stand in for biology.
Batch design is where multi-sample Hi-C studies succeed or fail. It's not glamorous, but it is one of the few things you can control before any data exist.
The design goal: don't let batch stand in for biology
If "all treated samples were processed in week 1" and "all controls were processed in week 3," you have built a confounded study. No downstream normalization can fully recover biological inference from a confounded design.
This isn't just a generic omics warning. Hi-C has a distance-dependent structure (contact probability decays with genomic distance), and between-sample unwanted variation can also be distance-dependent. The BNBC framework formalizes that point by correcting batches within distance bands in Removing unwanted variation between samples in Hi-C experiments (Briefings in Bioinformatics, 2024). You shouldn't treat batch as a simple additive nuisance you can always subtract later.
Practical batching rules for 20+ samples
1) Balance groups across every batch
Each prep batch should contain samples from every major group whenever possible.
- If you have 4 conditions and 3 batches, every batch should contain all 4 conditions.
- If one group has rare or fragile samples, distribute those across batches rather than clustering them.
2) Randomize within constraints
If you have to process by shipment arrival, don't give up on randomization. Randomize within what you control:
- Order within a day.
- Lane allocation on the sequencer.
- Library pooling.
3) Use a shared reference sample if comparison is central
For large projects, consider including one "reference" sample (or pooled reference) in multiple batches. This gives you an internal anchor for QC consistency.
The reference doesn't eliminate batch effects, but it does make them visible.
4) Track batch metadata like a first-class dataset
Record batch variables that become covariates:
- Operator, date, protocol version.
- Reagent lots.
- Fixation time/concentration.
- Enzyme lot, incubation times.
- PCR cycles.
- Sequencing run / flow cell.
If you can't reconstruct batch metadata later, you can't model it.
⚠️ Warning: A "clean" design is one where every biological group appears in every batch. If that's not true, be explicit that your downstream comparison will be limited.
Replicates: what matters in large cohorts
In multi-sample studies, teams sometimes confuse "more samples" with "more replicates." If you have 20 samples but each condition has only one sample, you do not have replication.
A service quote doesn't need your full statistical plan, but it does need your replication logic:
- Biological replicates per condition.
- Whether samples are paired (e.g., matched tumor/normal).
- Whether the comparison is within subject (time course) or between subjects.
QC gates before library construction: viability, fixation, nuclei integrity, digestion, ligation, library complexity
A multi-sample Hi-C project needs QC gates that are consistent across the cohort. If you "push through" marginal samples because you need to hit N=24, you usually pay later: higher duplicates, lower signal, poorer comparability.
The QC gates below are framed around two principles:
- Hi-C QC is about signal content and complexity, not just DNA quantity.
- You want to fail early, not after deep sequencing.
ENCODE publishes a standards-style QC checklist for Hi-C that includes explicit thresholds for duplicates, ligation motif frequency, and long-range contact fraction in its HiC Data Standards and Processing Pipeline (updated September 2022). While not every protocol must match ENCODE's specific numbers, the metrics are useful because they directly reflect the success of digestion, fill-in, and ligation.
Gate 1: Viability and cell state (before fixation)
For many sample types, viability is not a bureaucratic checkbox. It's a proxy for whether you are fixing intact nuclei or a stressed/degrading population.
For primary samples, decide upfront:
- Can you ship live, or do you need to fix on site?
- If you fix on site, can fixation be standardized across collection locations?
If you cannot standardize, your batch plan must explicitly model that as a source of variation.
Gate 2: Fixation suitability (under- vs over-fixation)
Fixation is where many Hi-C libraries become irrecoverable. Under-fixation can lose proximity information; over-fixation can reduce enzyme accessibility and ligation efficiency.
For 20+ samples, the main planning point is not the exact fixation protocol. It's this: you need one protocol and one acceptance logic across the cohort, or you'll create systematic quality differences that look like biology.
Gate 3: Nuclei integrity
Nuclei integrity is the bridge between "we have cells" and "enzymes can do their job." Broken nuclei and degraded chromatin can produce noisy libraries.
For tissue-derived samples, nuclei prep can become the dominant risk variable. If your cohort includes multiple tissue types, you may need separate nuclei prep SOPs.
Gate 4: Digestion check (is restriction working?)
Digestion affects fragment distribution, junction availability, and downstream complexity. In practice, digestion failures show up later as low valid contacts and low long-range fraction.
Your QC plan should include:
- A digestion efficiency check appropriate to your protocol.
- A rule for when a digestion is considered failed and the sample is reprocessed.
Gate 5: Ligation success (are proximity junctions being formed?)
One of the most actionable sequencing-derived QC metrics is the fraction of reads containing the ligation motif, because it reflects whether ligation worked.
ENCODE's in situ Hi-C standards include "minimum reads with ligation motif present: 15%" and note that low values can indicate ligation failure (ENCODE Hi-C Data Standards, updated 2022).
Even if you don't treat that as a universal threshold, the planning implication is clear: if your ligation motif rate is poor, deeper sequencing just buys you more low-signal reads.
Gate 6: Library complexity (will deeper sequencing still help?)
Low complexity is one of the most expensive failure modes. You can sequence a low-complexity library to billions of reads and still not get more unique contacts.
ENCODE includes a "maximum unique total duplicates: 40%" as a standard-style threshold for in situ Hi-C (updated 2022). High duplicate burden is often a sign of limited unique molecules and too much PCR.
To keep the language consistent across your internal documentation, treat these as Hi-C QC metrics rather than "sequencing stats": they're telling you whether the proximity-ligation chemistry worked, and whether the library can still yield new information if you sequence deeper.
Gate 7: Pilot sequencing QC before committing to full depth
If you want a practical gate that saves budget, it is this: pilot sequence first, then decide.
Two sources support this pilot-first logic:
- qc3C was developed specifically to estimate Hi-C signal content based on ligation artefacts, including reference-free estimation, to support early QC decisions in qc3C: Reference-free quality control for Hi-C sequencing data (PLOS Computational Biology, 2021).
- Yardımcı et al. benchmarked reproducibility/quality metrics and show that some methods can distinguish biological replicates at around 5 million interactions, which supports the idea that shallow data can still diagnose "is this working?" before you spend on depth in Measuring the reproducibility and quality of Hi-C data (Genome Biology, 2019).
ENCODE also demonstrates that meaningful QC metrics can be defined at the level of unique Hi-C contacts, long-range fractions, and ligation motifs. Together, these support a simple planning rule:
- If pilot QC indicates low signal content or very high duplicates, fix the wet lab first.
- If pilot QC is strong, then depth becomes a rational lever rather than a gamble.
Plan Hi-C sequencing depth and deliverables by project type
Depth planning is where teams want a single number. In reality, depth is a function of: goal, genome size, restriction strategy, library complexity, and what you call a "deliverable."
In this section I'll use the phrase Hi-C sequencing depth in the practical sense: how many usable contacts you can expect after filtering, and whether that is enough to support the feature class you plan to interpret.
A practical hierarchy: what's depth-hungry?
At a high level:
- Compartments are usually the least demanding.
- Domains/TADs are intermediate.
- Loops and specific mid-range contacts are typically the most demanding.
Yardımcı et al. provide empirical backing for this planning mindset: loop-like significant interactions are depleted as coverage drops, while some domain boundary signals can remain stable until coverage becomes very low (Genome Biology, 2019).
Deliverables you can define upfront (and tie to depth)
For multi-sample projects, define deliverables in three layers:
Layer 1: Raw + processed basics (should be standard)
- FASTQ and alignment summaries.
- Contact matrices at multiple resolutions.
- A QC report that includes duplicates, valid contact fraction, ligation motif rate, and long-range contact fraction.
ENCODE's QC metric set is a helpful reference for what belongs in that QC report (ENCODE Hi-C Data Standards, updated 2022).
Layer 2: Feature calling (depends on goal and depth)
- Compartments (A/B) and compartment switching.
- Domain/TAD calls and boundary scores.
- Loop calling (if depth supports it) and loop reproducibility.
Layer 3: Comparative statistics (depends on design)
- Differential contact analysis.
- Batch-aware normalization and covariate modeling.
- Cross-sample reproducibility benchmarking.
If your cohort includes strong batch structure, it is worth defining in advance whether batch will be treated as:
- a design factor you avoid through balancing, or
- a modeled covariate in downstream analysis.
BNBC is one example of how between-sample correction can be structured in a distance-aware way (Fletez-Brant et al., Briefings in Bioinformatics, 2024). You don't need to commit to BNBC specifically, but you should acknowledge the need for distance-aware thinking.
What to avoid promising
Avoid promising:
- "Loop calls for every sample" without defining depth, complexity expectations, and replicate logic.
- "kb-resolution maps" without specifying what "resolution" means (bin size vs effective resolution).
- "Cross-sample comparisons" when the batch plan makes condition and batch inseparable.
Pilot-first design versus full cohort sequencing
A pilot-first design is a budgeting tool and a credibility tool. It forces you to test reality before committing to the full spend.
What a pilot should answer (and what it shouldn't)
A pilot is not meant to prove your biological hypothesis. It is meant to validate:
- Sample-handling feasibility.
- Library quality across your sample types.
- Expected signal content and complexity.
- Reproducibility trends at low depth.
Yardımcı et al. show that shallow sequencing can support replicate classification for several reproducibility measures at around 5 million interactions (Genome Biology, 2019). qc3C reinforces that pilot-scale sequencing can be enough to estimate signal content to guide deeper sequencing decisions (PLOS Computational Biology, 2021).
A realistic pilot structure for 20+ samples
One practical structure:
-
Choose 4–8 representative samples.
- Include the most difficult sample type (e.g., primary tissue) and the easiest (e.g., cell line).
- Include at least one sample from each condition.
-
Process them in the same workflow you plan to scale.
-
Pilot sequence them shallowly.
-
Apply a consistent QC acceptance rubric.
If the pilot shows high variability, you've learned something important: your cohort will not behave like a tidy cell-line study. You can then decide whether to invest in protocol optimization, change the goal (e.g., domain-level instead of loops), or revise sample inclusion criteria.
Scaling logic: when to spend on depth vs more samples
In practice, there are three scaling moves:
- Increase depth per sample when your libraries are high quality and your goal requires finer-scale features.
- Increase replicate count when your between-sample variability is high and your comparison is the primary goal.
- Tighten sample inclusion criteria when your failure rate is high and poor samples would dilute the cohort.
What you should not do is mix all three without a gate. That's how large Hi-C projects quietly double in cost.
Quote-ready checklist for researchers
If you're trying to reduce back-and-forth, treat this as a Hi-C quote checklist: copy it into your email or project ticket. The items are phrased as binary inputs whenever possible.
1) Project definition
- Primary goal selected: scaffolding / 3D regulation / SV screening / condition comparison
- Secondary goals explicitly labeled as "nice-to-have"
- Feature deliverables defined (compartments, TADs, loops, differential contacts)
2) Sample list and metadata
- Sample table includes: ID, condition/group, tissue/cell type, organism, genome version (if known)
- Per-sample input: cell number or nuclei count
- Viability estimate available (or explicit statement if not applicable)
- Shipping constraints defined (live vs fixed; temperature; time window)
3) Batch design constraints
- Number of expected wet-lab batches declared
- Design rule agreed: each batch contains all conditions (or exceptions justified)
- Operator/protocol metadata will be recorded and delivered
- Reference sample plan defined (yes/no)
4) QC gates (pass/fail logic)
- Predefined rules for: re-fix / re-prep / re-library / resequence
- Pilot sequencing step included (yes/no)
- QC metrics required in report include ligation motif rate, duplicates/complexity, and long-range contact fraction (ENCODE-style metrics are acceptable references)
5) Sequencing and data delivery
- Sequencing mode specified (paired-end length, platform preference if any)
- Target depth logic agreed: "by deliverable" rather than a single number
- Deliverables list includes: FASTQ, processed matrices at stated resolutions, QC report, and analysis outputs tied to goals
- Data transfer method and storage expectations agreed
6) Analysis expectations
- Analysis scope includes batch-aware normalization/covariates if comparison is the goal
- Reproducibility reporting expected (replicate correlations / reproducibility metrics)
- Visualization expectations defined (contact maps, compartment tracks, domain/loop tracks)
Next steps
If you want a service team to quote efficiently, send the checklist above plus your sample sheet in the first email. If you want to reduce risk further, ask for a pilot-first structure with explicit QC gates.
For CD Genomics' 3D genomics service overview and study-design support, see the CD Genomics 3D Genomics hub.
Author
Dr. Yang H. is a Senior Scientist at CD Genomics. He supports study design, QC strategy, and deliverables planning for multi-sample 3D genomics projects (including Hi-C) across diverse sample types and cohort sizes. LinkedIn: Dr. Yang H.
