High-Resolution HLA Typing by NGS: A Comprehensive Guide to Class I and Class II Allele Identification for Research and Translational Applications

Meta Intent: A practical guide for defining the HLA resolution a study actually needs, selecting an NGS design that can support it, and interpreting allele calls without overstating what the sequence evidence proves.

High-resolution HLA typing is often requested as a single deliverable, but it is better understood as an evidence standard. A result is only as resolved as the loci, sequence regions, molecule-level phase information, reference database, and reporting policy behind it. This guide explains how to design and audit NGS-based HLA typing for research questions that require defensible Class I and Class II allele calls.

High Resolution Is an Evidence Claim, Not a Synonym for NGS

The first planning error in an HLA project is to treat "NGS-based" and "high-resolution" as interchangeable. NGS can reduce ambiguity substantially, but the result still depends on what was amplified or captured, which bases were covered, whether variants can be phased on the same chromosome, and how the analysis software maps reads to its reference set. A short-read assay that reports two fields and a full-gene long-read assay can both be described as NGS, yet they do not make the same claim.

For project design, define the output before the method. Is the study asking for a broad antigen-associated category, a two-field allele call, a sequence distinction within a coding region, or a full-gene allele context that includes intronic and untranslated variation? The answer determines whether a targeted design, a long-amplicon design, or an inference workflow from existing genomic data is appropriate.

The practical rule is simple: the report should state the resolution that the available sequence evidence supports, not the largest number of fields available in a database. This shifts the discussion from "Which platform is best?" to "What evidence must the experiment preserve?"

HLA gene structure with an evidence ladder showing field-level resolution and the boundary created by unphased distant variants.Figure 1: HLA resolution ladder and evidence boundaries.

Where HLA Ambiguity Comes From

Two technical problems create most residual HLA ambiguity. The first is incomplete observation. If an assay interrogates only selected exons, two alleles that differ in an uncovered intron, exon, or UTR cannot be distinguished by that assay. The second is missing phase. A heterozygous sample may contain two variants, but separate short reads may not show which variant belongs with which allele on the same chromosome.

These mechanisms matter because they require different corrections. An unobserved region calls for broader locus coverage. A cis-trans uncertainty calls for a design that links informative variants on the same molecule, such as a sufficiently long amplicon, a molecule-aware library, or longer reads. Adding more reads to the same disconnected fragments can improve confidence in individual bases without resolving a distant phase relationship.

Class I and Class II loci also do not behave as one uniform target. Many traditional strategies focus on exons 2 and 3 for Class I and exon 2 for Class II because they encode key peptide-binding regions. That is useful biology, but it does not make the remainder of each gene irrelevant to allele-level reporting. A project must distinguish a useful peptide-binding-domain result from a claim of complete allelic definition.

The HLA Typing service discussion should therefore start with the unresolved question: incomplete coverage, unphased polymorphisms, potential novel variation, or a need for a specific reportable group. Treating every ambiguity as an analytical software issue often leads to an unhelpful re-analysis of data that never contained the necessary evidence.

Split diagram distinguishing an unsequenced HLA interval from a phase gap caused by disconnected short-read fragments.Figure 2: Coverage gaps versus phasing ambiguity.

Read the Allele Name Before Comparing Results

An allele name such as HLA-A*02:01:01:01 is compact notation for several layers of sequence distinction. The first field groups alleles historically associated with a serological family. The second field distinguishes alleles that encode different protein sequences. The third field records synonymous coding differences, and the fourth field records differences outside the coding sequence. Suffixes can add expression or sequence-status information, so they should not be removed casually during data export.

These fields are not a universal scorecard. A four-field designation can be scientifically meaningful only when the assay observed or defensibly inferred the regions that distinguish the candidates. Conversely, a two-field result can be the appropriate end point if the study question is explicitly limited to protein-level allelic distinction and the uncertainty outside that scope is declared.

G groups and P groups are useful reporting tools, but they answer different questions. A G group combines alleles that share identical nucleotide sequence across the exons encoding the peptide-binding domains: exons 2 and 3 for Class I and exon 2 for Class II. A P group relates alleles that encode the same peptide-binding-domain protein sequence. Neither label proves that the full genomic sequence is identical. The report should include the database release because group membership and named alleles evolve as the official reference is updated.

HLA-A allele name mapped to the four nomenclature fields, a gene model, and expression-status suffixes.Figure 3: Four-field HLA nomenclature mapped to gene structure.

Define Loci and Reportable Resolution Before Choosing an Assay

Do not begin with a fixed panel just because it is familiar. Start with the biological or translational research question, then state which loci and which reporting level are needed to test it. Many projects need a focused subset of classical loci; others need a broader Class I and Class II package because their downstream comparison depends on both antigen-presentation contexts.

For example, a study of class I-restricted cellular recognition may prioritize HLA-A, HLA-B, and HLA-C. A question involving class II-associated immune context may require DRB1, DQB1, and DPB1, with DQA1 or DPA1 added when the planned interpretation depends on heterodimer context. The correct set is not the longest set. It is the minimum locus package that preserves the question's inference.

Write a short analysis contract before samples move to the laboratory. It should specify the loci, the intended field level or group-level output, whether phase must span the full gene, acceptable ambiguity classes, and what will trigger orthogonal confirmation. This contract prevents a common failure mode: receiving an apparently precise report that cannot answer the original question because the target locus or required phase boundary was never defined.

Evidence need Minimum design feature Defensible report language Escalation trigger
Peptide-binding-domain grouping Direct coverage of the relevant polymorphic exons for the class and locus G or P group when the observed coding region supports that grouping Candidate alleles differ outside the observed region and the distinction matters
Protein-level two-field call Evidence at the coding positions that distinguish the leading allele pair Two-field allele call, with residual ambiguity stated where present The observed positions do not separate the leading candidates
Full-gene distinction Continuous observation of every reportable region, including the positions that discriminate candidates Full-gene or field-level context only for the regions actually observed An unobserved intron, UTR, or distant exon separates plausible candidates
Cis-trans relationship Molecule-spanning reads or another validated linkage strategy across the relevant variants Phased allele-pair assignment Distant variants remain unlinked across fragments or amplicons

This Spoke article focuses on locus-level HLA evidence. The broader HLA Typing and Immune Repertoire Sequencing Services Hub connects HLA context with immune-repertoire and HLA-KIR questions; it should not be used as a substitute for a locus-specific reportability plan.

Research question branching to Class I, Class II, and combined HLA locus packages with field, phase, and coverage checks.Figure 4: Question-to-locus decision matrix.

Amplicon Boundaries Determine What Can Be Claimed

In HLA typing, primer and capture boundaries are part of the interpretation. A targeted assay may generate excellent depth over selected polymorphic exons while leaving the rest of the gene unobserved. A long-range PCR design may connect more variants across a locus, but its performance depends on DNA integrity, primer compatibility, and balanced amplification of both alleles. Capture-based designs can broaden observation but need explicit assessment of locus coverage and mapping specificity.

The planning question is not merely whether a locus is "covered." Ask where the amplicon begins and ends, which discriminating positions lie inside it, whether the design bridges the variants that must be phased, and whether an allele-specific primer mismatch could reduce representation. These questions are especially important in highly polymorphic loci, where a primer-binding difference may cause allele dropout or imbalance that looks like a clean homozygous call.

Targeted Region Sequencing is appropriate when the target intervals and reportable resolution are defined in advance. Amplicon Sequencing Services can support focused, high-depth interrogation, but the assay description must not imply that every full-gene difference is captured. When the decision depends on continuous locus-scale evidence, Long Amplicon Analysis is the more natural bridge between experimental design and phase-aware interpretation.

Build a coverage review into the report template. For each locus, retain target boundaries, covered intervals, low-confidence intervals, read balance, and the positions that separate the leading candidate alleles. This turns a call into an auditable measurement rather than a database label.

Full HLA locus shown with short exon targets, overlapping amplicons, and a long-range amplicon linked to different report claims.Figure 5: Amplicon design and reportable resolution.

Short Reads and Long Reads Solve Different HLA Problems

Short-read NGS is valuable when a project needs multiplexed, high-depth interrogation of defined regions across many samples. Read pairs and local assembly can resolve many common allele combinations, especially when the target design covers the informative regions and analysis is matched to the data type. Its limitation is structural, not a general defect: variants farther apart than the fragment or read-pair span may remain unlinked.

Long reads add a different kind of evidence. They can preserve the relationship among variants across a large amplicon or a full locus, which is particularly useful when the unresolved question is phase rather than base accuracy alone. Long reads are therefore an escalation tool for full-gene context, distant cis-trans ambiguity, or candidate alleles that differ outside conventional exon targets. They are not automatically required for every two-field research result.

The 2025–2026 trend is not simply "replace short reads." It is to align read architecture with the claim: use short reads for scalable focused calling when their span supports the needed distinction, and use long-read evidence where continuity across the molecule changes the interpretation. Nanopore Targeted Sequencing and Long-Read Sequencing Data Analysis are relevant when phase-aware, locus-scale evidence is part of the study design.

The validation standard should also reflect this distinction. Concordance at two fields does not demonstrate concordance at a full-gene level. A project should state the field level, loci, data source, and database release used for every comparison, rather than summarizing all agreement as one percentage.

Comparison of disconnected short reads and full-locus long molecules resolving cis-trans phase in a heterozygous HLA locus.Figure 6: Short-read assembly versus single-molecule phasing.

Build an Analysis Contract Before Generating Reads

HLA bioinformatics is a versioned inference problem. The same reads can produce different-looking reports when the reference database, caller, locus definitions, quality thresholds, or ambiguity policy changes. That does not automatically mean one report is wrong. It means the analytical contract must be visible enough to assess why the outputs differ.

The contract should record: sample identifier and input type; extraction and library strategy; target loci and regions; read or amplicon architecture; reference genome where relevant; IPD-IMGT/HLA release; caller and version; supported output resolution; allele-pair ranking or confidence logic; and all filters that can create a no-call or ambiguity flag. For existing WES or WGS data, add capture kit, read length, locus depth, and whether the caller was validated for that data geometry.

Software names are not interchangeable quality labels. Some callers are designed for raw short-read data, some for aligned reads, some for specific locus sets, and some for long-read assemblies. A useful comparison asks whether the tool's expected input and supported resolution match the experiment. A caller optimized for Class I from RNA-derived reads should not be used to make a full-gene Class II claim from sparse genomic data merely because it returns a formatted allele name.

For WGS or WES-derived data, accept an HLA inference as fit for the stated purpose only when the data geometry, locus completeness, caller validation, and requested report level all align. If any of those elements remains unsupported, label the result as screening or hypothesis-generating rather than presenting it as a resolution-equivalent substitute for a targeted design.

Use a two-level validation design when the study includes more than one data source. First, validate the experimental system with samples whose locus-level calls are independently established at the intended reporting level. This checks whether the primers, capture design, read processing, and caller work together for the locus package. Second, validate the project-specific data geometry: the exact read length, insert size, source material, and coverage distribution expected in the study. A caller that performs well on a deeply sequenced reference set may be a poor fit for shallower archival data or a capture panel that undersamples Class II intervals.

Pre-specify how discordant calls will be adjudicated. Record whether the comparison is at first field, second field, G group, three-field, or full-gene level; otherwise a result can look discordant simply because two pipelines report at different granularities. Then inspect whether the discordance is caused by database release, a missing interval, a mapping competition, low allelic balance, or truly conflicting sequence evidence. This workflow produces a meaningful discordance log instead of an undifferentiated list of mismatches.

When a study extends from a targeted HLA assay to wider genomic context, Whole Genome Sequencing, Gene Panel Sequencing Service, and Variant Calling should be framed as data-generation or analysis choices with specific coverage constraints—not as interchangeable routes to the same HLA evidence.

Versioned HLA calling trail from reads through reference, caller, quality control, and a locus-, phase-, and ambiguity-aware report.Figure 7: Versioned HLA calling audit trail.

A High-Resolution Result Needs an Orthogonal Verification Plan

Verification should be designed before the first run, not improvised after a difficult call. The goal is not to repeat every sample by a second method. It is to define the conditions under which the original evidence is insufficient for the intended claim.

Examples of escalation triggers include a locus with low or strongly imbalanced support, a candidate allele pair separated only by unobserved positions, a phase conflict across amplicons, an unexpected null or low-expression suffix, a possible novel sequence, or a sample identity inconsistency. The verification path should match the failure mode. Re-extraction helps when source DNA is compromised; a redesigned amplicon helps when primer compatibility or interval coverage is the issue; a longer contiguous read helps when cis-trans linkage remains unresolved.

Sanger Sequencing can be useful for targeted confirmation of a defined local variant, but it should not be described as a universal full-locus resolution fix. Its role is strongest when the precise region and alternative being tested are already known. The output should make clear whether confirmation addresses a local base, a segment, or the allele-pair phase itself.

An effective report labels each call as directly supported, inferred from a candidate combination, grouped because the observed region is identical, or pending confirmation. That language keeps an ambiguous result scientifically useful without presenting it as a complete allele assignment.

HLA verification decision tree routing low support, phase gaps, coverage gaps, and novel patterns to targeted follow-up actions.Figure 8: HLA call verification decision tree.

Sample Architecture Determines the Reliability of the Call

Sample type affects molecular evidence before sequencing begins. Blood-derived DNA, saliva, buccal swabs, and cord blood samples can all be suitable for research HLA projects, but they differ in DNA integrity, co-extracted material, cellular composition, and practical risk of low molecular weight DNA. A full-length or long-range strategy is more sensitive to fragmentation than a short targeted assay because a damaged template can reduce recovery of the long molecule needed for phase.

Use one sample intake specification for the entire study, then add a locus-aware QC layer. In addition to quantity and broad purity measures, document DNA integrity, extraction batch, sample identity controls, plate map, expected locus set, and re-test criteria. For a 300-sample academic study, the batch plan should distribute comparison groups across extraction, amplification, and sequencing batches. If all members of one biological group occupy one workflow batch, a locus-specific technical artifact can become indistinguishable from a biological association.

Allelic balance is especially important in amplicon-based HLA work. A low proportion of reads supporting one allele can arise from stochastic sampling, primer mismatch, poor template integrity, contamination, or a processing problem. The same observed imbalance does not have one universal interpretation, so the project plan should define a review zone rather than silently filtering the weaker allele. Retaining locus-level balance and coverage plots makes later troubleshooting possible when a result appears homozygous, unexpectedly ambiguous, or inconsistent with related samples.

For multi-batch studies, reserve a bridge sample or material equivalent that appears in more than one batch. The objective is not to claim that every run will be identical, but to detect shifts in locus recovery, allelic balance, and ambiguity patterns before comparing biological groups. This is particularly valuable when the work includes a mixture of newly extracted and archived DNA, because differences in sample condition can interact with amplicon length and create a non-biological pattern.

Include positive controls that exercise the difficult locus or resolution category, not only a sample that is easy to call. Include negative controls that detect contamination, and preserve duplicate or repeatable material for the small subset of samples meeting pre-defined escalation criteria. This is a project-design safeguard, not an optional add-on after results arrive.

How to Read a Research-Ready HLA Report

A research-ready HLA report is more than two allele strings per locus. At minimum, it should identify the locus, allele call or group assignment, field level, database release, analysis version, sequencing or coverage note, phase status, ambiguity description, and QC/review flag. It should also distinguish between a direct sequence-supported call and a conclusion inferred because multiple candidate alleles are indistinguishable across the observed interval.

Read the report in this order. First, check the intended locus set: a blank DPB1 result is not a minor omission if the project was designed around DPB1. Second, check the stated resolution and group notation. Third, identify whether the report records phase or only a compatible allele pair. Fourth, inspect whether coverage or allele balance makes one locus less reliable than the rest. Finally, check the reference release and caller version before comparing the result with a historical dataset.

The A Practical Guide to HLA Typing and Result Interpretation is useful background for reading report nomenclature. For a project-level decision, add the experimental boundaries described in this article: what was sequenced, what was phased, and what was left unresolved.

Isometric HLA report card annotated with locus, allele, field, phase, coverage, database release, and quality-control evidence.Figure 9: Research-ready HLA report checklist.

When to Escalate Beyond Conventional High Resolution

Escalation is justified when a specific evidence gap can change the study conclusion. Examples include a full-gene distinction that affects the intended analysis, a suspected novel sequence, a conflict between candidate allele pairs, an unresolvable phase boundary, or a locus with unusual mapping or copy-number context. The goal is to close that gap, not to pursue maximum detail across every sample by default.

Use a decision memo with three fields: the current call, the unresolved distinction, and the method that can observe the missing evidence. This makes it easy to separate scientifically necessary follow-up from technically impressive but uninformative extra sequencing. It also preserves budget and sample material for the cases where full-gene evidence genuinely matters.

For long-read implementation details, the existing Full-Length HLA Typing Method Using PacBio SMRT Sequencing Protocol and Full-Length HLA Genotyping with Nanopore Sequencing Protocol provide complementary background. The planned HLA-KIR Typing Spoke will address the separate question of combined HLA-KIR genetic interpretation; it is intentionally outside this article's scope.

Frequently Asked Questions

Is two-field typing always enough for research?

No. Two-field typing can be appropriate when it matches the pre-specified hypothesis, but it should not be presented as proof of unobserved intronic, UTR, or phase-resolved full-gene sequence.

Does four-field resolution prove the whole gene was sequenced?

Not by itself. The report must show that the assay covered or validly resolved the sequence regions that distinguish the reported four-field alleles.

What does a G group mean in an HLA report?

It identifies alleles sharing identical nucleotide sequence across peptide-binding-domain exons. It does not establish complete genomic identity outside those exons.

Can WGS data replace targeted HLA typing?

It can support HLA inference when read architecture, depth, capture design, and caller validation are adequate for the required loci and resolution. It is not automatically equivalent to a targeted full-locus assay.

Why can two valid pipelines report different allele calls?

They may use different database releases, reference assumptions, target loci, filters, or ambiguity policies. Compare inputs and versioned analysis settings before treating the difference as biological discordance.

When should an ambiguous result be re-tested?

Re-test when the ambiguity changes the planned interpretation, when sequence support is low or imbalanced, or when a phase or coverage gap remains at a critical locus.

Research Use Only. This content supports experimental planning and research interpretation; it is not diagnostic or treatment guidance.

References:

  1. Mayor NP, Robinson J, McWhinnie AJM, et al. HLA Typing for the Next Generation. PLOS ONE. 2015;10:e0127153. DOI: 10.1371/journal.pone.0127153.

  2. Liu C, et al. Benchmarking the Human Leukocyte Antigen Typing Performance of Three Assays and Seven Next-Generation Sequencing-Based Algorithms. Frontiers in Immunology. 2021;12:652258. DOI: 10.3389/fimmu.2021.652258.

  3. Dashti M, Malik MZ, Nizam R, et al. Evaluation of HLA Typing Content of Next-Generation Sequencing Datasets From Family Trios and Individuals of Arab Ethnicity. Frontiers in Genetics. 2024;15:1407285. DOI: 10.3389/fgene.2024.1407285.

  4. Dilthey AT, Gourraud P-A, Mentzer AJ, et al. High-Accuracy HLA Type Inference From Whole-Genome Sequencing Data Using Population Reference Graphs. PLOS Computational Biology. 2016;12:e1005151. DOI: 10.1371/journal.pcbi.1005151.

  5. Anukul N, Jenjaroenpun P, Sirikul C, et al. Ultrarapid and High-Resolution HLA Class I Typing Using Transposase-Based Nanopore Sequencing Applied in Pharmacogenetic Testing. Frontiers in Genetics. 2023;14:1213457. DOI: 10.3389/fgene.2023.1213457.

For research purposes only, not intended for clinical diagnosis, treatment, or individual health assessments.
Speak to Our Scientists
What would you like to discuss?
With whom will we be speaking?

* is a required item.