From a variant to a call: Recommended Track Sets, regulation, and the evidence
Worked examples in germline & somatic interpretation · genome.ucsc.edu
A thread for today: three cancer variants
A BRCA2 variant: germline, “uncertain significance”. How do experts resolve it? (morning, in the Recommended Track Sets)
The TERT promoter: a non-coding driver. The answer lives in the regulation. (morning, then the epigenetics section)
BRAF V600E: a coding driver in melanoma. Somatic, famous, druggable. (afternoon, the somatic worked example)
Watch them recur
We will come back to these variants throughout the session. The goal is that by the end you can interpret a variant, explore its regulatory context, load your own data, and share it as a link.
Interpreting a variant = asking questions (germline)
Databases: ask a question, know which track answers it:
Is it already classified? → ClinVar, ClinGen(germline pathogenicity, ClinVar also carries somatic oncogenicity)
Is it druggable? → CIViC (variant → disease → therapy → evidence)
Germline first
These questions fit a germline variant, and the Recommended Track Sets (next) bundle exactly these tracks. Later we revisit the same questions for somatic cancer variants.
Recommended Track Sets
Hundreds of tracks is overwhelming, start from a curated set
Recommended Track Sets: The problem they solve
The Browser has hundreds of tracks, beginners don’t know which to turn on.
Recommended Track Sets = pre-configured collections for a scenario.
One click turns on a themed set, without changing your locus.
Open via the “Recommended Track Sets” menu item.
The Recommended Track Sets menu: each link loads a curated, themed set of tracks at your current position.
Seven sets on hg38
Clinical SNVs: disease contribution of coding SNVs
Clinical CNVs: coding structural variants
Non-coding SNVs: functional context of non-coding variants
Determine Exon Relevance: is the variant in a required exon?
Problematic Regions: low-confidence / high-homology regions
Same codon / nearby seen before? (PS1 / PM5) → ClinVar
In a key protein domain? (PM1) → UniProt: the DNA-binding domain (OB1 fold) ✓
Computational? (PP3) → REVEL ≈ 0.93 + deep conservation ✓
Loss-of-function? (PVS1) → gene model (missense → N/A)
The point
Each line of evidence stacks in one view, here they all agree on pathogenic. You read the ACMG codes off the screen. (The harder, guideline-dependent cases come in the ENIGMA set.)
At p.Asp2723His: RefSeq/MANE, the BRCA2 OB1 UniProt domain (PM1), and ClinVar variants stacked.
Demo 2: Non-coding SNVs → epigenetics (1/2)
For variants outside coding exons, the regulatory evidence:
GeneHancer enhancers & enhancer→gene links
Hi-C / Micro-C 3D chromatin contacts
JASPAR TF binding sites · 100-way conservation
Our bridge to epigenetics later in this deck, and the set for our TERT promoter variant.
Non-coding SNVs near the TERT promoter: GeneHancer regulatory elements, JASPAR TF sites, and conservation, the context a non-coding variant needs.
Demo 2 · cont: a non-coding variant at TERT (2/2)
TERT promoter hotspot mutations (e.g. C228T · NM_198253.3(TERT):c.-124C>T) sit ~100–150 bp upstream of the start codon, in the core promoter. With the Non-coding set, ask:
In a regulatory element? → GeneHancer / ENCODE cCRE (promoter)
Creates / breaks a TF site? → JASPAR (TF motifs here; the C228T / C250T hotspots create a new ETS / GABPA site that switches TERT back on)
Contacts a distal gene in 3D? → Hi-C / Micro-C (most useful for enhancer variants; less so for this promoter)
Evolutionarily constrained? → conservation
No coding ACMG here
Non-coding variants aren’t scored by coding ACMG rules, you weigh the regulatory evidence instead.
The TERT promoter: ENCODE cCRE, GeneHancer, and dense JASPAR TF-binding sites, where the hotspot builds a new ETS site.
Demo 3: expert-panel gene sets
This is the home of our germline BRCA variant of uncertain significance.
ENIGMA BRCA1/BRCA2 VCEP: the exact evidence for ClinGen ENIGMA classification, per exon & variant.
InSiGHT Lynch Syndrome VCEP: same idea for MLH1/MSH2/MSH6/PMS2.
Same variant, different rules, ▶ open itBRCA2 c.830A>G (p.Asn277Ser): standard ACMG reached likely benign (REVEL → BP4); under ENIGMA it reverts to VUS (BayesDel not allowed here; SpliceAI → PP3). The spec can move a call either way; overall it cut VUS, but not for every variant.
Published co-authored
Benet-Pagès, Laner, Nassar… Genet Med Open 2025. Session (hg19): /s/abenet/BRCA.ENIGMA.hg19
The ENIGMA BRCA1/BRCA2 VCEP set at a BRCA1 exon (hg19): the exact evidence the panel rules use, per variant.
Somatic diagnosis
Somatic variants
acquired mutations: SNVs, de novo changes, and cancer drivers
Interpreting a variant = asking questions (somatic)
Databases: ask a question, know which track answers it:
Is it already classified? → ClinVar, ClinGen(germline pathogenicity, ClinVar also carries somatic oncogenicity)
Is it druggable? → CIViC (variant → disease → therapy → evidence)
Germline vs somatic
These questions fit a germline variant. For a somatic driver (BRAF V600E, next) lean on COSMIC & CIViC and ClinVar’s somatic oncogenicity / clinical-impact, and don’t read gnomAD frequency as “benign” (a true somatic variant is just absent). REVEL / conservation flag a damaging residue, not oncogenicity.
Worked example: BRAF V600E
The classic melanoma driver: a somatic variant. BRAF V600Ep.Val600GluNM_004333.6:c.1799T>Achr7:140,753,336
Try it, ▶ open the BRAF V600E sessionCOSMIC: how often in tumours? · CIViC: oncogenic & druggable? · ClinVar: its somatic oncogenicity / clinical-impact, not the germline label.
Somatic ≠ germlinegnomAD won’t “verify” it: a true somatic variant is simply absent from healthy-population data (gnomAD filters out inherited variants).
Thread tie-in
Our coding driver: the Recommended “Clinical SNVs” set assembles these in one click.
The BRAF V600E session: ClinVar, COSMIC and CIViC stacked at chr7:140,753,336.
Expression
where, and in which cell type, is a gene switched on?
Three expression datasets on hg38
“Where is my gene expressed, and in which cell type?” Two axes separate these tracks: bulk tissue vs single cell, and one uniform study vs many pooled together.
GTEx Gene V8
Bulk tissue reference.
54 tissues, 948 donors, bulk RNA-seq. On by default.
Each bar is a whole tissue, so every cell type is averaged together.
Best for: which organ is the gene expressed in?
Tabula Sapiens
One uniform single-cell atlas.
~480k cells across ~24 organs, one consortium, consistent processing.
Bars split by tissue and by cell type, so signal resolves to a cell type.
Best for: which cell type, answered cleanly within one atlas.
Merged Single-Cell
Many single-cell studies at once.
Pools many published atlases (incl. Tabula Sapiens) into one track.
Widest coverage, but heterogeneous: compare within a dataset, not across.
Best for: does it hold across the single-cell literature?
GTEx gives you the organ, the single-cell tracks refine it to the cell type, and the merged track checks whether it holds across many studies.
Expression: what tissue is it expressed in?
KLK3 on hg38 with GTEx, Tabula Sapiens, and single-cell tracks: the signal spikes in prostate luminal epithelium and is near-silent elsewhere.
KLK3 encodes PSA (prostate-specific antigen), the protein behind the prostate-cancer blood test. Expression is restricted to prostate luminal epithelium, so the contrast is clear.
Our thread genes (BRAF, TERT, BRCA2) are broadly expressed, so we pick a textbook tissue-specific gene to make the contrast obvious.
Try it
Look up KLK3; read its GTEx bars (prostate towers over the rest), then Tabula Sapiens for the cell type. Load the session.
Regulation & epigenetics
enhancers · histone marks · open chromatin · methylation
Picking up from “Non-coding SNVs”
That set pointed us at GeneHancer, Hi-C/Micro-C, JASPAR and conservation.
They all live in the Regulation group, largely from ENCODE.
Heads-up
Two “Regulation” super-tracks exist: ENCODE3 & ENCODE4. Use ENCODE4.
Try it, ▶ a real 3D loop at MYC
HFFc6 Micro-C: a stripe off the MYC promoter reaches a loop dot in the 8q24 enhancer desert (chr8:128.31–128.33 Mb).
The Regulation group: ENCODE cCREs, DNA Methylation, GeneHancer, Hi-C and Micro-C, JASPAR, VISTA Enhancers and more, all in one place.
Enhancers & promoters: cCREs
ENCODE4 cCREs: candidate cis-regulatory elements, on by default.
The TERT promoter: the red ENCODE4 cCRE = promoter-like; orange = enhancer-like. Layered H3K27Ac and conservation sit alongside.
TERT: the data say “active promoter” (2/2)
At the TERT TSS: red promoter cCRE + H3K27Ac + a DNase peak + GeneHancer: every track points to an active promoter.
cCRE coloured red = promoter-like element
H3K27Ac peak = active regulatory region
DNase hypersensitivity = open chromatin
GeneHancer marks the TERT TSS / regulatory element
Conclusion
Independent assays converge → a real, active promoter, exactly where a non-coding driver mutation bites.
A variant-interpretation toolkit
AI predictors and population frequencies
AlphaMissense: AI missense pathogenicity
Google DeepMind's deep-learning score for every possible missense substitution (Cheng et al., Science 2023).
One value per change, 0 to 1: likely benign → ambiguous → likely pathogenic, drawn at base resolution.
The AI counterpart to REVEL: another computational line of evidence (ACMG PP3) for a coding variant.
Try it
Open the AlphaMissense track at BRAF V600 or BRCA2; read the score for our variants.
Why it fits
We leaned on REVEL for PP3 earlier. AlphaMissense is the same idea, trained by an AI model, and it covers the whole protein so you can scan a gene for predicted hotspots.
SpliceAI: predicting splice disruption
Illumina's deep-learning predictor of whether a variant creates or breaks a splice site (Jaganathan et al., Cell 2019).
Delta scores (0 to 1) for acceptor / donor gain & loss, with the predicted base position.
Lives under the Splicing Impact super-track; SNVs and indels, raw and masked.
Try it
Open Splicing Impact and inspect a splice-region variant.
Closes a loop
The ENIGMA demo (Demo 3) cited SpliceAI → PP3 as the evidence that moved a BRCA2 call. This is that track: see the score that drove the reclassification.
Variant frequencies: how common, everywhere
SNV Frequencies: allele frequencies gathered from population-scale projects worldwide, ~1.7 million genomes / exomes / arrays.
One place to compare how common a variant is across populations, ancestries and cohorts, including national projects not in gnomAD.
Combined tracks aggregate the data, plus one subtrack per project (TOPMed, gnomAD, 1000 Genomes, and more).
note Collected as-is, not re-harmonized: pipelines and assays differ between projects.