Nolege News

Biotechnology

Contact tracing gives you a plausible story. Sequencing tells you which stories are impossible

By ·28 September 2026·7 min read

🌐 इस लेख को हिन्दी में पढ़ें
Contact tracing gives you a plausible story. Sequencing tells you which stories are impossible

In short: Genomic epidemiology compares pathogen sequences to test whether cases are linked. This guide explains how mutation accumulation works as a rough clock, why clock speed determines what a pathogen can resolve, why a phylogenetic tree relates sequences rather than people and cannot establish direction of transmission, where sequencing is genuinely decisive through exclusion, how resistance can be predicted from sequence and where that prediction fails, and the sampling and within-host limits that are rarely stated.

Six patients on the same ward develop the same infection within three weeks. Contact tracing produces a story: this patient was next to that one, a nurse moved between both bays, the timing works. The story is plausible, internally consistent, and there is no way to check it against anything.

Then someone sequences the six isolates, and three of them turn out to differ from the other three by far more mutations than a few weeks of transmission could possibly produce. There were never six linked cases. There were two unrelated clusters, and the investigation had been building a narrative to connect them.

This is what genomic epidemiology is actually for, and it is both less and more than the way it gets reported.

Copying errors as a rough clock

Every time a pathogen replicates, its copying machinery makes occasional mistakes. Most of these changes do nothing — they neither help nor hinder the organism — and so they simply persist and get inherited by everything descended from that cell.

That gives you a measuring device. Two samples that share a recent ancestor have had little time to accumulate differences, so they look nearly identical. Two samples separated by many transmission events and many months differ in more places. Count the differences and you have a rough estimate of how much replication separates them.

The word "rough" is doing real work. Mutations arrive at random, not on a schedule, so the count is a noisy estimate rather than a date. And clock speeds differ enormously between pathogens, which is the single most important fact about what this technique can do.

A fast-mutating RNA virus accumulates changes across its genome every few weeks, so samples taken a month apart are visibly different and fine-grained tracing is possible. Mycobacterium tuberculosis is at the opposite extreme: it changes by well under one mutation per genome per year. Two TB isolates from a genuine recent transmission may be identical, which sounds like a failure and is actually informative — because it means that isolates differing by a dozen mutations are almost certainly not from the same recent chain, whatever the contact history suggests.

What a tree is, and what it is not

From a set of sequences you can build a phylogeny — the branching structure that best explains the pattern of differences. This is where most misreporting happens, so it is worth stating plainly.

A phylogeny relates sequences. It does not relate people, and it does not have arrows on it. If two patients' isolates are identical, that is equally consistent with the first infecting the second, the second infecting the first, or both being infected by a third person nobody sampled. The tree constrains the possibilities; it does not select among them.

Direction and timing come from combining the tree with everything else — dates of symptom onset, ward movements, who was where. Sequencing is one input to an epidemiological investigation, not a replacement for it, and a transmission chain drawn confidently from genomes alone is usually a chain drawn from assumptions.

Sequencing rarely tells you who infected whom. It tells you which stories are impossible — and in an outbreak investigation, eliminating stories is most of the work.

Where it is decisive

The power is in exclusion, and exclusion changes decisions.

Splitting apparent outbreaks. The ward example above is routine. Two unrelated clusters presented as one produce a search for a shared cause that does not exist, and the sequencing result redirects effort immediately.

Joining apparent coincidences. The reverse is just as valuable. Near-identical isolates turning up in three hospitals in different cities implies a common source, which moves the investigation off local hygiene and onto something shared — a device, a product, a supplier, a colonised individual moving between sites. Food-borne outbreak investigation now runs largely on this logic: matching a patient isolate to an isolate from a production facility is what turns a suspicion into a recall.

Relapse versus reinfection. This one is clinically consequential and specific to slow-clocked pathogens like TB. A patient who completes treatment and becomes ill again has either relapsed with the same strain — meaning the treatment failed, and why it failed needs understanding — or been reinfected by a new one, meaning treatment worked and exposure continued. The two demand different responses, and without sequencing they are indistinguishable. With it, comparing the two isolates usually settles it.

Reading resistance off the sequence

There is a second, more direct use. For some pathogens and some drugs, resistance is caused by specific known mutations. If you have the genome, you can look them up.

For tuberculosis this matters enormously, because the conventional method — growing the organism in the presence of each drug — takes weeks for a slow-growing bacterium, during which the patient is being treated on a guess. Sequencing can return a predicted susceptibility profile in a small number of days, and curated catalogues of resistance-associated mutations exist precisely to standardise that interpretation.

The honest caveats matter as much as the capability. Prediction is strong for drugs whose resistance mechanisms are well characterised and weaker for others. And absence of a known resistance mutation is not proof of susceptibility — it may mean resistance by a mechanism not yet catalogued. A sequence-based prediction is a fast, usually-right input, not a verdict that retires culture entirely.

The limits that rarely get stated

Sampling. A tree contains only what was sequenced. Unsampled intermediates are invisible, so two isolates separated by several unsampled people can look like direct transmission. In settings where most cases are never captured — which is most settings — the tree is a sketch of a much larger structure.

Low diversity. If a pathogen has barely changed across a whole region, everything looks related to everything. Resolution runs out, and sequencing stops being able to distinguish chains at all.

Within-host variation. A person does not carry one genome; they carry a population. Standard reporting collapses that into a consensus sequence, which is a summary. Two samples from the same patient can differ, and a minority variant present in one host may be the one that transmits.

Artefacts. Contamination between samples, low-coverage regions and alignment errors all produce apparent mutations that are not real. In a technique whose conclusions turn on counting a handful of differences, a small number of false differences is not a rounding error — it is the answer.

Why it matters for students and researchers

India has the largest tuberculosis burden in the world and a substantial drug-resistant share of it, which makes both applications above directly consequential here rather than illustrative. Faster resistance prediction shortens the period a patient spends on a regimen that will not work. Distinguishing relapse from reinfection tells a programme whether its problem is treatment quality or continued transmission — two completely different interventions, currently often guessed at. And the country built substantial sequencing and surveillance capacity during the pandemic, which is infrastructure now looking for its next durable use.

The discipline worth learning from this field is a specific one: it reports what the data excludes. A well-written genomic epidemiology finding says that these cases cannot belong to one chain, or that a common source cannot be ruled out — statements that are weaker-sounding and far more defensible than the confident arrows readers want. That habit, of saying precisely what your evidence forbids rather than what it hints at, transfers to every kind of analysis, and it is rarer than it should be.

Frequently asked questions

How does sequencing trace an outbreak?

By comparing pathogen genomes from different cases. Because mutations accumulate as the organism replicates, closely related samples differ in few places and distantly related ones differ in more, so the pattern of differences indicates which cases could plausibly be linked.

Can sequencing prove who infected whom?

Usually not. A phylogenetic tree relates sequences rather than people and carries no direction, so identical isolates are equally consistent with transmission either way or with an unsampled third source. Direction requires epidemiological information as well.

Why does mutation rate matter so much?

Because it sets the resolution. Fast-mutating viruses differ measurably within weeks, allowing fine-grained tracing, while slow organisms such as tuberculosis may show identical genomes in genuinely linked cases — which still allows unlinked cases to be ruled out.

Can drug resistance be predicted from a genome?

For several drugs, yes: specific mutations are known to confer resistance and can be looked up in curated catalogues, giving an answer in days rather than the weeks culture requires. But absence of a known mutation does not guarantee susceptibility.

What is the difference between relapse and reinfection, and why does sequencing settle it?

Relapse is recurrence of the original strain, implying treatment failure; reinfection is a new strain, implying continued exposure. Comparing the genomes of the first and second isolates distinguishes them, and the two situations call for different responses.