Inside the Black Box: How Probabilistic Genotyping Software Turns a DNA Mixture Into a Number, and Where That Number Can Break

Blue DNA double helix representing forensic DNA evidence review for the defense

A jury heard that the DNA evidence was 21.4 trillion times more probable if the defendant contributed than if he did not. No analyst counted to 21.4 trillion. A computer did. It ran a model, made assumptions, and printed a number that sounded like certainty. The defense asked to see how the machine got there. The company that built it said no.

That is the fight, in one paragraph. Below is the science underneath it, laid out the way an expert would walk a defense team through it, step by step. If you want the shorter, courtroom-facing version of this story, the Ambeau Law Firm covers it for defendants and lawyers. This piece goes deeper into the mechanics.

What a DNA Mixture Actually Is

Start with the sample. A single-source sample comes from one person. A gun grip, a cigarette butt handled by one smoker, a bloodstain from one wound. Those are easy. The lab reads the profile and compares it. No software needed for the hard part.

A mixture is different. Two, three, four people leave their DNA on the same object. A steering wheel touched by a car’s owner, a passenger, and a thief. Now the lab does not see one profile. It sees everyone’s DNA piled on top of everyone else’s, at every location it tests.

Picture a specific case. Joe owns the car. Jill borrowed it last week. Then someone steals it and leaves it in a ditch. The crime lab swabs the wheel and finds a three-person mixture. Joe and Jill are in there for innocent reasons. The third person is the question. The lab wants to know if the police suspect, call him Ray, is that third contributor. Every genetic location on that swab now carries a jumble of Joe, Jill, and someone else, and the analyst has to pull them apart.

Where the Old Method Broke

For years, labs handled mixtures with a method called combined probability of inclusion, or CPI. It asks one blunt question at each location: can I rule this person out, yes or no. If the suspect’s genetic markers all appear somewhere in the jumble, he is included. Multiply the inclusion probabilities across every location, and you get a statistic.

CPI has two deep flaws. First, it is one-sided. It only asks whether Ray fits. It never asks how many other people would fit just as well. In a busy three-person mixture, a huge slice of the population might also be included, but CPI does not force the analyst to say so out loud. That tilts the method against the accused.

Second, CPI throws away information. It ignores how much DNA each person left. It ignores peak heights, the signal strength the instrument records for each genetic marker. Two contributors who left very different amounts of DNA look the same to CPI as two who left equal amounts. Dr. Mark Perlin, who built the TrueAllele system, put the criticism in stark terms in a 2015 interview with Forensic Magazine, calling CPI “a random number generator” and warning that convictions had rested on it. The field listened. CPI fell out of favor, and probabilistic genotyping moved in.

What the Software Does Instead

Probabilistic genotyping does not ask yes or no. It builds a model of the whole mixture and asks how well different explanations fit the data. Programs like TrueAllele and STRmix are the two names you will hear most, and Louisiana labs use them.

Go back to the swab from Joe’s car. The software does not just note which genetic markers are present. It reads the peak heights, the height of each signal the instrument recorded. Then it proposes a full explanation of the sample. It might guess: three contributors, one who left a lot of DNA, one who left a medium amount, one who left a trace, and here are the exact profiles I think each of them has. That is one hypothesis. The software scores how well that guess reproduces the actual peak heights on the swab.

Then it does it again. And again. Millions of times. Each run tweaks the guessed profiles, the guessed amounts, the guessed number of contributors, and rescores. The explanations that reproduce the real data well get kept and refined. The ones that fit badly get discarded. This is the engine, and it has a name.

Markov Chain Monte Carlo, in Plain English

The math behind those millions of guesses is called Markov Chain Monte Carlo, or MCMC. The name sounds forbidding. The idea is not.

Imagine you are dropped somewhere in a dark, hilly landscape and told to find the highest ground. You cannot see. So you take a step. If the ground rises, you tend to keep going that way. If it drops, you sometimes step back, but not always. You wander like this for a long time. Do it long enough, and you spend most of your time on the high peaks. Map where you walked, and you have mapped the high ground.

MCMC does that with DNA explanations. Height means fit. The peaks are the profiles that best explain the mixture. The software wanders through the space of possible explanations, spending more time on the ones that fit the data well. When it is done, it has a picture of which contributor profiles are plausible and how plausible each one is.

Two things follow from this, and both matter in court. First, the process is random. Run it twice and you will not get identical numbers, only close ones. That is expected, but a defense expert should confirm the runs are stable. Second, where the software starts and how long it walks affect the answer. Cut the walk short and it may never find the true high ground. These are not bugs. They are settings, and settings are choices a human made.

The Messy Parts: Drop-Out, Drop-In, and Stutter

Real evidence is dirty. Three problems come up in almost every hard case, and how the software handles each one changes the result.

Drop-out is when a person’s DNA is present but too faint to show up. Ray touched the wheel for two seconds. He left barely any cells. At some genetic locations, his markers never cross the detection threshold. The software has to decide how likely it is that a real contributor simply failed to register. Set that probability too high, and the model forgives missing evidence too easily, which helps the prosecution. Set it too low, and it excludes people it should not. This single assumption can swing a case.

Drop-in is the opposite. A stray bit of DNA from handling or lab contamination shows up where it does not belong. The model has to allow for that too, or it will treat noise as a fourth contributor.

Stutter is an artifact of the testing chemistry itself. When the machine copies a genetic marker, it sometimes produces a small false peak next to the real one. A stutter peak can look exactly like a faint contributor. The software must separate true signal from stutter, and it does that with, again, an assumption about how much stutter to expect. Get the stutter model wrong and a phantom contributor appears, or a real one vanishes.

Degraded DNA makes all of this worse. Sun, heat, water, and time break DNA into shorter fragments. The larger genetic markers fade first, so the profile tilts. A good model accounts for degradation. A model told to assume a clean sample will misread a degraded one.

The Number at the End: The Likelihood Ratio

After the millions of runs, the software produces a likelihood ratio, or LR. This is the number that reaches the jury, and it is widely misunderstood, so be precise about what it means.

The LR compares two competing stories. Story one: Ray and two unknown people left this DNA. Story two: three unknown, unrelated people left it, and Ray had nothing to do with it. The software asks how well each story explains the actual evidence. The LR is the ratio between them. An LR of 21.4 trillion means the evidence is 21.4 trillion times more expected under the first story than the second.

Read that sentence again, because here is what it does not say. It does not say there is a 21.4 trillion to one chance Ray is guilty. It does not say Ray is a match. It is a statement about the evidence under two hypotheses, not a statement about Ray. Confusing the two is called the prosecutor’s fallacy, and it has put innocent people in prison. A defense expert’s first job is often to stop a jury from hearing the number as a verdict.

The LR also depends entirely on that second story. Who counts as the random alternative. What population database the software compares against. A profile that is 21.4 trillion times more probable against one reference population might land at a very different number against another. The denominator is a choice, and choices can be examined.

Why the Number of Contributors Is a Fight of Its Own

Before the software runs, someone tells it how many people are in the mixture. Two? Three? Four? That input is not read off the machine. An analyst decides it, sometimes from the same messy data the software is about to interpret.

This matters more than almost anything else. Tell the model there are two contributors when there are really three, and it will force the evidence into the wrong shape. The reported LR can move by orders of magnitude depending on that one number. In Joe’s car, if the analyst assumes two contributors and forgets that Jill drove it last week, the whole analysis is built on sand. Always ask how the contributor count was chosen, and whether a different, defensible count changes the answer.

Validation Is Not the Same as Being Right in Your Case

Developers point to validation studies to show the software works. Those studies matter, and the developer-run and independent validation work catalogued by the National Institute of Justice is worth reading closely. But validation proves the tool performs well on the samples it was tested against. It does not prove the tool performed well on your sample.

Ask the pointed question. Was this software validated on mixtures like the one in this case? Same number of contributors, same low DNA quantity, same degradation, same ratio between the major and minor contributors? A tool validated on clean two-person mixtures is not automatically reliable on a degraded four-person trace. The lab’s own internal validation file is discoverable, and it often reveals the edges of what the tool was ever shown to handle.

These Programs Have Been Wrong Before

Software written by people carries the mistakes of people. Two examples are now part of the record. New York City’s own mixture tool, the Forensic Statistical Tool, was retired after a court-ordered look at its code revealed a function that could tilt results against defendants in a way that had never been disclosed. Separately, the maker of STRmix acknowledged a coding error that affected results in roughly sixty cases in Australia before it was caught.

Neither of those was found by reading a validation study. They were found by looking at the code and the outputs directly. That is the whole argument for access in one sentence.

The Source-Code Fight

When the defense asks to see the code, the company usually refuses and calls it a trade secret. Courts have split. Some order the code released to a defense expert under a protective order, so it can be examined without publishing it to competitors. Others rule that published validation and courtroom track record are enough, and refuse.

There is a middle path worth knowing about. The source code is not always the most useful thing to get. The maker of STRmix has argued that its extended output, the detailed record of what the program actually did on your specific sample, tells an expert more than the raw code would. That record shows the assumptions used, the contributor count, the drop-out and stutter settings, and how the model landed where it did. For most defense purposes, that is the gold. Fight for it whether or not you also get the code.

What a Defense Expert Actually Does

Put the pieces together and the attack has a shape. It is not “the computer is wrong.” It is more surgical than that.

Check the contributor count and rerun the analysis with a different, defensible number. Examine the drop-out and stutter assumptions and see whether reasonable alternatives change the LR. Confirm the MCMC runs were stable and long enough. Test the LR against a different reference population. Pull the lab’s internal validation file and compare the tested conditions to the conditions in this case. Read the extended output for anything the summary report left out. Any one of these can move a trillion-to-one number toward something a jury should weigh very differently.

Frequently Asked Questions

Is probabilistic genotyping junk science?

No. Used correctly, on the kind of sample it was validated for, it is a genuine advance over the old methods. The problem is not the science. The problem is treating its output as beyond question when the inputs and assumptions behind it were never examined.

Can two programs disagree on the same sample?

Yes. TrueAllele and STRmix use different models and different assumptions. Run on the same difficult mixture, they can report meaningfully different numbers. That alone tells you the result is not a fixed fact of nature.

Do I need the source code to win?

Often, no. The assumptions, the contributor count, the validation file, and the extended output are usually where a case is won or lost. Source-code access is one tool, not the only one.

Does Louisiana admit this evidence?

Louisiana labs use these tools, and Louisiana courts have handled cases built on them. Admissibility runs through the state’s standard for expert testimony, which means the reliability of how the tool was applied in your case is fair game.

The Bottom Line

A trillion-to-one number is not a fact handed down from the machine. It is the end of a long chain of human choices: how many contributors, how much drop-out, how much stutter, which population, how long to run the model. Break any weak link and the number changes. Sometimes it changes enough to change everything.

If a probabilistic genotyping result is driving a case, do not let it stand unexamined. Get the assumptions, get the validation file, get the extended output, and put an expert on it. The team at the Ambeau Law Firm builds these challenges. Until the software is examined, the number on the page is a claim, not a conviction.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top