Outliers in Laboratory Data: When Excluding a Value Is Justified

A laboratory runs three replicate HPLC purity determinations on the same peptide sample. Two come back at 98.6 and 98.4 percent. The third reads 96.1. The deadline is tomorrow, the report template has room for one number, and someone suggests dropping the low value because “it is obviously an outlier.” It may be. It may also be the most informative measurement of the three. What the laboratory is allowed to do with it depends far less on the number itself than on what the laboratory knew, and wrote down, before it saw that number.

Why removal is the most dangerous edit in a data set

Almost every other analytical choice leaves a trace. A different integration setting can be reproduced. A different statistical test can be rerun by a reviewer. A deleted observation leaves nothing behind. The reader of a final table or graph cannot see what was taken out, how many values went, or on what grounds. A statement such as “one outlier was excluded” records that something happened without showing that it was justified.

That asymmetry is the reason exclusion deserves more scrutiny than any other step. A data set reported whole, with one odd value flagged and left unexplained, is more honest and more useful to a reader than a tidy set with the awkward point quietly gone.

Rejecting a run is not the same as rejecting a value

Two actions are often lumped together under the word “outlier,” and they deserve different handling in a laboratory’s procedures.

  • Invalidating a whole run because a control failed, or because the system did not pass its suitability criteria, rests on evidence that sits outside the measurements of interest. That is a defensible decision, and it can be written into a procedure in advance.
  • Deleting one value from an otherwise acceptable run rests, almost always, on the value itself looking wrong. The number is both the suspect and the only witness.

A sensible policy therefore puts its effort into run-level acceptance criteria and sets a very high bar for individual deletions. In chromatography, those run-level criteria usually take the form described in HPLC system suitability.

The test of an assignable cause

An individual result can legitimately be removed when there is an assignable cause that was identified independently of the result. Examples include a documented pipetting error, a vial that was recorded as mislabeled, a pressure fault logged by the instrument during that sample, or a visibly contaminated well noted before reading the plate.

The key word is independent. “It is far from the others” is not a cause; it is the observation that prompted the search. A cause that was found only because the analyst went looking after seeing a strange number, and accepted because it would explain that number, is weak evidence. In our HPLC example, a note in the run log saying that the third vial’s cap was found loose before the sequence started would be an assignable cause. Discovering a vague possibility after the fact would not.

What outlier tests actually do

Statistical procedures such as the Grubbs or Dixon tests flag values that would be unlikely if the data came from an assumed distribution, usually a normal one. That is all they establish. Two limits follow.

First, the assumption may be wrong. Real analytical and biological data are often skewed or have heavier tails than a normal curve, and a value that looks improbable under the wrong model may be ordinary under the right one. Second, improbable is not the same as erroneous. An unusual value can signal a genuine subpopulation, a threshold effect, or real instability in the material. A test can tell you a value is unusual; it cannot tell you the value is a mistake.

What deletion does to the numbers

Removing an extreme value reduces the apparent spread. Standard deviations shrink, confidence intervals narrow and p-values fall. If exclusions are made habitually, every result in a laboratory’s records claims more precision than the work supports.

The effect is worst in small sets. Dropping one of three replicates leaves two, which cannot give a meaningful estimate of variability. Yet small sets are exactly where a single odd value is most tempting to remove. With twenty replicates, one extreme value barely moves the summary; with three, it dominates it.

Better options than deleting

ApproachWhat it achieves
Report the analysis with and without the valueIf the conclusion holds either way, the question stops mattering; if it flips, that instability is itself the finding
Use a robust summary such as the median and interquartile rangeDescribes skewed data without anyone deciding which points deserve to stay
Collect more dataThe only option that adds information instead of removing it
Investigate the value as a signalAn unexplained extreme may reveal carryover, contamination or a sample handling problem worth fixing

In the purity example, a re-analysis of a fresh preparation, alongside a look at the chromatogram for the low replicate, would answer more than any test statistic. Reading that trace for a new peak, a shoulder or a baseline disturbance is described in how to read an HPLC chromatogram.

Write the rules before the data exist

The solution is procedural. Before results are seen, record what will invalidate a run. Typical criteria include a system suitability failure, a control outside a range fixed in advance, a recognized plate edge effect, or cells outside the planned confluence range. Criteria set beforehand can be applied without the outcome pulling on them. Criteria invented afterward cannot, however reasonable they sound, because a plausible explanation can be built for nearly any single number.

Where no such criteria were set, two honest options remain: keep the value, or report openly that it was removed after the fact and why. Presenting a trimmed set as if it were the complete result is not an option. The same logic explains part of why different laboratories report different purity figures for similar material.

Questions

Is a Grubbs test enough to justify removing a value?

No. It shows that a value is unusual under an assumed distribution. Removal needs an assignable cause identified independently of the value.

Can I discard a whole run instead?

Yes, when it fails criteria that were defined before the data were examined, such as system suitability or control limits.

What if I only have three replicates?

Be especially cautious. Removing one leaves too little data to estimate variability, so collecting more measurements is usually the better course.


Research use only. All products supplied by Battle Born Peptides are laboratory reference materials for in-vitro research and analytical use by qualified professionals. They are not drugs, foods, dietary supplements, cosmetics or medical devices; they are not approved by the FDA or any other regulator for use in humans or animals; and they are not intended to diagnose, treat, cure, mitigate or prevent any disease, or to affect the structure or any function of the body of humans or animals. Nothing in this article is preparation, handling or dosing guidance. See our full research-use terms.