Error Bars Explained: SD, SEM and Confidence Intervals

A postdoc presents a bar chart at lab meeting: control versus peptide-exposed cells in a binding assay, each bar topped with a neat, short whisker. Someone asks what the whiskers are. Standard deviation? Standard error? A confidence interval? And what was n? The answer turns out to be standard error, with three wells pipetted from one dilution. By the end of the discussion the chart says considerably less than it appeared to, and nothing about the data has changed.

Error bars are one of the most common graphical elements in laboratory science and one of the least consistently labeled. This article explains what each type represents, why the choice matters, and what a reader should check before trusting a figure.

Start with what counts as one observation

Every error bar is calculated from a sample size, so the first question is not which formula was used but what n actually counts. Three wells filled from one dilution of one preparation are three measurements of the same thing. They capture pipetting and plate-reader scatter, not the variation you would see if the experiment were repeated from scratch.

That distinction, between technical replicates (repeat measurements of one sample) and independent replicates (separate preparations, separate days, separate cultures), sits upstream of every other issue here. Count technical replicates as independent and every bar on the chart shrinks for reasons that have nothing to do with the biology or chemistry.

Three bars, three different questions

BarHow it is calculatedQuestion it answersEffect of more data
Standard deviation (SD)Spread of individual observations around the meanHow variable is the system?Estimate gets more reliable; the spread itself does not shrink
Standard error of the mean (SEM)SD divided by the square root of nHow precisely is the mean known?Shrinks as n grows
95% confidence interval (CI)Mean plus or minus a multiplier times the SEMWhat range of values for the true mean is consistent with the data?Narrows as n grows; wide at small n

SD describes the data. SEM and the confidence interval describe the estimate of the mean. They are not interchangeable, and a figure that does not say which one is shown cannot be read correctly.

The multiplier used for a confidence interval depends on sample size. With many observations it approaches about two. With very few it is much larger: for n = 3, the multiplier for a 95% interval from the t distribution is about 4.3. That is why a confidence interval from three observations is often strikingly wide, and why it is an honest picture of how little three observations can tell you.

Why SEM is so popular

SEM is always the shortest of the three bars, smaller than SD by the square root of n. The same data look tighter when plotted with SEM. That is not automatically wrong: if the claim is about a mean, the precision of the mean is the relevant quantity. The problem arises when the bar is unlabeled, because a reader assuming SD will badly underestimate the true scatter in the system. With n = 4, SD is twice SEM; the same chart can imply two very different levels of variability depending on which the reader assumes.

Overlap is a poor significance test

Two rules of thumb circulate, and both fail. SEM bars that do not overlap do not prove that two groups differ significantly. SEM bars that do overlap do not prove that they do not. The relationship between visual overlap and any formal test depends on the sample sizes and on the test used.

Confidence intervals behave somewhat better: when two independent 95% intervals do not overlap at all, the difference would generally test as significant at the conventional level. The reverse does not hold, since intervals can overlap modestly while the difference is still significant. The most direct approach is to calculate a confidence interval for the difference between the groups, because the difference is what the claim is about.

A checklist for reading any figure

  • Which quantity is plotted? SD, SEM or CI, stated in the legend.
  • What is n, and what is one observation? Wells, plates, culture passages, days or independent preparations.
  • Are the observations independent? Technical replicates should be averaged before calculating the bar, not counted separately.
  • Is the center a mean or a median? This matters when data are skewed, as concentration-response data often are.
  • Are the individual points shown? At small n, plotting every point is more honest than any summary bar and takes no more space.

Planning enough independent replicates is also a budgeting question, because each replicate consumes reference material. Our note on budgeting peptide material for an assay series covers that side.

The same logic on an analytical report

The principle carries over to analytical documents. A purity figure reported with no indication of method variability is a point estimate presented as though it were exact. A figure printed with more decimal places than the method can support implies a precision that does not exist, as discussed in significant figures and rounding. And a difference between two documents may reflect method differences rather than material differences, which is the subject of why suppliers report different peptide purity.

At Battle Born, Across the catalog, each product’s independent reverse-phase HPLC result is published on its own page. Read that figure the same way you would read a well-labeled chart: as one measurement by one method, with the spread of the method sitting quietly behind it.

Questions

Should I plot SD or SEM?

Plot SD when you want to show how variable the observations are. Plot SEM or a confidence interval when the claim concerns the mean. Either way, say which one it is.

Is n = 3 enough for error bars?

It is enough to calculate them, but the SD from three values is itself a rough estimate. Showing the individual points is usually more informative.

Do non-overlapping error bars mean the result is significant?

Not for SEM bars. For independent 95% confidence intervals, no overlap generally does indicate a significant difference, but overlap does not rule one out.

Can technical replicates be counted in n?

No. They measure the precision of the measurement, not the variability of the experiment. Average them first, then count independent replicates.


Research use only. All products supplied by Battle Born Peptides are laboratory reference materials for in-vitro research and analytical use by qualified professionals. They are not drugs, foods, dietary supplements, cosmetics or medical devices; they are not approved by the FDA or any other regulator for use in humans or animals; and they are not intended to diagnose, treat, cure, mitigate or prevent any disease, or to affect the structure or any function of the body of humans or animals. Nothing in this article is preparation, handling or dosing guidance. See our full research-use terms.