A research group tests two lots of the same peptide in a cell-based assay. Lot B produces a slightly lower signal than lot A, and the software reports p = 0.04. Someone writes “lot B is significantly less active” on the slide, and a purchasing decision starts to form around it. Before anyone acts, it is worth asking exactly what that 0.04 was calculated to say, because it answers a much narrower question than the sentence on the slide suggests.
The question a p-value actually answers
A p-value starts from an assumption, the null hypothesis, which in this example says the two lots perform identically and any difference is noise. Under that assumption, the p-value is the probability of seeing a difference at least as large as the one observed, purely from random variation. That is all it is.
The direction of the conditional is the whole story. The calculation gives the probability of the data, given that the null is true. What people usually want is the probability that the hypothesis is true, given the data. Those two quantities are not interchangeable, and there is no way to convert one into the other without further information that the test never used. So p = 0.04 does not mean a 4% chance the result is a fluke, and it does not mean 96% confidence that lot B really differs.
Common readings and what is wrong with them
| What people say | What is actually the case |
|---|---|
| “p = 0.04, so there is a 4% chance the null is true.” | The calculation assumed the null was true from the start; it cannot also estimate that probability. |
| “Below 0.05 is real, above 0.05 is not.” | 0.05 is a convention. Nothing meaningful separates 0.049 from 0.051. |
| “Not significant, so there is no difference.” | A small or noisy study cannot distinguish “no effect” from “effect too small to detect here.” |
| “p = 0.001 means a bigger effect than p = 0.03.” | The p-value mixes effect size with sample size and noise. A tiny, precisely measured difference can give a very small p-value. |
How many tests stood behind the one you see
A 0.05 threshold means that, when nothing real is happening, about one comparison in twenty will still come out “significant” by chance. Run twenty independent comparisons of pure noise and the probability that at least one crosses the line is roughly 64% (one minus 0.95 to the twentieth power). Nobody needs to behave badly for this to happen; it is simply how thresholds work.
The trouble arrives when many comparisons are performed and only the ones that crossed the threshold are reported. Peptide experiments easily multiply comparisons without anyone noticing: several concentrations, several time points, several readouts, each pair a separate test. A reader who sees only the final significant result has no way to judge it without knowing the denominator.
Formal multiple-comparison corrections exist, but they reduce power and can be blunt. The more useful habit is to decide the primary comparison before collecting data, and to label everything else as exploratory. That also makes the result something another laboratory can meaningfully try to reproduce.
Small studies inflate the effects they find
There is a quieter bias that affects honest work too. In an underpowered study, a modest true effect will usually fail to reach significance. The occasions when it does cross the threshold tend to be those where random variation happened to push the observed effect upward. If significant results are the ones that get noticed and written up, the effects on record will be systematically larger than the true ones.
Laboratory peptide research is particularly exposed: sample sizes are often small, experiments track several readouts at once, and studies tend to appear individually rather than as coordinated series. Each of those features raises the share of positive findings that later fail to replicate. The practical implications for planning are discussed in planning a peptide budget across an assay series.
Report the size and the precision instead
The two things a reader really needs are how big the difference was and how precisely it was estimated. An effect size with a confidence interval supplies both. In the lot comparison, a statement such as “lot B signal 8% lower, 95% interval 1% to 15%” tells you the difference could be trivial or could be meaningful, which is far more honest than a bare verdict of significance. When an interval is wide and spans zero, it states plainly that the study did not settle the question.
This mirrors good practice in analytical chemistry, where a purity figure means more when the method and its variability are disclosed. The same thinking runs through why suppliers report different peptide purity. Before blaming a lot, it is also worth confirming the material itself is what you think it is; purity versus net peptide content explains why two vials with the same label mass can carry different amounts of peptide, which alone can shift an assay signal.
A short checklist before accepting a significant result
- Is an effect size reported, not only a p-value?
- Does it come with a confidence interval, and how wide is that interval?
- How many comparisons were run in total, including those not shown?
- Was the sample size decided before the data were collected?
- Was the analysis planned in advance, or chosen after seeing the numbers?
- Are the replicates independent, or repeated readings of the same preparation?
A result that passes these questions is worth building on. One that offers only “p < 0.05” with no magnitude and no count of tests has given a conclusion while holding back the evidence for it. For definitions of related terms, see the research peptide glossary.
Questions
Is a p-value useless, then?
No. It does indicate how surprising a pattern would be under one specific null model. It is simply the least informative input into a judgment, and it should sit behind the effect size and interval rather than in front of them.
Why is 0.05 used at all?
Largely by convention and habit. It is a reasonable default for flagging results worth attention, but it is not a boundary between real and imaginary effects.
Does a non-significant result prove two lots are equivalent?
No. Showing equivalence requires defining an acceptable margin in advance and demonstrating that the confidence interval falls within it.
Should I apply a multiple-comparison correction every time?
It depends on the design. Naming a single primary comparison up front often serves better than correcting a large set of unplanned tests after the fact.
Research use only. All products supplied by Battle Born Peptides are laboratory reference materials for in-vitro research and analytical use by qualified professionals. They are not drugs, foods, dietary supplements, cosmetics or medical devices; they are not approved by the FDA or any other regulator for use in humans or animals; and they are not intended to diagnose, treat, cure, mitigate or prevent any disease, or to affect the structure or any function of the body of humans or animals. Nothing in this article is preparation, handling or dosing guidance. See our full research-use terms.