Proficiency Testing and Inter-Laboratory Comparisons Explained

A laboratory can validate its methods, calibrate its instruments and pass every system suitability test and still be consistently wrong. Internal checks compare the laboratory with itself; they cannot reveal a bias that is built into its reference material, its integration habits or its understanding of a method. The only way to find that kind of error is to measure the same material as other laboratories and compare. Proficiency testing and inter-laboratory comparisons are the formal versions of that exercise.

This article explains how the schemes work, how results are scored, and what a purchaser can sensibly ask about them when relying on a laboratory’s numbers.

Three kinds of inter-laboratory study

The term inter-laboratory comparison covers any organized measurement of the same or similar items by two or more laboratories. Within it, three designs serve different purposes:

DesignQuestion it answersWhat is being tested
Proficiency testing (PT)Is this laboratory performing competently?The laboratory
Collaborative method studyHow reproducible is this method across laboratories?The method
Reference material characterizationWhat is the best estimate of this material’s value?The material

In proficiency testing, each participant uses its own routine method, exactly as it would for a customer sample. That is the point: the scheme evaluates everyday performance, not a special effort.

How a proficiency testing round runs

A provider prepares a homogeneous material, confirms that it is stable for the duration of the round, and sends portions to participating laboratories. Each laboratory analyzes its portion by its usual procedure within a deadline and reports the result, usually without knowing the expected value. The provider then compares every result with an assigned value and issues a confidential report to each participant, typically identifying laboratories only by code.

Providers of such schemes can themselves be accredited against ISO/IEC 17043, and the statistical treatment most schemes follow is described in ISO 13528. For laboratories accredited to ISO/IEC 17025, participation in PT or other comparisons, where available, is one of the expected ways of monitoring the validity of results. For how that accreditation works, see our piece on ISO 17025 accreditation for peptide testing.

Assigned values and z-scores in proficiency testing

The assigned value is the provider’s best estimate of the true value. It may come from a reference laboratory using a high-accuracy method, from formulation of the material with known amounts, or from a robust consensus of the participants’ results. Robust statistics are used for consensus values so that a few extreme results do not drag the estimate.

Each laboratory’s performance is then commonly expressed as a z-score:

z = (x − X) / σpt

Here x is the laboratory’s result, X the assigned value and σpt the standard deviation for proficiency assessment, which the provider sets to reflect fitness for purpose. The conventional interpretation is that |z| ≤ 2 is satisfactory, 2 < |z| < 3 is questionable, and |z| ≥ 3 is unsatisfactory and calls for investigation.

A worked example: suppose a scheme distributes a peptide sample with an assigned HPLC purity of 97.8% and a σpt of 0.5 percentage points. A laboratory reporting 98.2% scores z = 0.8. One reporting 98.9% scores z = 2.2, a warning signal. One reporting 96.2% scores z = −3.2, an action signal that the laboratory must investigate and document.

When the scheme uses En numbers instead

Some comparisons, especially those involving calibration or reference measurements, score results with an En number, which takes each participant’s stated uncertainty into account:

En = (x − X) / √(Ulab² + Uref²)

Here U is the expanded uncertainty of each value. |En| ≤ 1 is normally satisfactory. This score penalizes a laboratory that claims a smaller uncertainty than its results justify, which makes it a useful check on the uncertainty statements discussed in measurement uncertainty in peptide purity.

What a laboratory does with a poor score

An unsatisfactory result is not a failure of accreditation in itself; what matters is the response. A competent laboratory treats it like any nonconformity: it checks for transcription and calculation errors, reviews the raw data and integration, examines the calibration and reference material, and considers whether routine customer results reported in the same period could be affected. The investigation, root cause and corrective action are recorded, and a repeated pattern of questionable scores is treated as a trend even when no single result crossed the action limit. Many providers also publish the spread of all participants’ results grouped by method, which lets a laboratory see whether its whole method family runs high or low compared with other approaches.

Where peptide purity testing fits

Formal PT schemes are common for food, water and environmental analytes but scarce for the purity of individual synthetic peptides. Laboratories working in this area therefore rely more on other comparisons: split samples exchanged with a partner laboratory, repeat analysis of a characterized reference standard, and structured method comparisons of the kind described in HPLC method transfer between laboratories. A laboratory’s ability to show such evidence, together with the method validation behind its figures, set out in analytical method validation for peptide purity, says more than a logo on its report.

For a purchaser, reasonable questions are which comparisons the laboratory takes part in, how recent its last round was, and whether its scores for comparable analytes were satisfactory. The scores themselves are usually confidential, but a summary is often available on request. Battle Born uses independent laboratories for the reverse-phase HPLC result shown on each product listing; no lot numbers are used, and the crimp and cap colors on a vial point to the result that applies.

Frequently asked questions

What is the difference between proficiency testing and an inter-laboratory comparison?

Proficiency testing is one type of inter-laboratory comparison, used specifically to evaluate each participant’s performance against predetermined criteria.

Is a z-score of 2.5 a failure?

It is conventionally questionable rather than unsatisfactory. A single score in that zone warrants review; repeated scores there usually warrant investigation.

Can a laboratory assign its most experienced analyst to a PT sample?

It should not. PT samples are meant to be handled like routine work, and doing otherwise defeats the purpose of the exercise.


Research use only. All products supplied by Battle Born Peptides are laboratory reference materials for in-vitro research and analytical use by qualified professionals. They are not drugs, foods, dietary supplements, cosmetics or medical devices; they are not approved by the FDA or any other regulator for use in humans or animals; and they are not intended to diagnose, treat, cure, mitigate or prevent any disease, or to affect the structure or any function of the body of humans or animals. Nothing in this article is preparation, handling or dosing guidance. See our full research-use terms.