Peptide Hydrophobicity and HPLC Retention Time Prediction

Long before a peptide reaches a column, its sequence already says a good deal about where it will appear on a reverse-phase chromatogram. Residues with large nonpolar side chains pull a molecule toward the stationary phase; charged and polar residues keep it in the mobile phase. Turning that intuition into a number is the job of hydrophobicity scales, and the two most common tools are the GRAVY score and sets of empirically measured retention coefficients.

They are not interchangeable. GRAVY describes the average character of a sequence, while retention coefficients were fitted to actual elution data. Knowing which one answers which question keeps a prediction from being over-read.

What a GRAVY score calculates

GRAVY stands for grand average of hydropathy. Each residue is assigned a value from the Kyte–Doolittle hydropathy scale, published in 1982, and the values are summed and divided by the number of residues. Positive totals indicate a net hydrophobic sequence, negative totals a net hydrophilic one.

Residue groupKyte–Doolittle values
Most hydrophobicIle 4.5, Val 4.2, Leu 3.8, Phe 2.8, Cys 2.5
Mildly hydrophobicMet 1.9, Ala 1.8
Near neutralGly −0.4, Thr −0.7, Ser −0.8, Trp −0.9, Tyr −1.3, Pro −1.6
HydrophilicHis −3.2, Glu, Gln, Asp and Asn −3.5, Lys −3.9, Arg −4.5

The scale was built to find membrane-spanning segments in proteins, using water-to-vapor transfer energies and the observed distribution of residues between the interior and the surface of folded structures. It was never calibrated against a C18 column. That origin explains its main weakness as a chromatographic predictor, covered below.

A worked example with two pentapeptides

Leu-enkephalin (Tyr-Gly-Gly-Phe-Leu) and Met-enkephalin (Tyr-Gly-Gly-Phe-Met) differ only at the last position, which makes them a clean test case.

  • Leu-enkephalin: −1.3 − 0.4 − 0.4 + 2.8 + 3.8 = 4.5; divided by 5 gives a GRAVY of 0.90.
  • Met-enkephalin: −1.3 − 0.4 − 0.4 + 2.8 + 1.9 = 2.6; divided by 5 gives 0.52.

The prediction is that the leucine variant is retained longer, and on acidic reverse-phase conditions it is. The example works because the two sequences share everything except one residue. Comparing unrelated sequences of different lengths by GRAVY alone is far less reliable, because the score is an average and says nothing about how many hydrophobic contacts a longer chain can make at once.

Retention coefficients: scales fitted to elution data

Chromatographers took a different route. Starting around 1980, groups including Meek, Browne and colleagues, and Guo, Mant and Hodges measured the retention of many peptides of known sequence and solved for a contribution per residue. The Guo–Hodges set, determined with synthetic model peptides on C18 with aqueous trifluoroacetic acid and acetonitrile, is among the most cited. Later work by Krokhin and co-workers produced the Sequence Specific Retention Calculator, which adds corrections for length, residue position and neighboring residues.

In its simplest form the prediction is additive: the retention time under a given gradient is approximated as the sum of the residue coefficients plus a constant for the system. That works reasonably for short peptides and becomes rougher as sequences grow.

The most instructive difference from GRAVY is tryptophan. Kyte–Doolittle scores it slightly hydrophilic, because its indole nitrogen can hydrogen bond and it often sits near protein surfaces. On reverse phase, its large aromatic ring makes it one of the strongest contributors to retention, ranked alongside phenylalanine and leucine in the fitted coefficient sets. Tyrosine shows the same pattern to a lesser degree. A peptide rich in Trp can therefore have a modest GRAVY and still elute late.

Why predicted and observed elution diverge

Every coefficient set is tied to the conditions it was measured under. Several factors break the simple sum:

  • Ion pairing. Coefficients for Lys, Arg and His depend on the acid in the mobile phase. A more hydrophobic ion-pairing agent adds retention to each positive charge, so a basic peptide shifts more than a neutral one when the additive changes. Our article on TFA and formic acid as mobile-phase additives describes that dependence.
  • Secondary structure. A sequence that forms an amphipathic helix on contact with the stationary phase presents its nonpolar face as one continuous patch. It is retained more strongly than its composition suggests.
  • Length. Longer chains contact the bonded phase through only part of their surface, so each added residue contributes less than it would in a short peptide.
  • Terminal and positional effects. Residues near a charged terminus contribute differently from those in the middle, and C-terminal amidation removes a negative charge.
  • Nonstandard building blocks. D-amino acids, Aib, lipid chains and cyclic constraints have no entry in the classic tables, so any prediction for such a sequence is a qualitative estimate at most.
  • Column and temperature. Ligand density, pore size and column temperature move absolute retention, which is why coefficients predict order more dependably than minutes. The column chemistry guide covers those variables.

Using hydrophobicity to anticipate impurity positions

The most practical use of these scales is relative. When a known modification changes a sequence, the direction of the shift can usually be predicted even when its size cannot:

Change to the sequenceExpected direction on reverse phase
Loss of a Leu, Ile, Phe or Trp (deletion)Earlier than the main peak
Loss of a Gly, Ser or charged residueSmall shift, often close to the main peak
Methionine oxidized to the sulfoxideEarlier, because the sulfoxide is more polar
Residual side-chain protecting groupLater, because most protecting groups are nonpolar
Asn converted to Asp or isoAspSmall shift; direction depends on conditions

This reasoning helps when reading an impurity profile, and it links chromatography to the chemistry described in peptide synthesis impurities. It also shows why an impurity that removes a polar residue can land almost on top of the main peak, which is a problem for area-percent purity rather than something a scale can solve.

What a prediction cannot confirm

A calculated hydrophobicity is a plausibility check. If a sequence predicted to be strongly retained elutes in the first minutes of a gradient, something deserves a second look. Agreement, however, proves little: many different sequences share a similar retention, a limitation discussed in retention time as identity evidence. Mass spectrometry supplies the identity measurement that chromatography cannot.

For material we supply, the published reverse-phase HPLC result for each product lets a reader compare the observed elution with what the sequence suggests. Vials are not labeled with lot codes; each product is matched to its trace through its crimp and cap colors.

Frequently asked questions

Does a higher GRAVY score always mean a later retention time?

No. It is a reasonable guide between closely related sequences of similar length, but aromatic residues, helix formation and ion pairing can reverse the order for unrelated peptides.

Which scale predicts reverse-phase retention more accurately?

Coefficients derived from HPLC data, such as the Guo–Hodges set or SSRCalc, because they were fitted to measured elution under defined acidic conditions rather than to protein folding behavior.

Can retention coefficients be used with formic acid instead of TFA?

Only approximately. Coefficients for basic residues change with the acid, so an elution order predicted for TFA may shift under formic acid.

Why is tryptophan hydrophilic on one scale and strongly retained on another?

The hydropathy scale reflects protein surface behavior, where the indole ring often hydrogen bonds with water. On a bonded alkyl phase, its large aromatic surface dominates and drives retention.


Research use only. All products supplied by Battle Born Peptides are laboratory reference materials for in-vitro research and analytical use by qualified professionals. They are not drugs, foods, dietary supplements, cosmetics or medical devices; they are not approved by the FDA or any other regulator for use in humans or animals; and they are not intended to diagnose, treat, cure, mitigate or prevent any disease, or to affect the structure or any function of the body of humans or animals. Nothing in this article is preparation, handling or dosing guidance. See our full research-use terms.