Reading a Peptide Sequence: Notation, Direction and What It Tells You
One-letter and three-letter codes, why direction matters, and the four things a sequence tells you before any database lookup.
A peptide sequence written as GEPPPGKPADDAGLV and the same sequence written as Gly-Glu-Pro-Pro-Pro-Gly-Lys-Pro-Ala-Asp-Asp-Ala-Gly-Leu-Val describe the identical molecule. Knowing how to move between the two, and what a sequence tells you before you read anything else, is the most portable skill in handling this material.
Two notations, one molecule
The three-letter code is the older and more readable form. The one-letter code exists because sequences got long: a 30-residue peptide is 90 characters in three-letter form and 30 in one-letter, and databases needed the shorter one.
The one-letter assignments are not all intuitive. Some are the obvious initial — G for glycine, K for lysine is not. The awkward ones are worth memorising because they are where reading errors happen:
| Letter | Residue | Why this letter |
|---|---|---|
| K | Lysine | L was taken by leucine |
| R | Arginine | A was taken by alanine |
| W | Tryptophan | The double ring suggested a W |
| Y | Tyrosine | “tYrosine” |
| Q | Glutamine | E and G both taken |
| N | Asparagine | A taken; “asparagiNe” |
| F | Phenylalanine | “Fenylalanine” |
Direction matters
Sequences are written N-terminus to C-terminus, left to right, always. This is a convention with no exceptions in practice, and it matters because a peptide read backwards is a different molecule — the same residues in the opposite order.
This is also why numbering conventions like ACTH(4-10) are unambiguous: residues four through ten counting from the N-terminus of the parent.
What a sequence tells you immediately
- Oxidation-prone residues. M (methionine), C (cysteine) and W (tryptophan) are the three to look for. A sequence containing them needs more care with light and dissolved oxygen than one that does not.
- Deamidation candidates. N (asparagine) and Q (glutamine) convert slowly over time, changing mass by about 1 Da.
- Charge at neutral pH. K, R and H carry positive charge; D and E carry negative. A sequence heavy in one and light in the other is likely to have solubility behaviour worth checking before you commit a vial to a volume.
- Hydrophobicity. A run of I, L, V, F or W suggests a sequence that may not dissolve readily in plain aqueous solution.
A worked reading
Take BPC-157: GEPPPGKPADDAGLV. Fifteen residues. Three prolines in a row near the start, which is structurally rigid and unusually resistant to enzymatic cleavage. Two aspartates adjacent, contributing negative charge. No methionine, cysteine or tryptophan — so oxidation is less of a handling concern here than it would be for a sequence like DSIP, which opens with W.
None of that required a database lookup. That is the point of learning to read the string.
Where it connects to the certificate
The sequence is what the theoretical mass on a certificate of analysis is calculated from. If you can read the sequence, you can sanity-check that the theoretical mass is plausible — roughly 110 Da per residue as a rough average, minus 18 Da per bond for the water lost in forming it. Fifteen residues should land somewhere near 1400 Da, and BPC-157 comes in at about 1419.
That is not a precise check, but it catches the class of error where a certificate has been produced for a different compound entirely. Our guide to reading a certificate covers the rest of the document.
Every batch we supply is analysed by an independent laboratory for HPLC purity and mass-spectrometric identity, and full sequence data appears on each product page.
All products and information referenced are for in-vitro research and laboratory use only. Nothing here is medical advice, and no therapeutic claim is made or implied.