A peptide sequence is a compact piece of notation, and small conventions carry a lot of meaning. Getting peptide sequence notation wrong produces the wrong expected mass, the wrong molarity and, occasionally, the wrong molecule ordered. This guide sets out the standard conventions, how modifications and fragments are written, where ambiguity creeps in, and how to check a supplier's sequence against a certificate of analysis.
The standard codes
Amino acids have both a three-letter and a one-letter symbol, codified by the IUPAC-IUB Joint Commission on Biochemical Nomenclature in its 1983 recommendations, which remain the reference document for symbolism in this field [1,2]. A few points that trip people up:
- The one-letter codes are not all first letters. W is tryptophan, Y tyrosine, F phenylalanine, K lysine, R arginine, Q glutamine, E glutamate, N asparagine, D aspartate.
- Ambiguity symbols exist. B means Asx (Asp or Asn), Z means Glx (Glu or Gln), X means any or unknown. B and Z appear in older literature where amide assignment was not determined; they should never appear in a product specification.
- Non-standard residues have no one-letter symbol. Norleucine (Nle), ornithine (Orn), 2-aminoisobutyric acid (Aib) and D-residues are written in three-letter form, which is why many modified analogs are published only in three-letter code.
Three-letter code with explicit hyphens (Lys-Pro-Val) is unambiguous and is preferred for short and modified sequences. One-letter code (KPV) is compact and preferred for long ones.
Direction and numbering
Sequences are written N-terminus first, left to right, and residues are numbered from 1 at the N-terminus. This convention is so entrenched that it is rarely stated, which is exactly why reversed sequences cause trouble. A retro or retro-inverso analog is normally published in its own N-to-C direction, so the string looks nothing like the parent; see D-amino acids and retro-inverso peptides.
Free termini are implied. When they are not free, they are marked:
| Notation | Meaning | Mass effect |
|---|---|---|
| H-Xaa...Xaa-OH | Free amine, free acid | Reference |
| Ac-Xaa... | N-terminal acetyl | +42.011 Da |
| ...-NH2 | C-terminal amide | -0.984 Da vs free acid |
| pGlu-... | N-terminal pyroglutamate | -17.027 Da vs Gln |
| ...-OMe | C-terminal methyl ester | +14.016 Da |
The amide convention is a frequent source of error: an amidated peptide is about 1 Da lighter than the free acid, not heavier, because OH is replaced by NH2.
Fragments and substitutions
Two shorthand forms dominate the analog literature.
Fragment notation. Parentheses after a parent name give the residue range: hGRF(1-29) is residues 1 through 29 of human growth hormone-releasing hormone; alpha-MSH(4-10) is a seven-residue internal fragment. If the fragment is amidated, that is appended: hGRF(1-29)NH2.
Substitution notation. Square brackets before the parent name list replacements with their positions, for example [Ser8]hGRF(1-29)NH2, or [Nle4,D-Phe7]alpha-MSH. Multiple substitutions go in one bracket, separated by commas. Cyclization is written with "cyclo" and the bracketed ring residues.
Reading these correctly matters because the expected monoisotopic mass follows directly from the notation. An error in reading brackets is an error in the reference mass, which propagates into identity confirmation; see mass spectrometry and peptide identity confirmation.
Stereochemistry in notation
L-configuration is the default and is not written. D-residues are marked explicitly: D-Ala in three-letter code, and by a lower-case letter (a for D-Ala) in one-letter code, though the lower-case convention is not universal and should be defined wherever it is used. Unnatural residues such as Aib are achiral and carry no prefix.
Because stereochemistry is invisible to mass spectrometry, notation is often the only record of it short of chiral analysis. This is one reason short sequences with a single D-substitution should always be ordered and recorded in three-letter form.
A three-letter worked example
Pinealon is a tripeptide usually written Glu-Asp-Arg, or EDR in one-letter code. It is one of the short peptide bioregulators developed in St Petersburg and is discussed under that name across a Russian-language literature, which is itself a nomenclature problem: the same molecule appears as Pinealon, as EDR peptide and as the tripeptide Glu-Asp-Arg depending on the source. A review of this peptide's proposed mechanisms uses the EDR designation throughout and describes gene-expression and protein-synthesis effects reported in cell-culture and animal studies [3]; a biophysical study of the same tripeptide, using spectroscopy, NMR, viscosimetry and molecular dynamics, examined its interaction with DNA and reported partial penetration into the major groove with contacts at guanine N7 and O6, modulated by magnesium ions [4].
For the notation-minded, this short sequence demonstrates several conventions at once:
- Both acidic residues carry side-chain carboxyls, so the free peptide is strongly anionic apart from the Arg guanidinium. Writing it as EDR gives no indication of charge state or salt form; the certificate must specify the counterion.
- Direction is load-bearing even here. Glu-Asp-Arg and Arg-Asp-Glu are different molecules with identical mass and identical composition. A three-residue string is short enough that transcription errors go unnoticed.
- Free acid versus amide. The free-acid form and a C-terminal amide differ by 0.984 Da, a difference well within the resolving power of any modern instrument but easy to overlook when comparing a supplier's mass to a published one.
How a stated purity figure relates to what the notation describes is covered in peptide purity grades explained.
Common errors, and how to catch them
- Confusing the amide and free acid. Check for -NH2 in the name and expect a mass about 1 Da lower.
- Leu/Ile ambiguity. Identical masses; no standard MS method distinguishes them.
- Gln vs Lys. Differ by 0.036 Da; only high-resolution instruments separate them.
- Reading a reversed sequence as a parent. Always confirm direction for retro analogs.
- Dropping a D-prefix in transcription. The mass is unchanged, so nothing downstream catches it.
- Using B, Z or X in a specification. These are ambiguity symbols [1] and have no place on a certificate.
A simple habit prevents most of these: recompute the expected monoisotopic mass from the written sequence and compare it with the value on the certificate before using the material. Per-lot certificates and spectra for catalogue peptides are posted on our lab reports page, and impurity masses that appear alongside the main peak are catalogued in common peptide impurities.
Key takeaways
- IUPAC-IUB recommendations define the one- and three-letter symbols; B, Z and X denote ambiguity and should not appear in specifications.
- Sequences run N to C, numbered from 1 at the N-terminus; reversed analogs need an explicit statement of direction.
- Terminal modifications carry defined mass shifts; amidation makes a peptide about 1 Da lighter than the free acid.
- Fragment parentheses and substitution brackets determine the expected mass and must be read exactly.
- Stereochemistry is invisible to mass spectrometry, so the written record is the primary evidence of D-residues.
This article summarizes published research for informational purposes. All Ascent Sciences products are for laboratory research use only and are not for human or animal consumption.
References
- IUPAC-IUB Joint Commission on Biochemical Nomenclature (JCBN). Nomenclature and symbolism for amino acids and peptides. Recommendations 1983. Biochem J. 1984;219(2):345-373. PubMed
- IUPAC-IUB Joint Commission on Biochemical Nomenclature (JCBN). Nomenclature and symbolism for amino acids and peptides. Recommendations 1983. Eur J Biochem. 1984;138(1):9-37. PubMed
- Khavinson V, Linkova N, Kozhevnikova E, Trofimova S. EDR peptide: possible mechanism of gene expression and protein synthesis regulation involved in the pathogenesis of Alzheimer's disease. Molecules. 2020;26(1):159. PubMed
- Silanteva IA, Komolkin AV, Morozova EA, et al. Role of mono- and divalent ions in peptide Glu-Asp-Arg-DNA interaction. J Phys Chem B. 2019;123(9):1896-1902. PubMed
- D'Hondt M, Bracke N, Taevernier L, et al. Related impurities in peptide medicines. J Pharm Biomed Anal. 2014;101:2-30. PubMed
Frequently asked questions
Which direction is a peptide sequence written in?
By convention, N-terminus on the left, C-terminus on the right. Residue numbering starts at 1 at the N-terminus. Any deviation should be stated explicitly, which matters particularly for reversed-sequence analogs.
What does a bracketed residue such as [D-Ala2] mean?
It denotes a substitution relative to a named parent sequence: in this case, the residue at position 2 of the parent has been replaced by D-alanine. Multiple substitutions are listed inside one set of brackets.
How are cyclic peptides written?
Usually with 'cyclo' followed by the bracketed residues that form the ring, for example cyclo[Asp-His-D-Phe-Arg-Trp-Lys]. A disulfide is often shown by naming the bridged cysteines, such as Cys1-Cys6.
Why are Leu and Ile both problematic in sequence confirmation?
They are isomers with identical mass, so intact mass and standard fragmentation cannot distinguish them. Distinguishing the two requires specialised fragmentation methods or amino acid analysis combined with independent evidence.
All Ascent Sciences products are for laboratory research use only and are not for human or animal consumption. This article summarizes published research and is not medical advice. See our Research Use Agreement.