HELM notation for peptides
Several Azulene Studio tools accept a peptide written in HELM instead of a SMILES string or a plain sequence. HELM is a text format that spells out which amino acids are in the chain, in what order, and how they are joined, including unusual joins like rings. This page is the full syntax; each tool’s own input hint carries one sentence and one example and links here.
A linear peptide
Section titled “A linear peptide”PEPTIDE1{A.G.F.K.L}$$$$V2.0That string has three parts.
PEPTIDE1 is the polymer name. It is required and HELM has no default, so a string that starts straight at { will not parse.
{A.G.F.K.L} is the monomer block: the residues in order, separated by dots. Each of the twenty standard amino acids is its single letter. Anything else goes in square brackets, for example [Aib] for alpha aminoisobutyric acid or [dF] for D phenylalanine.
$$$$V2.0 is the four section tail, described next.
The four $ sections
Section titled “The four $ sections”After the monomer block, HELM2 always has exactly four $ separators, and then the version tag. Between them sit three sections, in this order:
| Section | What goes in it |
|---|---|
| Connections | Bonds that are not the ordinary backbone chain: cyclization, disulfides, lactams |
| Polymer groups | Groupings of polymers, which Azulene Studio does not use |
| Extended annotations | Free form metadata, which Azulene Studio does not use |
V2.0 after the last separator is the version tag, and it is not optional.
A linear peptide leaves all three sections empty, so the four separators end up next to each other and the string ends in $$$$V2.0.
Cyclic, disulfide and lactam peptides
Section titled “Cyclic, disulfide and lactam peptides”Put a bond in the connections section, the first of the three:
PEPTIDE1,PEPTIDE1,<pos1>:<Rgroup>-<pos2>:<Rgroup>pos1 and pos2 are residue positions counting from 1. Each R group names the attachment point on that residue:
| R group | Attachment point |
|---|---|
R1 | Backbone N terminus |
R2 | Backbone C terminus |
R3 | Side chain |
You do not declare what kind of cyclization you want. It follows from the two attachment points you connect: joining a C terminus to an N terminus is a head to tail macrocycle, joining two cysteine side chains is a disulfide, joining a lysine side chain to an aspartate or glutamate side chain is a lactam.
# head to tail macrocycle: C terminus of residue 5 to N terminus of residue 1PEPTIDE1{A.G.L.K.F}$PEPTIDE1,PEPTIDE1,5:R2-1:R1$$$V2.0
# disulfide bridge between the side chains of the two cysteinesPEPTIDE1{A.C.D.E.C.F}$PEPTIDE1,PEPTIDE1,2:R3-5:R3$$$V2.0
# lactam bridge between a lysine and an aspartate side chainPEPTIDE1{A.K.G.E.L.D.F}$PEPTIDE1,PEPTIDE1,2:R3-6:R3$$$V2.0Why a cyclic string ends in three $, not four
Section titled “Why a cyclic string ends in three $, not four”This is the easiest thing to get wrong, so it is worth stating on its own. The number of $ separators never changes: there are always four. What changes is how many of them you can see in a row at the end of the string.
PEPTIDE1{A.G.F.K.L}$$$$V2.0 linear: nothing between the separatorsPEPTIDE1{A.G.L.K.F}$PEPTIDE1,PEPTIDE1,5:R2-1:R1$$$V2.0 cyclic: the first separator is followed by the bondBoth strings contain four $. In the linear one all four sit together. In the cyclic one the first separator is followed by the connection, so only the remaining three sit together at the end. Copying $$$$V2.0 onto the end of a string that already has a connection gives you five, and copying $$$V2.0 onto a linear string gives you three. Either way the peptide fails to build.
Which tools take HELM
Section titled “Which tools take HELM”| Tool | Field | What it is |
|---|---|---|
| Peptide 3D Structure Generation (HELM) | helm | The peptide to build a structure for |
| Absolute Hydration Free Energy | helm | The peptide solute |
| Absolute Solvation Free Energy (Non-Aqueous) | helm | The peptide solute |
| Solvent Transfer Free Energy | helm | The peptide solute |
| Lipid Bilayer Permeation Free Energy | helm | The peptide that crosses the bilayer |
| Deprotonation Free Energy (pKa) | helm_protonated, helm_deprotonated | The protonated (HA) and deprotonated (A−) states, as two separate strings |
| Protein and Peptide Developability Prediction | helm | An alternative to a plain sequence |
| Protein Mutation Folding Stability (Physics-based ΔΔG) | mutant_chain_helm | The mutated chain in full, N to C |
In every case HELM is an alternative input: give the tool either the HELM string or the SMILES or sequence field it lists alongside, not both.
Two tools read HELM differently
Section titled “Two tools read HELM differently”Most of the tools above build the peptide exactly as written. Two do not, and the difference is silent, so it is worth knowing before you submit.
Protein and Peptide Developability Prediction converts the string to a plain sequence first. Bracketed non canonical residues are removed, not substituted, so PEPTIDE1{A.[Aib].G}$$$$V2.0 is scored as AG — one residue shorter for every bracketed monomer. Cyclization is ignored, and only the first PEPTIDE1 block is read. If your peptide relies on either, this tool is not measuring the molecule you wrote.
Protein Mutation Folding Stability (Physics-based ΔΔG) infers the mutations by lining mutant_chain_helm up against the PDB chain you picked, position by position, so the string has to be the same length as that chain. It is mutually exclusive with the mutations field.
Where HELM is resolved
Section titled “Where HELM is resolved”A HELM string is turned into a molecular structure inside Azulene Studio before the job starts, so every tool that accepts HELM accepts the same dialect. There is nothing to install and no separate conversion step to run.