Skip to content

Covalent Docking

Covalent Docking

Dock an inhibitor that bonds to a chosen residue, and score the fit in the active site.

Covalent Docking generates and ranks poses of an inhibitor bonded to a specified residue. Supported warhead classes include boronic acids, Michael acceptors, chloroacetamides, nitriles, aldehydes, epoxides, sulfonyl fluorides and phosphonates.

Each submission must specify five inputs that define the covalent bond: the chain, residue name, residue number, target_atom and covalent_element.

Use Covalent Docking for inhibitors that form a covalent bond with a known residue. The output includes the bound geometry and pose ranking. Use Docking for reversible binders. Use Cofactor & Ligand Docking to dock multiple molecules sequentially into the same site through covalent and non-covalent steps.

InputRequiredWhat it is
structure_fileyesProtein structure in PDB or CIF format.
drug_smilesyesSMILES string of the inhibitor. The structure must contain a reactive warhead.
chain_idyesChain containing the target residue, for example A.
target_resnameyesThree-letter name of the attacked residue. One of SER, CYS, LYS, THR, HIS, TYR.
target_residyesSequence number of the target residue within the chain.
target_atomyesName of the reactive atom on the target residue, for example OG for serine or SG for cysteine. This value is specified independently of the residue name.
covalent_elementyesElement of the ligand atom that forms the new bond. See Choosing the reacting element.
keep_cofactorsnoComma-separated cofactor residue names to retain, for example ZN or ZN,NDP, using PDB chemical component IDs. An empty value removes every non-standard residue. Catalytic metals should be retained to preserve active-site geometry. Waters are removed with the other non-standard residues. The value HOH retains all waters as a single residue class. In this field, CA denotes calcium, not the C-alpha atom.
extra_chainsnoComma-separated additional chain IDs to include, for example B or B,C.

Identify the intended covalent bond and specify the element of the ligand atom that participates in it. This atom may differ from the most prominent heteroatom in the warhead. For a vinyl sulfone, the residue attacks the beta carbon, so covalent_element is C.

Warhead on your ligandBond it formscovalent_element
Acrylamide, vinyl ketone, vinyl sulfone, maleimide and other Michael acceptorsC-S, C-OC
Chloroacetamide, halomethyl ketone, acyl halideC-SC
NitrileC-SC
Aldehyde, alpha-ketoamideC-S, C-OC
Epoxide, aziridine, beta-lactam, beta-lactoneC-S, C-O, C-NC
Boronic acid, benzoxaboroleB-O, B-NB
Sulfonyl fluoride, sulfinamide, disulfideS-O, S-N, S-SS
Phosphonate, organophosphateP-OP

Specify B, S or P when the attacked ligand atom is boron, sulfur or phosphorus, respectively.

target_atom specifies the PDB atom name of the nucleophile on the selected residue.

Residuetarget_atom
SEROG
CYSSG
THROG1
LYSNZ
TYROH
HISNE2, or ND1 if that is the reactive tautomer in your structure

Jobs may be submitted through Azulene Studio, the Python SDK or the CLI. The Get started page describes installation, authentication and a worked example. Every covalent submission must specify target_atom and covalent_element explicitly.

Open Covalent Docking from the tools list. In the Inputs and Parameters step, upload the protein structure and enter the inhibitor SMILES. Specify the chain ID and the target residue name, number and atom. Confirm the reacting element, then select Review and Submit.

from azulene import jobs
result = jobs.submit(
job_type="covalent_docking",
input_data={
"structure_file": "/path/to/your/protein.pdb",
"drug_smiles": "OB(O)c1ccccc1",
"chain_id": "A",
"target_resname": "SER",
"target_resid": 70,
"target_atom": "OG",
"covalent_element": "B",
},
)

Pass the inputs as a JSON string.

Terminal window
azulene jobs submit --job-type covalent_docking \
--input-data '{"structure_file": "/path/to/your/protein.pdb", "drug_smiles": "OB(O)c1ccccc1", "chain_id": "A", "target_resname": "SER", "target_resid": 70, "target_atom": "OG", "covalent_element": "B"}'

The primary reported metric is unified_score_dG_kcal, a predicted binding free energy in kcal/mol. More negative values indicate stronger predicted binding. unified_score_std reports the spread of the prediction. Larger values indicate greater uncertainty.

Review the following fields before interpreting the score:

  • is_covalent and confidence indicate whether the covalent route was executed. A confidence value of noncovalent_low_confidence indicates that it was not executed.
  • feature_completeness reports how much of the model input was derived from the submitted structure. Lower values indicate greater use of default values.
  • warnings reports detected conditions that may affect interpretation, including submission of a covalent anchor when no warhead was identified in the SMILES.

warhead_bond_type identifies the bond formed, for example B-O for a boronic acid on a threonine or C-S for a Michael acceptor on a cysteine. The reported bond type should correspond to the applicable row in the table above. An unintended bond type indicates an incorrect reacting element.

covalent_distance_a reports the distance in Angstroms between the ligand and the specified target atom in the top-ranked pose. This value indicates whether the warhead reached the target atom.

Each pose has an associated top_k_composite_kcal_mol score in kcal/mol. More negative values indicate more favorable poses within the same run.

Structure and result files are returned through protein_pdb_url, ligand_pose_pdb_url, ligand_sdf_url, complex_pdb_url, poses_zip_url and result_json_url. The complex is a multi-model file containing one model per pose.

The SMILES must contain a reactive warhead. The chain ID, residue name, residue number and atom name must match the submitted structure. Mismatched identifiers can produce an incorrect result. Structural cofactors, including catalytic zinc, should be retained with keep_cofactors.