Covalent Docking
Covalent Docking generates and ranks poses of an inhibitor bonded to a specified residue. Supported warhead classes include boronic acids, Michael acceptors, chloroacetamides, nitriles, aldehydes, epoxides, sulfonyl fluorides and phosphonates.
Each submission must specify five inputs that define the covalent bond: the chain, residue name, residue number, target_atom and covalent_element.
When to use it
Section titled “When to use it”Use Covalent Docking for inhibitors that form a covalent bond with a known residue. The output includes the bound geometry and pose ranking. Use Docking for reversible binders. Use Cofactor & Ligand Docking to dock multiple molecules sequentially into the same site through covalent and non-covalent steps.
Inputs
Section titled “Inputs”| Input | Required | What it is |
|---|---|---|
structure_file | yes | Protein structure in PDB or CIF format. |
drug_smiles | yes | SMILES string of the inhibitor. The structure must contain a reactive warhead. |
chain_id | yes | Chain containing the target residue, for example A. |
target_resname | yes | Three-letter name of the attacked residue. One of SER, CYS, LYS, THR, HIS, TYR. |
target_resid | yes | Sequence number of the target residue within the chain. |
target_atom | yes | Name of the reactive atom on the target residue, for example OG for serine or SG for cysteine. This value is specified independently of the residue name. |
covalent_element | yes | Element of the ligand atom that forms the new bond. See Choosing the reacting element. |
keep_cofactors | no | Comma-separated cofactor residue names to retain, for example ZN or ZN,NDP, using PDB chemical component IDs. An empty value removes every non-standard residue. Catalytic metals should be retained to preserve active-site geometry. Waters are removed with the other non-standard residues. The value HOH retains all waters as a single residue class. In this field, CA denotes calcium, not the C-alpha atom. |
extra_chains | no | Comma-separated additional chain IDs to include, for example B or B,C. |
Choosing the reacting element
Section titled “Choosing the reacting element”Identify the intended covalent bond and specify the element of the ligand atom that participates in it. This atom may differ from the most prominent heteroatom in the warhead. For a vinyl sulfone, the residue attacks the beta carbon, so covalent_element is C.
| Warhead on your ligand | Bond it forms | covalent_element |
|---|---|---|
| Acrylamide, vinyl ketone, vinyl sulfone, maleimide and other Michael acceptors | C-S, C-O | C |
| Chloroacetamide, halomethyl ketone, acyl halide | C-S | C |
| Nitrile | C-S | C |
| Aldehyde, alpha-ketoamide | C-S, C-O | C |
| Epoxide, aziridine, beta-lactam, beta-lactone | C-S, C-O, C-N | C |
| Boronic acid, benzoxaborole | B-O, B-N | B |
| Sulfonyl fluoride, sulfinamide, disulfide | S-O, S-N, S-S | S |
| Phosphonate, organophosphate | P-O | P |
Specify B, S or P when the attacked ligand atom is boron, sulfur or phosphorus, respectively.
Naming the target atom
Section titled “Naming the target atom”target_atom specifies the PDB atom name of the nucleophile on the selected residue.
| Residue | target_atom |
|---|---|
SER | OG |
CYS | SG |
THR | OG1 |
LYS | NZ |
TYR | OH |
HIS | NE2, or ND1 if that is the reactive tautomer in your structure |
How to run it
Section titled “How to run it”Jobs may be submitted through Azulene Studio, the Python SDK or the CLI. The Get started page describes installation, authentication and a worked example. Every covalent submission must specify target_atom and covalent_element explicitly.
In Azulene Studio
Section titled “In Azulene Studio”Open Covalent Docking from the tools list. In the Inputs and Parameters step, upload the protein structure and enter the inhibitor SMILES. Specify the chain ID and the target residue name, number and atom. Confirm the reacting element, then select Review and Submit.
From the Python SDK
Section titled “From the Python SDK”from azulene import jobs
result = jobs.submit( job_type="covalent_docking", input_data={ "structure_file": "/path/to/your/protein.pdb", "drug_smiles": "OB(O)c1ccccc1", "chain_id": "A", "target_resname": "SER", "target_resid": 70, "target_atom": "OG", "covalent_element": "B", },)From the CLI
Section titled “From the CLI”Pass the inputs as a JSON string.
azulene jobs submit --job-type covalent_docking \ --input-data '{"structure_file": "/path/to/your/protein.pdb", "drug_smiles": "OB(O)c1ccccc1", "chain_id": "A", "target_resname": "SER", "target_resid": 70, "target_atom": "OG", "covalent_element": "B"}'Reading the result
Section titled “Reading the result”The primary reported metric is unified_score_dG_kcal, a predicted binding free energy in kcal/mol. More negative values indicate stronger predicted binding. unified_score_std reports the spread of the prediction. Larger values indicate greater uncertainty.
Review the following fields before interpreting the score:
is_covalentandconfidenceindicate whether the covalent route was executed. Aconfidencevalue ofnoncovalent_low_confidenceindicates that it was not executed.feature_completenessreports how much of the model input was derived from the submitted structure. Lower values indicate greater use of default values.warningsreports detected conditions that may affect interpretation, including submission of a covalent anchor when no warhead was identified in the SMILES.
warhead_bond_type identifies the bond formed, for example B-O for a boronic acid on a threonine or C-S for a Michael acceptor on a cysteine. The reported bond type should correspond to the applicable row in the table above. An unintended bond type indicates an incorrect reacting element.
covalent_distance_a reports the distance in Angstroms between the ligand and the specified target atom in the top-ranked pose. This value indicates whether the warhead reached the target atom.
Each pose has an associated top_k_composite_kcal_mol score in kcal/mol. More negative values indicate more favorable poses within the same run.
Structure and result files are returned through protein_pdb_url, ligand_pose_pdb_url, ligand_sdf_url, complex_pdb_url, poses_zip_url and result_json_url. The complex is a multi-model file containing one model per pose.
The SMILES must contain a reactive warhead. The chain ID, residue name, residue number and atom name must match the submitted structure. Mismatched identifiers can produce an incorrect result. Structural cofactors, including catalytic zinc, should be retained with keep_cofactors.