Skip to content

Docking

Docking

Fit a small molecule into a protein pocket and score how well it sits there.

Docking generates and ranks noncovalent ligand poses within a specified binding site. Accepted ligand inputs include small molecules represented as SMILES, peptides represented in HELM notation, and uploaded ligand structures. Macrocycles and cyclic peptides can be processed with macrocycle-specific sampling.

The required binding_site_center input defines the coordinate that directs the docking search.

Ligands that form a covalent bond with the protein should be submitted to Covalent Docking.

Docking is appropriate for initial assessment of ligand, fragment, or cofactor occupancy within a known binding site and for ranking candidates by predicted fit. Its computational cost is substantially lower than that of a free energy calculation. Quantitative affinity predictions for selected compounds can subsequently be obtained with Absolute Binding Free Energy (ABFE) or Relative Binding Free Energy (RBFE).

InputRequiredWhat it is
structure_fileyesProtein structure in PDB or CIF format.
drug_smilesone of these threeSMILES string for the ligand or cofactor.
helmone of these threeCyclic or linear peptide in HELM2 notation, for example PEPTIDE1{A.G.F.K.L}$$$$V2.0. Ring closures must be declared in the connection section. See the HELM reference.
ligand_fileone of these threeUploaded ligand structure.
chain_idyesProtein chain containing the binding site, for example A.
binding_site_centeryesBinding-site centre specified as [x,y,z] in Angstroms, using the coordinate frame of the uploaded structure. This coordinate directs the docking search.
binding_site_residuesnoBinding-site residues specified as [name, number] pairs, for example [["SER",64],["HIS",57]]. Studio calculates their mean coordinates and assigns the result to binding_site_center. This field does not independently direct the docking search. Residue numbers must correspond to those in the uploaded structure. Matching includes every occurrence of each residue number across all chains.
macrocycle_samplingno, default autoSelection of macrocycle and cyclic-peptide sampling. auto activates macrocycle-specific sampling for eligible ring systems. This setting incurs greater computational cost than small-molecule docking. off applies the small-molecule settings to macrocycles.
keep_cofactorsnoComma-separated cofactor residue names to retain, for example ZN or ZN,NDP, specified using PDB chemical component IDs. An empty value removes all nonstandard residues. Waters are removed unless HOH is specified, which retains all waters. In this field, CA denotes calcium, not the C-alpha atom.

Jobs can be submitted through Azulene Studio, the Python SDK, or the CLI. Installation, authentication, and a worked example are documented on the Get started page. Studio can calculate the binding-site centre from specified residues. SDK and CLI submissions require binding_site_center directly.

Select Docking from the tools list. In the Inputs and Parameters step, upload the protein structure and specify the ligand as SMILES, HELM, or an uploaded file. Enter the chain ID and either the binding-site centre or the residues defining the site. Studio calculates the centre from the specified residues. Select Review and Submit to submit the job.

from azulene import jobs
result = jobs.submit(
job_type="docking",
input_data={
"structure_file": "/path/to/your/protein.pdb",
"drug_smiles": "c1ccc(cc1)C(=N)N",
"chain_id": "A",
"binding_site_center": "[12.5, 8.3, -4.1]",
},
)

Pass the inputs as a JSON string.

Terminal window
azulene jobs submit --job-type docking \
--input-data '{"structure_file": "/path/to/your/protein.pdb", "drug_smiles": "c1ccc(cc1)C(=N)N", "chain_id": "A", "binding_site_center": "[12.5, 8.3, -4.1]"}'

The primary result is unified_score_dG_kcal, a predicted binding free energy in kcal/mol. unified_score_std reports the associated spread. More negative values indicate stronger predicted binding.

binding_site_distance_a reports the distance between the docked ligand and the specified binding-site centre. This value confirms whether the predicted pose occupies the intended site.

Pose-level results include top_k_labels, top_k_composite_kcal_mol, top_k_interaction_kcal_mol, top_k_ligand_strain_kcal_mol, and top_k_smina_kcal_mol, with energetic quantities reported in kcal/mol. Ligand strain should be assessed together with the pose score because highly strained poses may be unsuitable for subsequent analysis despite favourable scores.

For macrocycles and cyclic peptides, is_macrocycle, macrocycle_ring_size, and n_ring_confs report macrocycle classification, ring size, and the number of sampled ring conformers. ligand_source identifies the ligand input used for the calculation.

Output structures and results are returned as URLs in protein_pdb_url, ligand_pose_pdb_url, ligand_sdf_url, complex_pdb_url, assembled_complex_pdb_url, and result_json_url.

The binding-site centre directs the docking search and should be verified against the coordinate frame of the uploaded structure before submission. Structural cofactors, including catalytic metals and hemes, should be retained with keep_cofactors when they define the binding-site environment. Docking supports rapid relative ranking of candidate compounds. Quantitative affinity prediction for selected compounds requires a subsequent free energy calculation.