Skip to content

PROTAC Ternary Complex Pose Prediction

PROTAC Ternary Complex Pose Prediction

Rigid-body ternary pose search, GPU minimization, and a trained pose ranker.

This tool predicts ternary complex poses for a PROTAC, its target protein, and an E3 ligase. Supply two binary co-complexes: the target with its warhead bound, and the E3 ligase with its recruiter bound. Also supply the linker as a SMILES fragment and identify the two attachment atoms. The job returns a ranked set of full ternary poses with the PROTAC built and relaxed in place.

The two binary interfaces remain fixed. The output shows where the E3 ligase can sit relative to the target within the reach of the linker. Binary docking does not provide this geometry.

Use this tool when both binary binding modes are known or modelled, but the ternary geometry is not. The predicted poses show whether a linker permits a compatible protein-protein arrangement and what that arrangement looks like. Comparing poses across a linker series shows how linker length and composition affect the predicted geometry. These results can help prioritize designs before synthesis.

Supply a known ternary structure as ground truth to benchmark the predictions. The job reports RMSDs and writes overlay structures.

Scope:

  • Validation is concentrated on VHL systems, because that is where most published ternary structures are. Other E3 ligases, CRBN in particular, have less supporting data behind them.
  • Both proteins are held rigid. Induced fit at the ternary interface is not modelled.
InputRequiredWhat it is
target_pdb_contentyesPDB file of the target protein with its warhead bound.
e3_pdb_contentyesPDB file of the E3 ligase with its recruiter bound.
linker_smilesyesLinker SMILES fragment with two attachment points, [*:1] and [*:2], marking the bonds to the warhead and the recruiter.
warhead_attach_atom_idxyes0-based atom index in the warhead that bonds to the linker.
recruiter_attach_atom_idxyes0-based atom index in the E3 recruiter that bonds to the linker.
target_chainno, default AChain ID or comma-separated chain IDs of the target protein. Use several for a multimeric target.
e3_chain_labelno, default BChain ID or comma-separated chain IDs of the E3 complex. VHL with Elongin C and Elongin B is typically B,C,D.
warhead_chainno, default XChain ID of the warhead ligand in the target PDB.
recruiter_chainno, default YChain ID of the recruiter ligand in the E3 PDB.
presetno, default mediumAngular resolution of the rigid-body scan. See the preset table below.
n_prescan_chunksno, default 8Number of distance bins scanned in parallel. Range 1 to 16. Increasing it shortens the wall time of the scan stage. Each chunk adds about 30 seconds of setup, so do not exceed the number of distance bins.
warhead_smilesnoSMILES template for the warhead with correct bond orders. It may also include an attachment point. Supply it when bond orders cannot be perceived reliably from the PDB coordinates alone.
e3_anchor_smilesnoThe same for the E3 recruiter.
gt_complex_pdb_contentnoExperimental ternary complex PDB for RMSD scoring and overlays.
gt_ligand_resnamenoResidue name of the PROTAC in the ground-truth file. Required when ground truth is supplied.
gt_target_chainno, default ATarget chain ID in the ground-truth file.
gt_e3_chainno, default BE3 chain ID or IDs in the ground-truth file.
keep_dirsno, default truePersist poses, overlays, and the summary as a downloadable ZIP. Set false to return only the scalar JSON summary.
PresetRelative costUse it for
quickAbout a third of mediumA fast, coarse pass. It gives only a rough view of the pose landscape. Use it as a demonstration rather than a prediction.
mediumReferenceThe default setting and the setting matched to the ranker. Start here.
longSeveral times mediumA finer, wider search. Use it when medium returns no pose with a plausible interface. This can happen when the E3 sits far around the target.

Submit the job from Azulene Studio, the Python SDK, or the CLI. Local file paths are uploaded automatically. The Get started page covers installation, login, and a worked example.

Open PROTAC Ternary Complex Pose Prediction from the tools list. On the Inputs and Parameters step, upload the target and E3 PDB files. Enter the linker SMILES with its two attachment points. Set the two attachment atom indices. Check that the protein and ligand chain IDs match your files. Leave the preset at medium, then select Review and Submit.

from azulene import jobs
result = jobs.submit(
job_type="protac_pose_prediction",
input_data={
"target_pdb_content": "target_with_warhead.pdb",
"e3_pdb_content": "vhl_with_recruiter.pdb",
"linker_smiles": "[*:1]CCOCCOCC[*:2]",
"warhead_attach_atom_idx": 12,
"recruiter_attach_atom_idx": 7,
"target_chain": "A",
"e3_chain_label": "B,C,D",
"preset": "medium",
},
)

Pass the inputs as a JSON string. File paths are uploaded automatically.

Terminal window
azulene jobs submit --job-type protac_pose_prediction \
--input-data '{"target_pdb_content": "target_with_warhead.pdb", "e3_pdb_content": "vhl_with_recruiter.pdb", "linker_smiles": "[*:1]CCOCCOCC[*:2]", "warhead_attach_atom_idx": 12, "recruiter_attach_atom_idx": 7, "target_chain": "A", "e3_chain_label": "B,C,D", "preset": "medium"}'

The result contains the ten highest-ranked poses in three fields: top1, top3, and top10. These are nested subsets of one ranked list, not separate calculations. top1 contains the best pose. top3 contains the best three. top10 contains all ten.

Each entry includes a pose rank and a predicted_RMSD. This value estimates how far the pose is from the true ternary geometry, in Angstroms. Lower is better, and the list is sorted by this value. Use it to compare poses within the run. It is not a calibrated error bar for an individual pose.

The ten poses often span several Angstroms. Use them as a shortlist. Check whether several poses agree on the protein-protein interface. Agreement across several poses is more informative than one score.

The downloadable ZIP contains the minimized ternary structures, the per-pose minimization JSON files for each of the four solvent and protonation conditions, and the run summary. Retrieve it with azulene jobs download.

When ground truth is supplied, each pose also reports protac_rmsd_vs_gt_A. This is the heavy-atom RMSD of the PROTAC after alignment on the target C-alpha atoms. The pocket_rmsd_vs_gt_A field applies the same measure to residues within 8 Angstroms of the reference PROTAC. Each pose also has two multi-model overlay PDBs. One contains the predicted and experimental PROTAC. The other also contains the pocket residues. Both can be opened in PyMOL or ChimeraX.

Chain IDs are the most common cause of failed runs. The target, E3, warhead, and recruiter chain IDs must match the PDB records. Supply ligand SMILES templates for unusual ligands. Bond orders cannot always be perceived reliably from coordinates alone.

Start with medium. This is the reference setting and the setting matched to the ranker. Use long when no medium pose has a plausible interface. Cost rises steeply between presets, so reserve long for a specific system rather than using it by default.