PROTAC Ternary Complex Pose Prediction
This tool predicts ternary complex poses for a PROTAC, its target protein, and an E3 ligase. Supply two binary co-complexes: the target with its warhead bound, and the E3 ligase with its recruiter bound. Also supply the linker as a SMILES fragment and identify the two attachment atoms. The job returns a ranked set of full ternary poses with the PROTAC built and relaxed in place.
The two binary interfaces remain fixed. The output shows where the E3 ligase can sit relative to the target within the reach of the linker. Binary docking does not provide this geometry.
When to use it
Section titled “When to use it”Use this tool when both binary binding modes are known or modelled, but the ternary geometry is not. The predicted poses show whether a linker permits a compatible protein-protein arrangement and what that arrangement looks like. Comparing poses across a linker series shows how linker length and composition affect the predicted geometry. These results can help prioritize designs before synthesis.
Supply a known ternary structure as ground truth to benchmark the predictions. The job reports RMSDs and writes overlay structures.
Scope:
- Validation is concentrated on VHL systems, because that is where most published ternary structures are. Other E3 ligases, CRBN in particular, have less supporting data behind them.
- Both proteins are held rigid. Induced fit at the ternary interface is not modelled.
Inputs
Section titled “Inputs”| Input | Required | What it is |
|---|---|---|
target_pdb_content | yes | PDB file of the target protein with its warhead bound. |
e3_pdb_content | yes | PDB file of the E3 ligase with its recruiter bound. |
linker_smiles | yes | Linker SMILES fragment with two attachment points, [*:1] and [*:2], marking the bonds to the warhead and the recruiter. |
warhead_attach_atom_idx | yes | 0-based atom index in the warhead that bonds to the linker. |
recruiter_attach_atom_idx | yes | 0-based atom index in the E3 recruiter that bonds to the linker. |
target_chain | no, default A | Chain ID or comma-separated chain IDs of the target protein. Use several for a multimeric target. |
e3_chain_label | no, default B | Chain ID or comma-separated chain IDs of the E3 complex. VHL with Elongin C and Elongin B is typically B,C,D. |
warhead_chain | no, default X | Chain ID of the warhead ligand in the target PDB. |
recruiter_chain | no, default Y | Chain ID of the recruiter ligand in the E3 PDB. |
preset | no, default medium | Angular resolution of the rigid-body scan. See the preset table below. |
n_prescan_chunks | no, default 8 | Number of distance bins scanned in parallel. Range 1 to 16. Increasing it shortens the wall time of the scan stage. Each chunk adds about 30 seconds of setup, so do not exceed the number of distance bins. |
warhead_smiles | no | SMILES template for the warhead with correct bond orders. It may also include an attachment point. Supply it when bond orders cannot be perceived reliably from the PDB coordinates alone. |
e3_anchor_smiles | no | The same for the E3 recruiter. |
gt_complex_pdb_content | no | Experimental ternary complex PDB for RMSD scoring and overlays. |
gt_ligand_resname | no | Residue name of the PROTAC in the ground-truth file. Required when ground truth is supplied. |
gt_target_chain | no, default A | Target chain ID in the ground-truth file. |
gt_e3_chain | no, default B | E3 chain ID or IDs in the ground-truth file. |
keep_dirs | no, default true | Persist poses, overlays, and the summary as a downloadable ZIP. Set false to return only the scalar JSON summary. |
Presets
Section titled “Presets”| Preset | Relative cost | Use it for |
|---|---|---|
quick | About a third of medium | A fast, coarse pass. It gives only a rough view of the pose landscape. Use it as a demonstration rather than a prediction. |
medium | Reference | The default setting and the setting matched to the ranker. Start here. |
long | Several times medium | A finer, wider search. Use it when medium returns no pose with a plausible interface. This can happen when the E3 sits far around the target. |
How to run it
Section titled “How to run it”Submit the job from Azulene Studio, the Python SDK, or the CLI. Local file paths are uploaded automatically. The Get started page covers installation, login, and a worked example.
In Azulene Studio
Section titled “In Azulene Studio”Open PROTAC Ternary Complex Pose Prediction from the tools list. On the Inputs and Parameters step, upload the target and E3 PDB files. Enter the linker SMILES with its two attachment points. Set the two attachment atom indices. Check that the protein and ligand chain IDs match your files. Leave the preset at medium, then select Review and Submit.
From the Python SDK
Section titled “From the Python SDK”from azulene import jobs
result = jobs.submit( job_type="protac_pose_prediction", input_data={ "target_pdb_content": "target_with_warhead.pdb", "e3_pdb_content": "vhl_with_recruiter.pdb", "linker_smiles": "[*:1]CCOCCOCC[*:2]", "warhead_attach_atom_idx": 12, "recruiter_attach_atom_idx": 7, "target_chain": "A", "e3_chain_label": "B,C,D", "preset": "medium", },)From the CLI
Section titled “From the CLI”Pass the inputs as a JSON string. File paths are uploaded automatically.
azulene jobs submit --job-type protac_pose_prediction \ --input-data '{"target_pdb_content": "target_with_warhead.pdb", "e3_pdb_content": "vhl_with_recruiter.pdb", "linker_smiles": "[*:1]CCOCCOCC[*:2]", "warhead_attach_atom_idx": 12, "recruiter_attach_atom_idx": 7, "target_chain": "A", "e3_chain_label": "B,C,D", "preset": "medium"}'Reading the result
Section titled “Reading the result”The result contains the ten highest-ranked poses in three fields: top1, top3, and top10. These are nested subsets of one ranked list, not separate calculations. top1 contains the best pose. top3 contains the best three. top10 contains all ten.
Each entry includes a pose rank and a predicted_RMSD. This value estimates how far the pose is from the true ternary geometry, in Angstroms. Lower is better, and the list is sorted by this value. Use it to compare poses within the run. It is not a calibrated error bar for an individual pose.
The ten poses often span several Angstroms. Use them as a shortlist. Check whether several poses agree on the protein-protein interface. Agreement across several poses is more informative than one score.
The downloadable ZIP contains the minimized ternary structures, the per-pose minimization JSON files for each of the four solvent and protonation conditions, and the run summary. Retrieve it with azulene jobs download.
When ground truth is supplied, each pose also reports protac_rmsd_vs_gt_A. This is the heavy-atom RMSD of the PROTAC after alignment on the target C-alpha atoms. The pocket_rmsd_vs_gt_A field applies the same measure to residues within 8 Angstroms of the reference PROTAC. Each pose also has two multi-model overlay PDBs. One contains the predicted and experimental PROTAC. The other also contains the pocket residues. Both can be opened in PyMOL or ChimeraX.
Chain IDs are the most common cause of failed runs. The target, E3, warhead, and recruiter chain IDs must match the PDB records. Supply ligand SMILES templates for unusual ligands. Bond orders cannot always be perceived reliably from coordinates alone.
Start with medium. This is the reference setting and the setting matched to the ranker. Use long when no medium pose has a plausible interface. Cost rises steeply between presets, so reserve long for a specific system rather than using it by default.