Skip to content

Aqueous Solubility (logS) Prediction

Aqueous Solubility (logS) Prediction

Predict how well a molecule dissolves in water, from its SMILES string.

Aqueous Solubility (logS) Prediction estimates logS, defined as the base ten logarithm of intrinsic aqueous solubility in mol/L, from a SMILES string. Each prediction includes a solubility class derived from the predicted value.

The model was trained using AqSolDB, a public compilation of measured solubility values. Predictions for chemical structures outside the domain represented in this training set warrant reduced confidence.

The tool supports candidate triage, screening-library filtration and structure assessment before more resource-intensive studies. A SMILES string is the only required input, permitting predictions before structural or assay data are available.

InputRequiredWhat it is
smilesyesSMILES string of the molecule.

Submit a single SMILES string or a batch through Azulene Studio, the Python SDK or the CLI. The Get started page describes installation, authentication and an example submission.

Open Aqueous Solubility (logS) Prediction from the tools list. In the Inputs and Parameters step, enter a single SMILES string, paste a list, or upload a CSV or SDF file for batch submission. Select Review and Submit.

from azulene import jobs
result = jobs.submit(
job_type="predict_solubility",
input_data={
"smiles": "CCO",
},
)

For multiple molecules, submit a batch of SMILES strings in one job.

Pass the inputs as a JSON string.

Terminal window
azulene jobs submit --job-type predict_solubility \
--input-data '{"smiles": "CCO"}'

The CLI also accepts a batch of SMILES strings in a single job.

The result contains the following fields for each molecule:

  • predicted_logS, the predicted logS in mol/L. Higher values indicate greater solubility.

  • solubility_class, a classification derived from predicted_logS:

    ClasslogS
    highly soluble0 and above
    soluble-2 to 0
    moderate-4 to -2
    poorly soluble-6 to -4
    insolublebelow -6
  • smiles, the submitted molecule.

  • descriptors, the molecular descriptors calculated from the structure for the prediction.

To rank molecules by predicted solubility, sort predicted_logS in descending order. If a SMILES string cannot be parsed, the corresponding result contains an error message in place of a prediction.

The predicted value is a machine-learning estimate and does not constitute an experimental measurement. It is appropriate for screening and prioritization. Predictive accuracy may decrease for chemical structures outside the domain represented in the AqSolDB training set. The output does not include a confidence estimate, so predictions for unusual scaffolds require additional scrutiny. A single batch job can process a large molecular library.