Skip to content

deeporigin.drug_discovery.metabolism

Metabolism drives platform tool deeporigin.metabolism. It scores ligands for sites of metabolism on cytochrome P450 (CYP) isoforms.

Use run() for small batches (fewer than 30 ligands); it blocks and returns a table of sites (every enzyme the tool scored). For 30 or more ligands, call start(), then wait() or watch(), then get_results() / get_molecules(). There is no cost quote and no client-side ligand cap.

Metabolism -- predict sites of metabolism for ligands.

Backed by the platform tool deeporigin.metabolism. One :class:Metabolism instance is configured with ligands, then executed with a blocking :meth:run (small batches) or asynchronous :meth:start (larger batches). :meth:run returns a :class:pandas.DataFrame of Metabolism site rows (atom, enzyme, site probability). :meth:get_molecules returns molecule-level confidence_tier rows for this execution.

Class-level :meth:fetch_results / :meth:fetch_molecules load indexed rows for a ligand set from the data platform (any past jobs) by ligand_id. Before run / start, if every ligand has a platform id and every id already has a Metabolism molecule, the client refuses (no execution). If the job still proceeds and any id is already indexed, it warns. Pass force=True to :meth:run or :meth:start to recompute ligands that already have indexed MetabolismMolecule rows (workflow skip filter).

The tool scores every cytochrome P450 isoform it supports; the client does not select or filter enzymes. tool_version stays "latest". Ligands are not mutated. Payload id is sent only when :attr:~deeporigin.drug_discovery.structures.ligand.Ligand.id is already set.

Batches larger than the platform Inline ligand cap (:data:~deeporigin.utils.constants.INLINE_LIGAND_CAP) dump a Ligand list file to UFA and submit inputs.ligands_file on :meth:start (transparent to the caller).

Sync usage (blocking; fewer than 30 ligands)::

from deeporigin.drug_discovery import Metabolism, Ligand

job = Metabolism(ligands=Ligand.from_smiles("CCO"))
sites = job.run()
mols = job.get_molecules()

Async usage (30 or more ligands, or any size)::

job = Metabolism(ligands=ligands)
job.start()
await job.watch()  # or job.wait()
sites = job.get_results()
mols = job.get_molecules()

Fetch indexed rows without starting a job::

sites = Metabolism.fetch_results(ligands=ligands)
mols = Metabolism.fetch_molecules(ligands=ligands)

Attributes

Classes

Metabolism

Bases: Execution, SyncExecutableMixin, AsyncExecutableMixin, NotebookWatchMixin

Predict sites of metabolism for ligands via deeporigin.metabolism.

The tool scores every cytochrome P450 isoform it supports. Ligands are not mutated (contrast with :class:~deeporigin.drug_discovery.molprops.Molprops).

Use :meth:run for fewer than :data:~deeporigin.utils.constants.METABOLISM_WORKFLOW_LIGAND_THRESHOLD ligands (blocking). For larger batches, call :meth:start, then :meth:wait or :meth:watch, then :meth:get_results / :meth:get_molecules (data platform first, jobOutputs fallback). Batches above :data:~deeporigin.utils.constants.INLINE_LIGAND_CAP upload a Ligand list file and pass ligands_file instead of inline ligands.

Use :meth:fetch_results / :meth:fetch_molecules to read indexed rows for a ligand set without starting a job.

Attributes:

Name Type Description
ligands list[Ligand]

Ligands whose SMILES are sent to the tool.

name

Execution label, set from the ligand count unless overridden.

Attributes

USER_LOG_COLUMNS class-attribute
USER_LOG_COLUMNS: list[str] = [
    "log_level",
    "tool_key",
    "timestamp",
    "message",
]
app instance-attribute
app: str | None = None
approve_amount instance-attribute
approve_amount: int | None = None
client instance-attribute
client: DeepOriginClient = client
completed_at instance-attribute
completed_at: str | None = None
cost property
cost: float | None

Actual cost in dollars, set after execution completes.

This property cannot be set manually.

created_at instance-attribute
created_at: str | None = None
created_by instance-attribute
created_by: str | None = None
dto property
dto: dict[str, Any] | None

Last tools execution DTO from the platform, if any.

estimate property
estimate: float | None

Cost estimate in dollars, populated when the platform returns a quotation.

Set after run(quote=True), start(quote=True), or any call with approve_amount=-1. None until a quotation result is received. This property cannot be set manually.

id property
id: str | None

Platform execution ID when set (read-only).

ligands property
ligands: list[Ligand]

Ligands targeted by this run (read-only).

name instance-attribute
name = (
    name
    if name is not None
    else _metabolism_default_name(len(self._ligands))
)
progress instance-attribute
progress: dict | None = None
runtime property
runtime: float | None

Seconds from DTO startedAt to completedAt or current UTC time.

Uses :attr:dto (same shape as client.executions.get). When completedAt is present, it is the end time; otherwise the end time is datetime.now(timezone.utc). Returns None if there is no DTO or startedAt is missing or empty.

session instance-attribute
session: str | None = None
started_at instance-attribute
started_at: str | None = None
status instance-attribute
status: PlatformStatus | None = None
tool_key class-attribute instance-attribute
tool_key: str = TOOL_KEYS_AND_VERSIONS["metabolism"][
    "tool_key"
]
tool_version class-attribute instance-attribute
tool_version: str = TOOL_KEYS_AND_VERSIONS["metabolism"][
    "tool_version"
]

Methods:

cancel
cancel() -> None

Cancel a running or queued execution.

Raises:

Type Description
ValueError

If the job has no execution ID.

ValueError

If the job is not in a cancellable state.

confirm
confirm() -> None

Confirm a quoted tools execution on the platform.

Requires :attr:id and status equal to "Quoted". Uses :meth:~deeporigin.platform.executions.Executions.confirm with :data:~deeporigin.utils.constants.TOOL_EXECUTION_POST_TIMEOUT_SECONDS (10 minutes) and retry=False, then applies the returned DTO via :meth:update_from_dto so status, :attr:cost, and :attr:dto reflect the platform response. For direct (blocking) tools the response is typically terminal (Completed); for async tools it may be Created or Running — call :meth:sync until the job finishes.

Raises:

Type Description
ValueError

If there is no platform execution id or status is not "Quoted".

duplicate
duplicate(
    *, client: DeepOriginClient | None = None
) -> Self

Create a fresh copy with the same configuration but no execution state.

Useful after from_id() to re-run the same calculation. The returned instance has no id, status, estimate, or cost — it is ready for run() / start().

Parameters:

Name Type Description Default
client DeepOriginClient | None

Optional API client for the new instance. Falls back to the current instance's client.

None

Returns:

Type Description
Self

A new instance sharing the same domain-specific configuration.

fetch_molecules classmethod
fetch_molecules(
    ligands: Ligand | list[Ligand] | LigandSet,
    *,
    client: DeepOriginClient | None = None
) -> DataFrame

Load indexed Metabolism molecule rows for ligands (any past jobs).

Queries the data platform by platform ligand_id. Does not start an execution. Partial or empty tables are OK when some ligands lack an id or have no indexed MetabolismMolecule. Missing smiles on indexed rows is filled from ligands by ligand_id.

Parameters:

Name Type Description Default
ligands Ligand | list[Ligand] | LigandSet

A ligand, list, or :class:LigandSet.

required
client DeepOriginClient | None

Optional API client. Uses the default if not provided.

None

Returns:

Type Description
DataFrame

DataFrame with preferred columns ligand_id, smiles, and

DataFrame

confidence_tier.

fetch_results classmethod
fetch_results(
    ligands: Ligand | list[Ligand] | LigandSet,
    *,
    client: DeepOriginClient | None = None,
    top_k: int | None = None,
    min_prob: float | None = None
) -> DataFrame

Load indexed Metabolism site rows for ligands (any past jobs).

Queries the data platform by platform ligand_id. Does not start an execution. Ligands without an id contribute no filter keys; missing indexed rows are omitted (partial or empty tables are OK). Missing smiles on indexed rows is filled from ligands by ligand_id.

Parameters:

Name Type Description Default
ligands Ligand | list[Ligand] | LigandSet

A ligand, list, or :class:LigandSet.

required
client DeepOriginClient | None

Optional API client. Uses the default if not provided.

None

Returns:

Type Description
DataFrame

DataFrame with preferred columns ligand_id, smiles,

DataFrame

atom_index, enzyme, and probability. Optional top_k or

DataFrame

min_prob filter rows client-side without rerunning the tool.

DataFrame

Returns rows from all indexed jobs (history-preserving).

from_dto classmethod
from_dto(
    dto: dict[str, Any],
    *,
    client: DeepOriginClient | None = None
) -> Self

Construct a Metabolism from a tools execution DTO.

Restores ligands from userInputs (falling back to inputs). When only ligands_file is present, downloads and parses that UFA Ligand list file.

Parameters:

Name Type Description Default
dto dict[str, Any]

Execution payload (same shape as client.executions.get).

required
client DeepOriginClient | None

Optional API client. Uses the default if not provided.

None

Returns:

Name Type Description
A Self

class:Metabolism with id, lifecycle fields, and ligands set.

Raises:

Type Description
ValueError

If stored inputs have no ligands, or ligands_file cannot be downloaded or parsed.

from_id classmethod
from_id(
    id: str, *, client: DeepOriginClient | None = None
) -> Self

Construct an instance from an existing platform execution ID.

Fetches the execution DTO via client.executions.get and delegates to :meth:from_dto. Concrete subclasses override :meth:from_dto to attach domain state from userInputs.

Parameters:

Name Type Description Default
id str

Platform execution ID.

required
client DeepOriginClient | None

Optional API client. Uses the default if not provided.

None

Returns:

Type Description
Self

A partially-hydrated instance with common fields populated.

Raises:

Type Description
NotImplementedError

If cls has no tool_key (bare :class:Execution).

from_last_run classmethod
from_last_run(
    *, client: DeepOriginClient | None = None
) -> Self

Construct an instance from the most recently created execution of this tool.

Scoped to the client's project, like :meth:list. Calls client.executions.list with tool_key, order set to :data:~deeporigin.utils.constants.EXECUTION_LIST_ORDER_CREATED_DESC, the client's project_id, and page_size=1, then delegates to :meth:from_dto. Concrete subclasses inherit this method; domain state is restored via their from_dto overrides.

Parameters:

Name Type Description Default
client DeepOriginClient | None

Optional API client. Uses the default if not provided.

None

Returns:

Type Description
Self

A partially-hydrated instance for the newest execution by

Self

createdAt.

Raises:

Type Description
NotImplementedError

If cls has no tool_key (bare :class:Execution).

ValueError

If no executions exist for this tool type.

get_molecules
get_molecules(
    dto: dict[str, Any] | None = None,
) -> DataFrame

Return molecule-level confidence_tier rows as a DataFrame.

Prefers data-platform result-explorer rows for this execution (result_type=metabolismmolecule), then falls back to jobOutputs.molecules. One row per scored SMILES.

For indexed molecules across any past jobs (no execution required), use :meth:fetch_molecules instead of calling this on the class with ligands.

Parameters:

Name Type Description Default
dto dict[str, Any] | None

Optional execution payload from executions.create / executions.get used only for the jobOutputs fallback.

None

Returns:

Type Description
DataFrame

DataFrame with ligand_id, smiles, and confidence_tier.

Raises:

Type Description
TypeError

If called as Metabolism.get_molecules(ligands).

ValueError

If :attr:id is unset and dto is omitted.

DeepOriginException

If no molecule rows could be parsed.

get_results
get_results(
    dto: dict[str, Any] | None = None,
    *,
    top_k: int | None = None,
    min_prob: float | None = None
) -> DataFrame

Return this execution's Metabolism site rows as a DataFrame.

Prefers data-platform result-explorer rows for this execution (result_type=metabolismsite), then falls back to jobOutputs.sites. Includes every enzyme the tool scored.

For indexed sites across any past jobs (no execution required), use :meth:fetch_results instead of calling this on the class with ligands.

Parameters:

Name Type Description Default
dto dict[str, Any] | None

Optional execution payload from executions.create / executions.get used only for the jobOutputs fallback.

None

Returns:

Type Description
DataFrame

DataFrame with ligand_id, smiles, atom_index,

DataFrame

enzyme, and probability. Optional top_k or min_prob

DataFrame

filter rows client-side. This execution's rows are authoritative

DataFrame

after a forced rerun.

Raises:

Type Description
TypeError

If called as Metabolism.get_results(ligands).

ValueError

If :attr:id is unset and dto is omitted.

DeepOriginException

If no site rows could be parsed.

get_user_logs
get_user_logs(
    *,
    limit: int | None = None,
    offset: int | None = None,
    select: list[str] | None = None,
    with_total_count: bool = False
) -> DataFrame | None

Search data-platform user_logs rows for this execution.

Uses :meth:deeporigin.platform.user_logs.UserLogs.search with this execution's id (tools executionId), stored as execution_id on user_logs rows — the same string passed to :meth:get_results as compute_job_id.

When no execution id is assigned yet, returns None without calling the API.

Parameters:

Name Type Description Default
limit int | None

Max rows to return (forwarded to UserLogs.search).

None
offset int | None

Skip offset (forwarded).

None
select list[str] | None

Columns to select (forwarded).

None
with_total_count bool

Request total count from the server (forwarded).

False

Returns:

Type Description
DataFrame | None

A DataFrame with columns log_level, tool_key, timestamp,

DataFrame | None

and message. tool_key omits the deeporigin. prefix;

DataFrame | None

timestamp is a compact humanized relative time. Returns None

DataFrame | None

if this instance has no execution id yet or the client has no

DataFrame | None

user_logs API.

list classmethod
list(
    *,
    client: DeepOriginClient | None = None,
    status: list[str] | None = None
) -> list[Self]

List executions of this tool, newest first, scoped to the client's project.

Parameters:

Name Type Description Default
client DeepOriginClient | None

Optional API client. Uses the default if not provided.

None
status list[str] | None

Optional list of statuses to keep.

None

Returns:

Type Description
list[Self]

Instances of this class, newest first.

Raises:

Type Description
NotImplementedError

If cls has no tool_key (bare :class:Execution).

run
run(
    *,
    force: bool = False,
    top_k: int | None = None,
    min_prob: float | None = None
) -> DataFrame

Execute metabolism synchronously and return site rows.

Blocks until the job finishes. Requires fewer than :data:~deeporigin.utils.constants.METABOLISM_WORKFLOW_LIGAND_THRESHOLD ligands; use :meth:start for larger batches. There is no quote=True path. The sites table includes every enzyme the tool scored.

Refuses before create when every ligand has a platform id and every id already has a Metabolism molecule (use :meth:fetch_results instead).

Returns:

Name Type Description
A DataFrame

class:pandas.DataFrame of Metabolism site rows.

Raises:

Type Description
DeepOriginException

If all ligands are already scored, the execution did not complete successfully, or no site rows could be parsed.

ValueError

If there are 30 or more ligands.

show
show() -> None

Display the current execution in Jupyter using the execution card HTML view.

If no platform execution ID exists yet, shows the same card with a short notice instead of raising (see :meth:~deeporigin.platform.execution_display.ExecutionDisplay.from_pending).

start
start(
    *,
    quote: bool = False,
    approve_amount: int | None = None,
    **kwargs
) -> None

Submit a persisted async execution to the platform.

Only valid when status is None (no execution exists yet). All other statuses raise immediately to prevent re-submission.

Pass quote=True or approve_amount=-1 to request a cost estimate without running. If the platform returns a Quoted DTO the instance is left in that state — call :meth:~deeporigin.drug_discovery.execution.Execution.confirm explicitly to proceed.

Parameters:

Name Type Description Default
quote bool

Shorthand for approve_amount=-1. Takes precedence when both quote and approve_amount are provided.

False
approve_amount int | None

Spend cap passed to the platform as approveAmount. -1 (or any negative value) forces a quote-only create. None omits the field (platform may auto-confirm).

None
**kwargs

Forwarded verbatim to _start_impl.

{}

Raises:

Type Description
ValueError

If the current status is not None.

stop_watching
stop_watching() -> None

Cancel an in-flight watch loop if one is running.

Safe to call when no watch is active. Does not cancel the caller's current task when invoked from inside the watch loop.

sync
sync() -> None

Fetch the latest tools execution from the platform and refresh fields.

Calls client.executions.get for :attr:id and applies the response with :meth:update_from_dto. Use when the job may have changed outside this process (for example after submission from the web UI), to poll lifecycle state, or to refresh an instance built from an older DTO. Available on sync-only and async execution types alike.

Rejected executions are refreshed without raising for their status, so history inspection and notebook displays remain available. HTTP failures while fetching the execution still raise.

If executions.get returns a falsy value, this instance is left unchanged.

Raises:

Type Description
ValueError

If this instance has no execution id yet.

NotImplementedError

If type(self).tool_key is empty (bare :class:Execution).

ValueError

If the returned DTO tool.key does not match this class (see :meth:update_from_dto).

update_from_dto
update_from_dto(dto: dict[str, Any]) -> None

Apply tools execution fields from dto onto this instance.

Updates id, pricing, lifecycle fields, and _dto the same way as :meth:from_dto for a newly created instance. Use after a live executions.create / sync() response to refresh state without constructing a new object (domain inputs on self are unchanged).

Parameters:

Name Type Description Default
dto dict[str, Any]

Execution payload (same shape as client.executions.get).

required

Raises:

Type Description
NotImplementedError

If type(self) has no tool_key (bare :class:Execution).

ValueError

If the DTO tool.key does not match tool_key.

wait
wait(
    *,
    poll_interval: float = 5.0,
    timeout: float | None = None
) -> dict[str, Any] | None

Block until this execution reaches a terminal state.

Parameters:

Name Type Description Default
poll_interval float

Seconds to sleep between polling cycles.

5.0
timeout float | None

Maximum total seconds to wait. If None, waits indefinitely.

None

Returns:

Type Description
dict[str, Any] | None

The latest execution DTO, or None if the platform returned no DTO.

Raises:

Type Description
ValueError

If this instance has no execution id yet.

TimeoutError

If timeout elapses before the execution terminates.

watch async
watch(
    *, interval: float = 5.0, blocking: bool = False
) -> Task | None

Start live notebook updates; optionally block until the job finishes.

By default, awaiting this coroutine finishes in one event-loop turn, so the cell returns while the display keeps updating in the background. Use this when you need to run other cells while the job runs::

task = await abfe.watch()

Set blocking=True or export JOB_WATCH_BLOCK=1 (truthy values: 1, true, yes, on) to run the poll loop inline so the cell does not return until a terminal state — useful for nbconvert --execute and doc CI (see :data:~deeporigin.utils.constants.JOB_WATCH_BLOCK_ENV)::

await abfe.watch(blocking=True)
# or: export JOB_WATCH_BLOCK=1

Parameters:

Name Type Description Default
interval float

Seconds between polls.

5.0
blocking bool

When True, await the watch loop instead of returning a background task. Also blocks when JOB_WATCH_BLOCK is truthy.

False

Returns:

Type Description
Task | None

The background task when not blocking; None when blocking.

Task | None

Cancel a background watch with :meth:stop_watching or task.cancel().

Functions: