deeporigin.drug_discovery.metabolism¶
Metabolism drives platform tool deeporigin.metabolism. It scores ligands
for sites of metabolism on cytochrome P450 (CYP) isoforms.
Use run() for small batches (fewer than 30 ligands); it blocks and returns
a table of sites (every enzyme the tool scored). For 30 or more ligands, call
start(), then wait() or watch(), then get_results() /
get_molecules(). There is no cost quote and no client-side ligand cap.
Metabolism -- predict sites of metabolism for ligands.
Backed by the platform tool deeporigin.metabolism. One :class:Metabolism
instance is configured with ligands, then executed with a blocking
:meth:run (small batches) or asynchronous :meth:start (larger batches).
:meth:run returns a :class:pandas.DataFrame of Metabolism site rows
(atom, enzyme, site probability). :meth:get_molecules returns
molecule-level confidence_tier rows for this execution.
Class-level :meth:fetch_results / :meth:fetch_molecules load indexed
rows for a ligand set from the data platform (any past jobs) by
ligand_id. Before run / start, if every ligand has a platform id
and every id already has a Metabolism molecule, the client refuses (no
execution). If the job still proceeds and any id is already indexed, it
warns. Pass force=True to :meth:run or :meth:start to recompute ligands
that already have indexed MetabolismMolecule rows (workflow skip filter).
The tool scores every cytochrome P450 isoform it supports; the client does
not select or filter enzymes. tool_version stays "latest". Ligands
are not mutated. Payload id is sent only when
:attr:~deeporigin.drug_discovery.structures.ligand.Ligand.id is already set.
Batches larger than the platform Inline ligand cap
(:data:~deeporigin.utils.constants.INLINE_LIGAND_CAP) dump a
Ligand list file to UFA and submit inputs.ligands_file on
:meth:start (transparent to the caller).
Sync usage (blocking; fewer than 30 ligands)::
from deeporigin.drug_discovery import Metabolism, Ligand
job = Metabolism(ligands=Ligand.from_smiles("CCO"))
sites = job.run()
mols = job.get_molecules()
Async usage (30 or more ligands, or any size)::
job = Metabolism(ligands=ligands)
job.start()
await job.watch() # or job.wait()
sites = job.get_results()
mols = job.get_molecules()
Fetch indexed rows without starting a job::
sites = Metabolism.fetch_results(ligands=ligands)
mols = Metabolism.fetch_molecules(ligands=ligands)
Attributes¶
Classes¶
Metabolism
¶
Bases: Execution, SyncExecutableMixin, AsyncExecutableMixin, NotebookWatchMixin
Predict sites of metabolism for ligands via deeporigin.metabolism.
The tool scores every cytochrome P450 isoform it supports. Ligands are
not mutated (contrast with
:class:~deeporigin.drug_discovery.molprops.Molprops).
Use :meth:run for fewer than
:data:~deeporigin.utils.constants.METABOLISM_WORKFLOW_LIGAND_THRESHOLD
ligands (blocking). For larger batches, call :meth:start, then
:meth:wait or :meth:watch, then :meth:get_results /
:meth:get_molecules (data platform first, jobOutputs fallback).
Batches above
:data:~deeporigin.utils.constants.INLINE_LIGAND_CAP upload a
Ligand list file and pass ligands_file instead of inline ligands.
Use :meth:fetch_results / :meth:fetch_molecules to read indexed rows
for a ligand set without starting a job.
Attributes:
| Name | Type | Description |
|---|---|---|
ligands |
list[Ligand]
|
Ligands whose SMILES are sent to the tool. |
name |
Execution label, set from the ligand count unless overridden. |
Attributes¶
USER_LOG_COLUMNS
class-attribute
¶
USER_LOG_COLUMNS: list[str] = [
"log_level",
"tool_key",
"timestamp",
"message",
]
cost
property
¶
cost: float | None
Actual cost in dollars, set after execution completes.
This property cannot be set manually.
estimate
property
¶
estimate: float | None
Cost estimate in dollars, populated when the platform returns a quotation.
Set after run(quote=True), start(quote=True), or any call with
approve_amount=-1. None until a quotation result is received.
This property cannot be set manually.
name
instance-attribute
¶
name = (
name
if name is not None
else _metabolism_default_name(len(self._ligands))
)
runtime
property
¶
runtime: float | None
Seconds from DTO startedAt to completedAt or current UTC time.
Uses :attr:dto (same shape as client.executions.get). When
completedAt is present, it is the end time; otherwise the end time is
datetime.now(timezone.utc). Returns None if there is no DTO or
startedAt is missing or empty.
tool_key
class-attribute
instance-attribute
¶
tool_key: str = TOOL_KEYS_AND_VERSIONS["metabolism"][
"tool_key"
]
tool_version
class-attribute
instance-attribute
¶
tool_version: str = TOOL_KEYS_AND_VERSIONS["metabolism"][
"tool_version"
]
Methods:¶
cancel
¶
cancel() -> None
Cancel a running or queued execution.
Raises:
| Type | Description |
|---|---|
ValueError
|
If the job has no execution ID. |
ValueError
|
If the job is not in a cancellable state. |
confirm
¶
confirm() -> None
Confirm a quoted tools execution on the platform.
Requires :attr:id and status equal to "Quoted". Uses
:meth:~deeporigin.platform.executions.Executions.confirm with
:data:~deeporigin.utils.constants.TOOL_EXECUTION_POST_TIMEOUT_SECONDS
(10 minutes) and retry=False, then applies the returned DTO via
:meth:update_from_dto so status, :attr:cost, and :attr:dto
reflect the platform response. For direct (blocking) tools the response
is typically terminal (Completed); for async tools it may be
Created or Running — call :meth:sync until the job finishes.
Raises:
| Type | Description |
|---|---|
ValueError
|
If there is no platform execution id or status is not
|
duplicate
¶
duplicate(
*, client: DeepOriginClient | None = None
) -> Self
Create a fresh copy with the same configuration but no execution state.
Useful after from_id() to re-run the same calculation. The
returned instance has no id, status, estimate, or
cost — it is ready for run() / start().
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
client
|
DeepOriginClient | None
|
Optional API client for the new instance. Falls back to the current instance's client. |
None
|
Returns:
| Type | Description |
|---|---|
Self
|
A new instance sharing the same domain-specific configuration. |
fetch_molecules
classmethod
¶
fetch_molecules(
ligands: Ligand | list[Ligand] | LigandSet,
*,
client: DeepOriginClient | None = None
) -> DataFrame
Load indexed Metabolism molecule rows for ligands (any past jobs).
Queries the data platform by platform ligand_id. Does not start an
execution. Partial or empty tables are OK when some ligands lack an id
or have no indexed MetabolismMolecule. Missing smiles on
indexed rows is filled from ligands by ligand_id.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
ligands
|
Ligand | list[Ligand] | LigandSet
|
A ligand, list, or :class: |
required |
client
|
DeepOriginClient | None
|
Optional API client. Uses the default if not provided. |
None
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
DataFrame with preferred columns |
DataFrame
|
|
fetch_results
classmethod
¶
fetch_results(
ligands: Ligand | list[Ligand] | LigandSet,
*,
client: DeepOriginClient | None = None,
top_k: int | None = None,
min_prob: float | None = None
) -> DataFrame
Load indexed Metabolism site rows for ligands (any past jobs).
Queries the data platform by platform ligand_id. Does not start an
execution. Ligands without an id contribute no filter keys; missing
indexed rows are omitted (partial or empty tables are OK). Missing
smiles on indexed rows is filled from ligands by ligand_id.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
ligands
|
Ligand | list[Ligand] | LigandSet
|
A ligand, list, or :class: |
required |
client
|
DeepOriginClient | None
|
Optional API client. Uses the default if not provided. |
None
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
DataFrame with preferred columns |
DataFrame
|
|
DataFrame
|
|
DataFrame
|
Returns rows from all indexed jobs (history-preserving). |
from_dto
classmethod
¶
from_dto(
dto: dict[str, Any],
*,
client: DeepOriginClient | None = None
) -> Self
Construct a Metabolism from a tools execution DTO.
Restores ligands from userInputs (falling back to inputs).
When only ligands_file is present, downloads and parses that UFA
Ligand list file.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dto
|
dict[str, Any]
|
Execution payload (same shape as |
required |
client
|
DeepOriginClient | None
|
Optional API client. Uses the default if not provided. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
A |
Self
|
class: |
Raises:
| Type | Description |
|---|---|
ValueError
|
If stored inputs have no ligands, or |
from_id
classmethod
¶
from_id(
id: str, *, client: DeepOriginClient | None = None
) -> Self
Construct an instance from an existing platform execution ID.
Fetches the execution DTO via client.executions.get and delegates to
:meth:from_dto. Concrete subclasses override :meth:from_dto to attach
domain state from userInputs.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
id
|
str
|
Platform execution ID. |
required |
client
|
DeepOriginClient | None
|
Optional API client. Uses the default if not provided. |
None
|
Returns:
| Type | Description |
|---|---|
Self
|
A partially-hydrated instance with common fields populated. |
Raises:
| Type | Description |
|---|---|
NotImplementedError
|
If |
from_last_run
classmethod
¶
from_last_run(
*, client: DeepOriginClient | None = None
) -> Self
Construct an instance from the most recently created execution of this tool.
Scoped to the client's project, like :meth:list. Calls
client.executions.list with tool_key, order set to
:data:~deeporigin.utils.constants.EXECUTION_LIST_ORDER_CREATED_DESC,
the client's project_id, and page_size=1, then delegates to
:meth:from_dto. Concrete subclasses inherit this method; domain
state is restored via their from_dto overrides.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
client
|
DeepOriginClient | None
|
Optional API client. Uses the default if not provided. |
None
|
Returns:
| Type | Description |
|---|---|
Self
|
A partially-hydrated instance for the newest execution by |
Self
|
|
Raises:
| Type | Description |
|---|---|
NotImplementedError
|
If |
ValueError
|
If no executions exist for this tool type. |
get_molecules
¶
get_molecules(
dto: dict[str, Any] | None = None,
) -> DataFrame
Return molecule-level confidence_tier rows as a DataFrame.
Prefers data-platform result-explorer rows for this execution
(result_type=metabolismmolecule), then falls back to
jobOutputs.molecules. One row per scored SMILES.
For indexed molecules across any past jobs (no execution required),
use :meth:fetch_molecules instead of calling this on the class
with ligands.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dto
|
dict[str, Any] | None
|
Optional execution payload from |
None
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
DataFrame with |
Raises:
| Type | Description |
|---|---|
TypeError
|
If called as |
ValueError
|
If :attr: |
DeepOriginException
|
If no molecule rows could be parsed. |
get_results
¶
get_results(
dto: dict[str, Any] | None = None,
*,
top_k: int | None = None,
min_prob: float | None = None
) -> DataFrame
Return this execution's Metabolism site rows as a DataFrame.
Prefers data-platform result-explorer rows for this execution
(result_type=metabolismsite), then falls back to
jobOutputs.sites. Includes every enzyme the tool scored.
For indexed sites across any past jobs (no execution required), use
:meth:fetch_results instead of calling this on the class with
ligands.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dto
|
dict[str, Any] | None
|
Optional execution payload from |
None
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
DataFrame with |
DataFrame
|
|
DataFrame
|
filter rows client-side. This execution's rows are authoritative |
DataFrame
|
after a forced rerun. |
Raises:
| Type | Description |
|---|---|
TypeError
|
If called as |
ValueError
|
If :attr: |
DeepOriginException
|
If no site rows could be parsed. |
get_user_logs
¶
get_user_logs(
*,
limit: int | None = None,
offset: int | None = None,
select: list[str] | None = None,
with_total_count: bool = False
) -> DataFrame | None
Search data-platform user_logs rows for this execution.
Uses :meth:deeporigin.platform.user_logs.UserLogs.search with this
execution's id (tools executionId), stored as execution_id on
user_logs rows — the same string passed to :meth:get_results as
compute_job_id.
When no execution id is assigned yet, returns None without calling
the API.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
limit
|
int | None
|
Max rows to return (forwarded to |
None
|
offset
|
int | None
|
Skip offset (forwarded). |
None
|
select
|
list[str] | None
|
Columns to select (forwarded). |
None
|
with_total_count
|
bool
|
Request total count from the server (forwarded). |
False
|
Returns:
| Type | Description |
|---|---|
DataFrame | None
|
A DataFrame with columns |
DataFrame | None
|
and |
DataFrame | None
|
|
DataFrame | None
|
if this instance has no execution id yet or the client has no |
DataFrame | None
|
|
list
classmethod
¶
list(
*,
client: DeepOriginClient | None = None,
status: list[str] | None = None
) -> list[Self]
List executions of this tool, newest first, scoped to the client's project.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
client
|
DeepOriginClient | None
|
Optional API client. Uses the default if not provided. |
None
|
status
|
list[str] | None
|
Optional list of statuses to keep. |
None
|
Returns:
| Type | Description |
|---|---|
list[Self]
|
Instances of this class, newest first. |
Raises:
| Type | Description |
|---|---|
NotImplementedError
|
If |
run
¶
run(
*,
force: bool = False,
top_k: int | None = None,
min_prob: float | None = None
) -> DataFrame
Execute metabolism synchronously and return site rows.
Blocks until the job finishes. Requires fewer than
:data:~deeporigin.utils.constants.METABOLISM_WORKFLOW_LIGAND_THRESHOLD
ligands; use :meth:start for larger batches. There is no
quote=True path. The sites table includes every enzyme the tool
scored.
Refuses before create when every ligand has a platform id and every
id already has a Metabolism molecule (use :meth:fetch_results
instead).
Returns:
| Name | Type | Description |
|---|---|---|
A |
DataFrame
|
class: |
Raises:
| Type | Description |
|---|---|
DeepOriginException
|
If all ligands are already scored, the execution did not complete successfully, or no site rows could be parsed. |
ValueError
|
If there are 30 or more ligands. |
show
¶
show() -> None
Display the current execution in Jupyter using the execution card HTML view.
If no platform execution ID exists yet, shows the same card with a short notice
instead of raising (see :meth:~deeporigin.platform.execution_display.ExecutionDisplay.from_pending).
start
¶
start(
*,
quote: bool = False,
approve_amount: int | None = None,
**kwargs
) -> None
Submit a persisted async execution to the platform.
Only valid when status is None (no execution exists yet).
All other statuses raise immediately to prevent re-submission.
Pass quote=True or approve_amount=-1 to request a cost estimate
without running. If the platform returns a Quoted DTO the instance
is left in that state — call :meth:~deeporigin.drug_discovery.execution.Execution.confirm
explicitly to proceed.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
quote
|
bool
|
Shorthand for |
False
|
approve_amount
|
int | None
|
Spend cap passed to the platform as |
None
|
**kwargs
|
Forwarded verbatim to |
{}
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If the current status is not |
stop_watching
¶
stop_watching() -> None
Cancel an in-flight watch loop if one is running.
Safe to call when no watch is active. Does not cancel the caller's current task when invoked from inside the watch loop.
sync
¶
sync() -> None
Fetch the latest tools execution from the platform and refresh fields.
Calls client.executions.get for :attr:id and applies the response
with :meth:update_from_dto. Use when the job may have changed outside
this process (for example after submission from the web UI), to poll
lifecycle state, or to refresh an instance built from an older DTO.
Available on sync-only and async execution types alike.
Rejected executions are refreshed without raising for their status, so history inspection and notebook displays remain available. HTTP failures while fetching the execution still raise.
If executions.get returns a falsy value, this instance is left
unchanged.
Raises:
| Type | Description |
|---|---|
ValueError
|
If this instance has no execution id yet. |
NotImplementedError
|
If |
ValueError
|
If the returned DTO |
update_from_dto
¶
update_from_dto(dto: dict[str, Any]) -> None
Apply tools execution fields from dto onto this instance.
Updates id, pricing, lifecycle fields, and _dto the same
way as :meth:from_dto for a newly created instance. Use after a live
executions.create / sync() response to refresh state without
constructing a new object (domain inputs on self are unchanged).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dto
|
dict[str, Any]
|
Execution payload (same shape as |
required |
Raises:
| Type | Description |
|---|---|
NotImplementedError
|
If |
ValueError
|
If the DTO |
wait
¶
wait(
*,
poll_interval: float = 5.0,
timeout: float | None = None
) -> dict[str, Any] | None
Block until this execution reaches a terminal state.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
poll_interval
|
float
|
Seconds to sleep between polling cycles. |
5.0
|
timeout
|
float | None
|
Maximum total seconds to wait. If |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, Any] | None
|
The latest execution DTO, or |
Raises:
| Type | Description |
|---|---|
ValueError
|
If this instance has no execution id yet. |
TimeoutError
|
If |
watch
async
¶
watch(
*, interval: float = 5.0, blocking: bool = False
) -> Task | None
Start live notebook updates; optionally block until the job finishes.
By default, awaiting this coroutine finishes in one event-loop turn, so the cell returns while the display keeps updating in the background. Use this when you need to run other cells while the job runs::
task = await abfe.watch()
Set blocking=True or export JOB_WATCH_BLOCK=1 (truthy values:
1, true, yes, on) to run the poll loop inline so the cell
does not return until a terminal state — useful for nbconvert --execute
and doc CI (see :data:~deeporigin.utils.constants.JOB_WATCH_BLOCK_ENV)::
await abfe.watch(blocking=True)
# or: export JOB_WATCH_BLOCK=1
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
interval
|
float
|
Seconds between polls. |
5.0
|
blocking
|
bool
|
When True, await the watch loop instead of returning a
background task. Also blocks when |
False
|
Returns:
| Type | Description |
|---|---|
Task | None
|
The background task when not blocking; |
Task | None
|
Cancel a background watch with :meth: |