Run dataset jobs
A Job runs member Tasks from the fixed Dataset revision resolved by the platform, with configurable attempts per member. Aggregate results belong to the Job; do not pass a Job ID to task commands.
CLI workflow
el job submit --dataset envloop/terminal-bench-2-1-lite --attempts 1
el job list --limit 25
el job status JOB_ID
el job result JOB_ID
el job stop JOB_ID
An APIServer with catalog support lets you specify the runner and queue parameters explicitly:
el job submit --runner envhub/harbor@0.23.0 --dataset envloop/terminal-bench-2-1-lite \
--queue-id default --priority 5 --tag nightly
A runner can include @version; optional --profile-id references a platform Profile and --job-id specifies this Job's identifier. Higher priority dispatches waiting tasks sooner without preempting running tasks; defaults are queue default and priority 0. A bare JobMaster / LocalServer without a catalog rejects named runner requests.
Replace JOB_ID with the response value. The dataset must be deployed on the target platform; you can substitute your own available dataset. Use a returned cursor with el job list --cursor CURSOR to read the next page.
Dataset selection
| Option | Behavior |
|---|---|
name@version | Select the exact version declared by the author |
--namespace | Complete the namespace for a short name |
--config | Select a dataset subset; uses the declared default if omitted |
--split | Select one complete group; uses all groups in the config if omitted |
--revision | Source Git branch / tag / commit; must match the platform's bound snapshot |
--attempts | Attempts per member, default 1 |
Multiple configurations without a default or ambiguous version matches produce explicit errors rather than guessing latest. Split slicing expressions are unsupported. Remote users usually do not need repository-root / repository-commit; source management belongs to the platform.
Python workflow
from envloop import Client
with Client.remote() as client:
job = client.jobs.submit(
"envloop/terminal-bench-2-1-lite",
attempts=1,
request_id="lite-evaluation-001",
)
job_id = job["job_id"]
final = client.jobs.wait(job_id, timeout=600)
print(final["status"]) # Job terminal states: completed / failed / cancelled
print(client.jobs.result(job_id))
jobs.wait() returns status details, not aggregate scores. Call jobs.result() separately to read scores. Even when a Job is completed, check for member failures and missing scores.
The SDK also accepts a fixed dataset_ref mapping supplied by the platform; fixed references cannot be combined with namespace / config / split / revision selectors. Retrying a request_id preserves the first frozen selection; use a new ID after changing the dataset or parameters.
Local Dataset Jobs require an initialized, bound repository containing Dataset definitions; an empty Client.local() does not automatically provide datasets. See CLI reference and SDK reference for full parameters.