Tasks, Trials and Jobs
A Runner defines the execution entrypoint, a Task is an evaluation problem in a Dataset, a Trial records one run, and a Job organizes multiple task runs from a fixed Dataset. The CLI / SDK are clients; the platform validates and freezes inputs, executes tasks, scores results, and manages artifacts.
| Concept | Meaning | Usage |
|---|---|---|
| Runner | A reusable execution package containing runner.toml and main.py / main.sh | Two-file directory or runner="envhub/harbor" |
| Task | An evaluation problem in a Dataset, including instructions, environment, reference solution, and verifier | dataset={"name": "org/bench", "task_id": "org/task"} |
| Trial | One run of a fixed Task, identified by trial_id | el task status TRIAL_ID |
| Dataset | A versioned task collection; members are frozen after selecting config / split | el job submit --dataset NAME |
| Job | One collection run, identified by job_id, organizing member Trials and attempts | client.jobs.submit(...) |
| Profile | An execution parameter template, snapshotted on submission; not a running resource | profile={...} or a named Profile |
| Artifact | Output files and hashes beyond stdout / stderr | client.tasks.result(...) |
State and score
Trial wait terminal states are succeeded, failed, and cancelled. Job wait terminal states are completed, failed, and cancelled; a Job's completed status means collection execution has ended, so still inspect each Trial and its scores.
Execution success does not mean evaluation success: benchmark verifier rewards are recorded separately. Missing scores cannot be inferred as 0 or 1, and scores should not be guessed from stdout. Read Job aggregate results with jobs.result().
Frozen inputs and retries
The platform freezes task contents, Dataset selection, and parameters when it first accepts a submission. Reusing request_id retries the original request and returns the first execution without updating its inputs. Use a new request_id when changing parameters or intentionally rerunning; the SDK generates one if omitted.
When Dataset version is omitted, the platform resolves it using its fixed source and default rules; this does not mean the client automatically selects latest. For reproducible runs, record the response's version, source snapshot, Profile, dependency environment, and artifact hashes.
Local and remote
Remote mode connects to an existing platform; closing the client does not stop the server or submitted work. Local mode creates an in-process service and waits for accepted work to finish on close; local Tasks execute trusted scripts directly and are not a secure isolated sandbox.
Without a repository binding, local state exists only in the current service instance. With a repository binding, completed receipts can be read by a new client; running state does not recover across processes.
Training boundary
A Train Job is a separate training resource with train_job_id and is not an Agent Job / Task / Trial. The source CLI supports independent Train Job submission and status queries over a remote API through el train submit/status; Client still has no client.train entrypoint or direct Tinker SDK integration. Real GPU training has not been verified. See Resources for support status.