Skip to main content

Tasks, Trials and Jobs

A Runner defines the execution entrypoint, a Task is an evaluation problem in a Dataset, a Trial records one run, and a Job organizes multiple task runs from a fixed Dataset. The CLI / SDK are clients; the platform validates and freezes inputs, executes tasks, scores results, and manages artifacts.

ConceptMeaningUsage
RunnerA reusable execution package containing runner.toml and main.py / main.shTwo-file directory or runner="envhub/harbor"
TaskAn evaluation problem in a Dataset, including instructions, environment, reference solution, and verifierdataset={"name": "org/bench", "task_id": "org/task"}
TrialOne run of a fixed Task, identified by trial_idel task status TRIAL_ID
DatasetA versioned task collection; members are frozen after selecting config / splitel job submit --dataset NAME
JobOne collection run, identified by job_id, organizing member Trials and attemptsclient.jobs.submit(...)
ProfileAn execution parameter template, snapshotted on submission; not a running resourceprofile={...} or a named Profile
ArtifactOutput files and hashes beyond stdout / stderrclient.tasks.result(...)

State and score​

Trial wait terminal states are succeeded, failed, and cancelled. Job wait terminal states are completed, failed, and cancelled; a Job's completed status means collection execution has ended, so still inspect each Trial and its scores.

Execution success does not mean evaluation success: benchmark verifier rewards are recorded separately. Missing scores cannot be inferred as 0 or 1, and scores should not be guessed from stdout. Read Job aggregate results with jobs.result().

Frozen inputs and retries​

The platform freezes task contents, Dataset selection, and parameters when it first accepts a submission. Reusing request_id retries the original request and returns the first execution without updating its inputs. Use a new request_id when changing parameters or intentionally rerunning; the SDK generates one if omitted.

When Dataset version is omitted, the platform resolves it using its fixed source and default rules; this does not mean the client automatically selects latest. For reproducible runs, record the response's version, source snapshot, Profile, dependency environment, and artifact hashes.

Local and remote​

Remote mode connects to an existing platform; closing the client does not stop the server or submitted work. Local mode creates an in-process service and waits for accepted work to finish on close; local Tasks execute trusted scripts directly and are not a secure isolated sandbox.

Without a repository binding, local state exists only in the current service instance. With a repository binding, completed receipts can be read by a new client; running state does not recover across processes.

Training boundary​

A Train Job is a separate training resource with train_job_id and is not an Agent Job / Task / Trial. The source CLI supports independent Train Job submission and status queries over a remote API through el train submit/status; Client still has no client.train entrypoint or direct Tinker SDK integration. Real GPU training has not been verified. See Resources for support status.