labtasker
Use Labtasker v2 to queue and run independent ML inference, evaluation, or experiment Tasks; migrate pipelines; design routes and Workers; inspect Task demand and Worker activity; and recover Tasks. Do not use it as a GPU allocator, cluster scheduler, workflow DAG, or artifact st
Install
npx skills add https://github.com/luocfprime/labtasker/tree/main/skills/labtasker
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install luocfprime-labtasker@llmmart
git clone https://github.com/luocfprime/labtasker.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole luocfprime/labtasker collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Labtasker
Use Labtasker when many independent ML jobs should be distributed across processes the user already controls, with progress, retries, and small structured results kept in one place. Keep GPU allocation, process launching, cluster management, dependent workflows, and artifact storage outside Labtasker.
Prefer the documented public path even when a custom workaround is technically possible. Do not add a Server, shell wrapper, Queue, or compatibility mechanism unless the workload needs it.
Documentation map: https://raw.githubusercontent.com/luocfprime/labtasker/refs/heads/main/docs/llms.txt
Read the relevant reference
- Read deployment-and-capabilities.md for installation, local versus shared operation, HTTP authentication, Windows, Unix-socket requests, package selection, version warnings, or “does it support this?” questions.
- Read workers-and-workloads.md when submitting Tasks, converting an experiment, choosing routes or Queues, binding Task args, wrapping a command, reusing a loaded model, reporting progress, or using a distributed launcher.
- Read operations-and-recovery.md for idempotent submission, priority, filtering, fuzzy name search, pagination, updates, cancellation, external early-stop decisions, retries, interruption, and rerunning work.
- Read observations-and-counts.md for online Workers, route presence, busy/idle activity, Worker metadata/latest telemetry, grouped counts, and paginated monitoring queries.
Read every reference relevant to the request before proposing commands. Before
the first Labtasker CLI operation in an environment, run the selected
executable's --version. Before the first use of each distinct command path,
run its exact --help (or the narrowest command-group help that exposes the
needed subcommand). Use the same launcher and environment for these checks as
for the real operation (for example, uv run labtasker). Treat the installed
help as authoritative when it differs from this Skill; do not guess an option,
subcommand, default, or output shape from memory.
Adapt an existing pipeline on the user's terms
When the user asks to migrate, convert, or adapt an existing pipeline, first set an internal working preference from the conversation: either lead from their existing project or review a concrete migration design they already proposed. This preference controls the agent's behavior; never name it, present it as a mode, or ask the user to select it.
- If they are new to Labtasker, work project-first. Ask about their existing command or function, what varies between runs, expensive setup, independent failure and retry units, current resource launching, dependencies, and output storage. Do not ask them to choose a Task, Worker, route, Queue, or Labtasker deployment. Make those mappings yourself and explain them after the relevant project facts are known.
- If they already propose a concrete Labtasker design, collaborate at that level, correct mistaken mappings, and still recommend a complete design rather than returning the decisions to them.
- If neither is clear, ask naturally whether this is their first Labtasker integration or whether they already have a concrete migration design to work from. Never offer “use Labtasker concepts” as a conversation mode.
Do not ask a classification question when the context already answers it. Ask only one or two decision-changing project questions at a time, inspect the current pipeline when available, then present the existing flow, what stays unchanged, what Labtasker coordinates, and what remains externally owned. Read workers-and-workloads.md for the detailed migration interview and mapping rules.
Use the managed-local path first
Labtasker requires Python 3.10 or newer. In an ordinary POSIX experiment project, install the complete package:
python -m pip install labtasker
# or in a uv project
uv add labtasker
No MongoDB, configuration file, or TCP port is needed. A default Client selects
exact CWD/.labtasker and connects to its derived Unix socket, but it does not
start a process. Explicitly authorize startup on the first operation:
labtasker --auto-start-local-server queue list
Submit one Task:
labtasker task submit \
--name sample-1 \
--args '{"prediction":"cat","reference":"cat"}' \
--route text-eval
The flag is invocation-scoped and idempotent. Later commands connect without
it while the daemon remains healthy. Use global --labtasker-root PATH or
LABTASKER_ROOT to select another exact root; never search parent or VCS
directories.
Run an existing program once for every compatible Task:
CUDA_VISIBLE_DEVICES=0 labtasker loop --route text-eval -- \
python evaluate.py \
--prediction '%{prediction}' \
--reference '%{reference}'
Start another Worker process on each additional resource already allocated by the user. Each Worker executes one Task at a time and asks for another when it finishes. Labtasker does not select the GPU.
Inspect the recorded state and result:
labtasker task list --status succeeded
labtasker task get t_ABCDEFGHIJKL
With a uv project, run these commands through uv run. Use
labtasker config show to inspect the selected endpoint without starting or
contacting a Server.
Keep the working model small
- A Task is one independent job plus its JSON inputs, state, retry count, metadata, and small result.
- A Worker is one user-started process that repeatedly executes compatible Tasks. The Server stores authoritative Tasks and supplementary expiring Worker observations with optional user-defined resource details; it does not discover, allocate, or manage processes or GPU capacity.
- A route is an exact, case-sensitive compatibility label shared by a Task and the implementation allowed to run it.
- A Queue is an independently managed body of Tasks, not a Worker, GPU, model, or route.
Put executable inputs in args, searchable grouping data in metadata, and
compact JSON outputs in result. Save images, videos, checkpoints, trajectories,
and detailed reports outside Labtasker and return their paths, URLs, checksums,
or summaries.
Use progress for one replace-only snapshot of provisional work position,
metrics, or early-stop diagnostics. Keep result for the final successful
output. Progress is not a history series and does not decide or complete the
Task. When a determinate display is useful, use top-level finite numeric
completed and total values with 0 <= completed <= total and total > 0;
all other keys remain workload-defined.
Use a command Worker for an existing executable. Use a Python Worker when a model, dataset, simulator, or evaluator should be initialized once and reused.
Preserve explicit behavior
- Do not infer Worker eligibility from Task args; use routes.
- Do not invent v1 aliases or implicit coercion. CLI objects are strict JSON.
- Inspect before mutating. Use
cancel,requeue, anddeleterather than patching status. - Do not silently start, stop, or reconfigure an externally managed Server. Confirm its ownership and deployment scope first.
- Unix sockets and HTTP are both public Server transports. Require an explicit
labtasker-server serve --connection socket|httpchoice; do not infer one. Prefer direct argv; add a wrapper only when the workload itself needs shell or multi-step logic. - Treat the Server as authoritative. Local run journals are diagnostic records, not a second source of Task state.
Files (labtasker)
-
references
-
deployment-and-capabilities.md 9.1 KB
# Deployment and capability decisions Choose the smallest supported path that fits the deployment. A mechanism being possible with custom code does not make it a Labtasker interface. ## Choose the endpoint | Situation | Canonical path | | --- | --- | | One POSIX project on one machine | Install `labtasker`; explicitly authorize the managed-local daemon once, then use its root-derived socket. | | Several machines or users share work | Run one explicit HTTP Server and point every Client and Worker at it. | | SQLite database on NFS, WekaFS, Lustre, or uncertain storage | Run one explicit Server with `--database-filesystem shared`; externally guarantee one owner across nodes. | | Windows Client or Python Worker | Use an explicit HTTP Server running on a POSIX host; the Windows Client path is best effort. | | Client-only environment | Install `labtasker-client`. | | Dedicated Server environment | Install `labtasker-server`. | | One same-user host needs an explicitly operated Unix endpoint | Run `serve --connection socket`; use its root-derived default or an explicit `--socket`. | The `labtasker` convenience package installs matching Client and Server distributions and is the default for local use. All packages require Python 3.10 or newer. Managed local mode is selected only when no URL or socket is configured. It is bound to the exact canonical Labtasker root when the Client is constructed. The root resolves from explicit input, `LABTASKER_ROOT`, then exact `CWD/.labtasker`; it does not search parent directories or a repository root. Importing `labtasker`, constructing a Client, displaying help, and running `labtasker config show` do not create state, start a process, or contact a Server. Ordinary operations only connect. `--auto-start-local-server` or `Client(auto_start_local_server=True)` explicitly authorizes one managed-local Client invocation to start or recover the standard daemon. Local management commands address one exact root: ```bash labtasker-server status --labtasker-root .labtasker labtasker-server logs --labtasker-root .labtasker labtasker-server stop --labtasker-root .labtasker ``` There is no public `start`. To launch directly, use `labtasker-server serve --connection socket --daemon --labtasker-root .labtasker`. Repeating an identical detached launch is idempotent. A conflicting launch fails and requires an explicit stop first. ## Run a multi-host HTTP Server Run the Server in the foreground under a process supervisor owned by the user: ```bash # Configure LABTASKER_SERVER_TOKEN through the supervisor's secret mechanism. labtasker-server serve \ --connection http \ --host 0.0.0.0 \ --port 8000 \ --database /data/labtasker.db \ --database-filesystem auto ``` A non-loopback bind requires `LABTASKER_SERVER_TOKEN`; there is no token CLI flag. Configure every Client and Worker with the matching endpoint and token: ```bash export LABTASKER_URL=http://server.example:8000 export LABTASKER_QUEUE=default # Supply LABTASKER_TOKEN through the environment's secret mechanism. ``` The same fields may be placed in the current directory's strict `.labtasker/config.toml`: ```toml url = "http://server.example:8000" queue = "default" ``` The config file also accepts `token`, but prefer a protected environment or secret manager so a credential is not committed with the project. Authentication is one Server-wide bearer token for every application API call; `/health` and `/openapi.json` remain unauthenticated for discovery. Labtasker does not provide users, roles, separate user tokens, per-Queue ACLs, or authenticated Worker identities. Observation IDs identify loop invocations, not security principals. A Queue is a scheduling namespace, not a security boundary. Use separate Server trust domains or external network/authentication controls when different groups require isolation. Endpoint selection is atomic: explicit URL/socket overrides the environment URL/socket layer, which overrides the root config URL/socket layer, then managed local. Queue and token follow ordinary explicit, environment, config, default precedence. Tokens are sent only for HTTP. Do not commit tokens or put them in commands, Task data, or logs. Top-level Python functions share one lazily created default Client. When one process must operate against several endpoints or Queues, construct explicit `Client(url=..., queue=..., token=...)` instances instead of changing global environment variables between calls. An explicit HTTP or socket Client never starts, stops, restarts, or otherwise supervises the Server. Every public `serve` requires `--connection http|socket`; `--daemon` changes only lifecycle. Run exactly one Server process for each SQLite database file and do not use multiple Uvicorn workers. Transport, lifecycle, and database filesystem are independent choices. For example, one cluster node may own a database on NFS while same-host Clients use an explicitly managed socket daemon: ```bash labtasker-server serve \ --connection socket \ --daemon \ --labtasker-root /var/tmp/my-run/labtasker \ --database /shared/project/server.db \ --database-filesystem shared ``` Use HTTP instead when Clients run on other nodes. In both forms, an external policy must prevent another node from starting a Server for the same database. `--database-filesystem auto|local|shared` selects the SQLite strategy. Local uses WAL/FULL. Shared uses DELETE/EXTRA, one pooled connection, and serializes all read and write transactions. Auto maps recognized local filesystems to local, recognized NFS/WekaFS/Lustre-style storage to shared, and unknown storage to shared with a warning. Detection is not a correctness proof: the operator must still prevent cross-node duplicate Servers and validate storage locking and durability behavior. ## Diagnose version differences `Client.server_version` reports the normalized Server package version from the latest ordinary API response, or `None` if its version header is absent or invalid. Reading it makes no request; `/health` is not a feature handshake. When the Server is older than the Client, each Client emits an advisory warning to stderr once per distinct older version. The warning does not change success, exit status, retries, or Worker execution, and does not prove why a request failed. There is no automatic feature fallback or generic version gate. Inspect the actual API error when an operation fails; upgrade a user-owned deployment when needed, respecting shared Server ownership. A missing version is unknown, not evidence of incompatibility. ## Respect platform boundaries - Linux is the fully supported and release-gated platform. - Ordinary HTTP Client and Python Worker behavior is best effort on macOS and Windows. - Every Server mode requires POSIX advisory file locking. Server operation is best effort on macOS and unsupported on Windows, including foreground HTTP. Run the Server on a POSIX host and connect Windows Clients over HTTP. - Managed local mode and external Unix-socket Clients require POSIX and are unsupported on Windows. - Command Workers are unsupported on Windows because Labtasker cannot guarantee whole-process-tree cancellation there. They fail before Server access, Task claim, journal creation, or child startup. - Single-node `torchrun` and Accelerate use the Command Worker boundary and are therefore also unsupported on Windows. Do not recommend trying an explicitly unsupported path and waiting for a system call to fail. Choose the HTTP/Python alternative or move command execution to a supported POSIX host. ## Answer capability questions directly | Request | Labtasker answer | | --- | --- | | Allocate, reserve, or discover a free GPU | No. The user, shell, or cluster scheduler starts Workers on allocated resources. | | See online Workers, route presence, or user-reported latest resource details | Yes. Bundled Workers automatically report expiring observations and may include static metadata plus latest telemetry; use `list_workers` / `count_workers` or `worker list` / `worker count`. Read [observations-and-counts.md](observations-and-counts.md). This is approximate observation, not discovered capacity, allocation, or process control. | | Schedule machines, pods, or multi-node rendezvous | No. Use SLURM, Kubernetes, or another external scheduler. | | Express Task dependencies or a workflow DAG | No. Use a workflow engine and submit independent leaves to Labtasker. | | Store checkpoints, images, videos, or datasets | No. Use project or artifact storage and record references in Task data. | | Run one Task with single-node `torchrun` or Accelerate | Yes, through one outer Command Worker on supported POSIX platforms. | | Provide an asynchronous Python Client | No. The public Python API is synchronous. | | Run several Tasks concurrently in one Worker process | No. Start more Worker processes; each executes at most one Task at a time. | | Keep an AI agent in the runtime loop | No. Agents may configure and operate Labtasker, but Worker execution is autonomous. | Labtasker coordinates independent work after processes exist. If the main problem is resource allocation, dependent pipelines, or artifact management, select another primary tool rather than building those concepts around Labtasker. -
observations-and-counts.md 9.4 KB
# Worker observations and grouped counts ## Inspect presence without inferring ownership Bundled Python and Command Workers automatically report one observation per Worker loop invocation, including idle periods. No registration setup or stable Worker name is needed. A fresh invocation gets a fresh `w_` ID, even in the same process; reconnects retain that ID. An outer Command Worker is observed once; its command child and distributed ranks are not separate Workers. Restarted and old observations can briefly overlap. Observations renew every 60 seconds and expire 300 seconds after the last accepted report, with best-effort activity-change reports. These timings are independent of Task heartbeats and ownership leases. Only unexpired observations appear in Worker queries. Use the Server-provided `expires_at` for freshness. There is no offline history, Worker get, automatically discovered hostname/PID/GPU inventory, remote start/stop/restart command, or resource allocator. A Worker may instead report user-defined invocation metadata and one latest telemetry snapshot. - `idle`: waiting for work, including an unconfirmed claim response. - `busy`: a confirmed claim occupies the Worker through execution, terminal reporting, and cleanup. Accepted `finish()` may make the Task succeeded while its Worker stays busy until local execution and cleanup end. Public fields are `id`, `queue`, `route`, `status`, nullable advisory `task_id`, `metadata`, nullable `telemetry` and `telemetry_updated_at`, `last_seen_at`, and `expires_at`. Timestamps are Server-generated UTC. `metadata` is fixed for one Worker invocation; telemetry is the latest complete user-defined object. The Task reference may be terminal, stale, or deleted; it is not an ownership token. Route presence means any idle or busy observation for that exact route in that Queue. It does not prove spare capacity, compatibility of undocumented settings, or the absence of other execution processes when zero Workers are observed. Report failures never block startup, claims, execution, Task reports, or the failure guard. A lost observation must not trigger cancellation or requeue: Tasks remain authoritative under their own lease and `run_id` fencing. Once the loop has independently decided to exit, it attempts withdrawal, waiting at most one second. This does not limit execution/cleanup or initiate exit. Crashes and failed withdrawals fall back to expiry. Worker observations do not prevent Queue deletion; Task-based deletion rules still apply. ## Read Workers ```bash labtasker worker list --queue experiments --filter 'status == "busy"' --limit 100 labtasker worker list --queue experiments --filter 'metadata.node == "node-a"' labtasker worker list --queue experiments --filter 'telemetry.gpu_util_pct < 20' labtasker worker count --queue experiments --group-by route,status ``` ```python from labtasker import Client with Client(queue="experiments") as client: page = client.list_workers(filter='status == "busy"', limit=100) total = client.count_workers() # int groups = client.count_workers(group_by=["route", "status"]) ``` Module-level `labtasker.list_workers()` and `labtasker.count_workers()` expose the same interface. Listing returns `WorkerPage(items, next_cursor)` containing `WorkerObservation` values. Follow every non-null cursor with the same Queue and exact filter. Lists are ordered by ID lexicographically ascending, with no `order_by` option. Page sizes default to 100, maximum 1000. Reads are live, not a snapshot spanning pages. Worker filters use the existing expression language over fixed observation fields plus nested `metadata.*` and `telemetry.*` paths. For example, `route == "judge" and status == "idle"`, `metadata.node == "node-a"`, or `telemetry.gpu_util_pct < 20`. Use `filter=...`, not a separate Worker `status=` selector. Task paths such as `args`, `progress`, and `result` are not Worker fields. Telemetry keys are user-defined; enumerate the normally small Worker set and process it locally when a richer analysis is needed. An API/transport error means the query failed, never an empty Worker inventory. ## Interpret and report Worker resource details Worker metadata and telemetry are abstractions, not a built-in NVIDIA, hostname, SLURM, or Kubernetes schema. Interpret only keys the workload defines. For example, metadata may record `node` and `gpu_ids`, while telemetry may report `gpu_util_pct` and `memory_used_gb`. A low latest value suggests underuse at that sample; it does not establish spare schedulable capacity. Telemetry reports synchronously replace the complete previous object. They do not merge fields, retain history, renew observation expiry, affect Task outcome, or influence scheduling. `telemetry_updated_at` is the Server acceptance time. Missing fields in the latest object are absent, not inherited from an earlier sample. Command descendants and distributed ranks share the outer Worker's ID; the last report committed by the Server is visible. Python code in an active Worker execution can call: ```python labtasker.report_worker_telemetry({"gpu_util_pct": 92, "memory_used_gb": 38}) ``` The helper performs one best-effort synchronous report and returns whether it was accepted. Labtasker does not sample, retry, throttle, merge, or store history; callers own periodic scheduling. Static placement belongs in Worker metadata, supplied through Python `loop(metadata={...})` or Command Worker `labtasker loop --metadata JSON -- COMMAND`. ## Count selected Tasks and Workers ```bash labtasker task count --status pending --group-by routes labtasker task count --filter 'metadata.batch == "pilot"' --group-by status,routes labtasker worker count --group-by route ``` ```python with Client() as client: demand = client.count_tasks(status="pending", group_by=["routes"]) presence = client.count_workers(group_by=["route", "status"]) ``` To diagnose waiting work, fully read pending Task groups by `routes` and independently read Worker groups by `route`, then align those results locally. There is no joined summary or Route registry. Worker-only routes may also be shown. Worker presence and Task demand can change between these reads. | Operation | Allowed grouping dimensions | | --- | --- | | `count_tasks` / `task count` | `routes`, `status`, or both in either order | | `count_workers` / `worker count` | `route`, `status`, or both in either order | Task `routes` means compatible route membership, including for running Tasks; it does not mean the route that actually executed the Task. There is no grouping by `last_route`, metadata, telemetry, arbitrary expressions, or additional metrics. Dynamic Worker metadata/telemetry may be filtered but not grouped. Filter on supported fields before aggregation instead. Without grouping, Python returns an integer and HTTP/CLI return `{"count": n}`. Do not pass `limit` or `cursor` for an ungrouped count. Python grouping takes an ordered list or tuple, never a comma-separated string. CLI/HTTP grouping takes one comma-separated value without spaces, for example `--group-by routes,status`. Empty values, duplicates, whitespace, unsupported fields, and repeated CLI options are errors rather than normalized inputs. Grouped Python calls return `GroupCountPage` with `CountGroup` items. HTTP/CLI return the same structure as JSON. Example for two pending Tasks accepting `["a", "b"]` and `["a"]`: ```json { "group_by": ["routes"], "count": 2, "items": [ {"key": {"routes": "a"}, "count": 2}, {"key": {"routes": "b"}, "count": 1} ], "next_cursor": null } ``` The top-level `count` is the complete deduplicated selected total, not a page subtotal. Never sum overlapping Task route groups to obtain a Task total. Each key is an object with string values; `group_by` defines dimension order. Groups sort lexicographically by those dimensions, including status (not lifecycle order). Only nonzero groups are returned; a missing group can mean zero only after all relevant pages have been read. ## Follow group pages Grouping is computed over the complete Server-side selection before pagination. `limit` counts groups, defaults to 100, and must be 1–1000. CLI returns one page; it does not automatically fetch the rest. For example: ```python with Client() as client: selection = dict( status="pending", filter='metadata.batch == "pilot"', group_by=["status", "routes"], ) cursor = None while True: page = client.count_tasks(**selection, limit=100, cursor=cursor) for group in page.items: print(group.key, group.count) cursor = page.next_cursor if cursor is None: break ``` Use the same operation, Queue, exact filter/Task selectors, and ordered grouping fields for continuation. Page size may change. Do not reuse list cursors for counts, Task cursors for Workers, or cursors with reordered dimensions or a rewritten filter. Cursors are opaque. Invalid/mismatched cursors are errors. Each page reads current data: its total and groups may change while paging, so combining pages does not give an atomic snapshot. Do not add page totals. Old Clients do not report observations. New Workers can execute against older Servers even when observation calls fail. A grouped request receiving an old scalar count response raises `TransportError`; do not display zero or silently aggregate a partial Task list as a replacement. Report the unavailable query and use matching Client/Server versions when these inspection features are needed. -
operations-and-recovery.md 15.2 KB
# Task operations and recovery ## Submit repeatably Before submission, follow the route-record and user-confirmation workflow in [workers-and-workloads.md](workers-and-workloads.md#record-route-settings-and-confirm-reuse-before-submission). CLI `--args`, `--metadata`, and `--changes` each accept one strict JSON object. The CLI does not infer types from repeated `key=value` options. Use a caller-chosen ID when submission must be safe to repeat across process restarts: ```bash labtasker task submit \ --id t_AbCdEf0123-_ \ --name baseline-seed-1 \ --args '{"seed":1,"enabled":true}' \ --route train ``` Task IDs are opaque: `t_` followed by exactly 12 ASCII letters, digits, underscores, or hyphens. Put the readable experiment label in `name`. Submitting the same normalized complete definition with the same ID is idempotent and returns the Task's current representation. Reusing the ID with a different definition is a conflict, never an update. The Client generates an ID before its own first network attempt when one is omitted, but an external tool that must retry after its own restart should persist and reuse its chosen ID. JSON object-key order, Task route input order, and explicitly spelling a submission default do not change the normalized definition. Later lifecycle or user-field updates also do not change the original creation definition used to recognize a retry. Changing a real submitted value such as `max_attempts` is a conflict. The synchronous Python counterparts are `submit_task`, `get_task`, `list_tasks`, `count_tasks`, `update_task`, `update_tasks`, `cancel_task`, `requeue_task`, and `delete_task`. They follow the same defaults and lifecycle rules as the CLI. For a large Python submission loop, reuse one explicit Client and close it deterministically. There is no bulk-submission endpoint: ```python from labtasker import Client with Client() as client: for task_id, seed in persisted_work: client.submit_task( {"seed": seed}, task_id=task_id, routes=["train"], ) ``` Persist each caller-chosen ID with its complete Task definition before the submission attempt when the loop must resume safely after its own process restarts. Package-level functions already reuse one lazy default Client, but an explicit context-managed Client is the canonical choice for a large loop, deterministic cleanup, tests, or more than one endpoint. Do not invent a `submit_tasks` API. ## Inspect all selected Tasks ```bash labtasker task count --status pending labtasker task list \ --filter 'status == "failed" and missing(result.score)' \ --limit 100 labtasker task get t_ABCDEFGHIJKL ``` `task list` returns one page. If `next_cursor` is non-null, request the next page with the same Queue, filter, selectors, and ordering: ```bash labtasker task list \ --filter 'status == "failed" and missing(result.score)' \ --limit 100 \ --cursor OPAQUE_CURSOR ``` Python returns a `TaskPage` with `items` and `next_cursor`: ```python page = labtasker.list_tasks(filter=filter_expr, limit=100) tasks = list(page.items) while page.next_cursor is not None: page = labtasker.list_tasks( filter=filter_expr, limit=100, cursor=page.next_cursor, ) tasks.extend(page.items) ``` Filters support comparisons, `and`/`or`, scalar candidate lists, array containment, and `exists(path)`/`missing(path)`. Literal null and Booleans use `None`, `True`, and `False`. Every comparison requires the path to exist, so use `missing(result.score) or result.score < 0.5` to include absent scores. Use `exists(metadata.owner)` to check an object key. `"owner" in metadata` is not object-key membership and is invalid. `"baseline" in metadata.tags` means array containment. General unary `not (...)` is unsupported; use explicit `!=`, `not in`, `exists`, or `missing` forms. For a remembered experiment name, use `labtasker task list --name-fuzzy 'tr ev'` or Python `list_tasks(name_fuzzy="tr ev")`. The same selector works with `task count` and `count_tasks`, including grouped counts. Matching case-folds Unicode, splits the query into whitespace-separated words, and requires each word's characters to occur in order in the name. Words match independently: `tr ev` and `EV TR` both match `train_model_eval`. This is subsequence search, not substring search, regex, wildcard matching, or relevance ranking. Exact `name`, `name_fuzzy`, `status`, and `filter` combine with AND. Empty or whitespace-only fuzzy input adds no restriction; nonempty input excludes absent or empty names. Keep the raw fuzzy input unchanged across pagination, including case and whitespace. Fuzzy search is a list/count selector, not a filter-language function or mutation selector. Inspect the matches and use explicit IDs or a supported filter for subsequent changes. For online Worker queries and Server-side Task grouping by routes/status, read [observations-and-counts.md](observations-and-counts.md). ## Apply an external early-stop policy from progress Labtasker stores the latest `progress` snapshot but does not choose an early-stop policy. A controller must use the experiment's explicit rule and scope, inspect current authoritative Task state, and cancel only the exact running Task IDs that meet that rule. Never infer a threshold, compare unrelated experiment groups, or cancel Tasks merely because progress is missing or stale. In Python, inspect every page of the intended running selection before mutating: ```python import labtasker page = labtasker.list_tasks( status="running", filter='metadata.experiment == "sweep-7"', limit=100, ) tasks = list(page.items) while page.next_cursor is not None: page = labtasker.list_tasks( status="running", filter='metadata.experiment == "sweep-7"', limit=100, cursor=page.next_cursor, ) tasks.extend(page.items) # Apply the user-defined policy to task.progress, then review exact IDs. selected = [task for task in tasks if user_policy(task.progress)] for task in selected: labtasker.cancel_task(task.id) ``` The same inspection is available through `labtasker task list` and `labtasker task get`; `labtasker task cancel TASK_ID` requests cancellation for one selected Task. Filters may address dynamic paths such as `progress.metrics.validation_loss`, but missing paths do not match comparisons. Use the Server-provided `progress_attempt` and `progress_updated_at` when the policy requires attempt or freshness checks. A retained terminal snapshot is diagnostic, not evidence that the Task is still running. Cancellation immediately changes the authoritative Task to `cancelled` and fences its `run_id`. Python Worker code should poll `cancellation_requested()` at safe boundaries, save any external checkpoint it needs, and return. Command Workers terminate the child process group under their configured force-stop contract. The last accepted progress snapshot remains available for diagnosis; it is not copied into `result`. ## Prioritize and update pending work Workers claim higher `priority` first. Equal-priority pending Tasks keep stable pending order. Priority affects only future claims; it never interrupts a running Task. Inspect a selection before changing it. Update one non-running Task: ```bash labtasker task update t_ABCDEFGHIJKL --changes '{"priority":20}' ``` Or atomically update all matching non-running Tasks on the Server: ```bash labtasker task update \ --filter 'status == "pending" and "clip-openai" in routes' \ --changes '{"routes":["clip-openai","clip-openclip"]}' ``` `args`, `metadata`, `result`, and `routes` are complete replacements, not merges. Unspecified fields remain unchanged. Running Tasks cannot be updated. Do not implement a bulk change as a client-side list/update loop when the filtered bulk operation expresses it directly. Bulk update is one Server transaction. A concurrent claim either sees the complete new values, or wins first and excludes that now-running Task. The result's `matched` count includes rows that satisfied the filter and remained non-running at execution time; `updated` counts only rows whose stored values actually changed. All matching non-running rows are validated before any write, so one state-dependent conflict rolls back the complete batch rather than silently skipping an invalid row. There is no Server-side expression that merges a different existing object for each Task. When preserving those differences is required, read each Task, construct its complete replacement, and use ID-addressed `update_task`. This read-modify-write path accepts last-write-wins if another caller updates the same Task concurrently; v2 has no revision or compare-and-swap field. ## Use explicit lifecycle actions ```bash labtasker task cancel TASK_ID labtasker task requeue TASK_ID labtasker task delete TASK_ID ``` - Cancel accepts pending or running Tasks. Server cancellation is immediate and fences an active run even if local cleanup continues. - Requeue accepts pending, failed, or cancelled Tasks, returns them to pending, resets `attempt` to zero, and clears the last error. - A succeeded Task cannot be requeued. Submit a new Task to rerun successful work. - A running Task cannot be updated, requeued, or deleted. Cancel it first, then requeue if a new execution is wanted. - Delete permanently removes one non-running Task. Do not delete unless the user's intent is explicit. Queue deletion is separate. A non-empty Queue requires explicit cascade deletion, and even cascade is rejected while any Task in the Queue is running. Cancel running Tasks first. A successful cascade atomically deletes the Queue and all of its Tasks. If Queue `default` is explicitly deleted, later requests do not recreate it; create it again explicitly if it is still wanted. Create and select another independently managed Queue only when needed: ```bash labtasker queue create paper-a labtasker queue list labtasker task list --queue paper-a labtasker queue delete paper-a --cascade ``` The Python Queue API is: ```python labtasker.create_queue("paper-a") queues = labtasker.list_queues() labtasker.delete_queue("paper-a", cascade=True) ``` Use `--queue`, the Python `queue=` argument, `LABTASKER_QUEUE`, or the config file to select it. Do not use Queues to represent Workers, GPUs, models, or routes. ## Choose the failure level In a Python Worker: | Situation | Raise | Task effect | Worker effect | | --- | --- | --- | --- | | Temporary infrastructure incident | `TransientError` | Return to pending without charging the incident | Continue | | Bad Task or ordinary execution failure | `TaskError` or an ordinary exception | Charge the attempt; retry or become failed | Continue | | Worker process is no longer trustworthy | `FatalWorkerError` | Charge the attempt; retry or become failed | Exit | The Continue outcomes above are subject to the local [consecutive-failure guard](workers-and-workloads.md#choose-worker-lifetime-deliberately). A charged failure returns to pending while `attempt < max_attempts`; otherwise it becomes failed. `TransientError` rolls back only the current claim's attempt increment; it does not erase older charged failures. Both a retryable charged failure and a transient return re-enter the end of their priority group, so already-waiting equal-priority Tasks run first. ## Recover safely A healthy long Task has no execution timeout. Every claim gets a private `run_id`, renews a five-minute lease with a heartbeat once per minute, and may run as long as heartbeats continue. Stopping or restarting the Server does not itself rewrite running Tasks. While an active Server is temporarily unavailable, Workers keep their local execution and retry heartbeat and terminal-report transport. A restart shorter than the remaining lease can therefore be transparent. Before a restarted Server begins serving, it applies the ordinary expiry transition to leases already past their deadline; heartbeat loss is a charged failure, not a special restart state. Startup checks retain short bounded retries so bad configuration fails promptly. After startup succeeds, claim has a separate fixed five-minute recovery window. Transport failures, `database_busy`, and Server 5xx responses preserve the same private `run_id`, pause `idle_timeout`, and use bounded jittered backoff. Never interpret them as an empty Queue. Only an explicit empty claim advances healthy idle time and causes the next poll to use a fresh token. If a claim succeeds after an uncertain response, the Worker renews that run before starting user code or a command child. A claim confirmed stale or finalized is discarded, and the Worker continues claiming without charging a workload failure. Five minutes of continuous claim unavailability exits the Worker for an external supervisor to handle; a long `idle_timeout` does not extend that fault window. Observation errors remain isolated from all phases and cannot stop startup, claiming, execution, or alter outcomes. Confirmed cancellation or ownership loss during execution still requires stopping or cooperatively cancelling the old execution. Network errors inside user code remain the workload's responsibility. When a Worker disappears, lease recovery returns the Task to pending or marks it failed according to its retry budget. Recovery is normally committed roughly five to six minutes after the last accepted heartbeat. A later heartbeat or completion from the stale Worker has the wrong `run_id` and cannot overwrite a newer run. For cancellation or other revocation, Python code may poll `cancellation_requested()`. A Command Worker terminates the child process group. The default `force_stop_timeout=None` waits indefinitely for safe cleanup; set a finite timeout only when forced termination is acceptable. Once `finish(result)` is accepted, a later exception or nonzero command exit is only a local diagnostic and cannot change succeeded state. Call `finish()` when the result must be accepted before cleanup continues. The local `.labtasker/runs/` journal contains the Task snapshot, output, state, and prepared terminal report for diagnosis. The Server remains authoritative; v2 does not reconstruct Server state from journals. It also does not provide automatic journal retention, compression, or cleanup, and deleting a Task or Queue does not delete project files or external artifacts. ## Automate the CLI safely Successful finite Task, Queue, and configuration commands keep requested data on stdout as one two-space-indented JSON document with no ANSI styling. Endpoint selection, local daemon startup or reconnection, and other diagnostics go to stderr, so redirecting stdout remains machine-readable. Successful delete commands are quiet on stdout. Handled configuration, transport, and API errors write one stable structured error envelope to stdout and exit `1` without an application traceback. Diagnostics remain on stderr. For finite commands, parse stdout as the response channel and use the exit status and top-level `error` key to distinguish a successful value from an error. Do not assume JSON stdout implies success. CLI argument or usage errors instead write natural-language stderr and exit `2`; an interrupted Worker retains the conventional `130`. `labtasker loop` is different: it is a long-running supervised process that writes ordinary timestamped operational logs and relays user-code output. It does not produce one JSON document or a JSONL event stream. -
workers-and-workloads.md 19.5 KB
# Workers and workload design ## Lead a pipeline migration For an existing pipeline, learn the workload before naming Labtasker abstractions. Inspect the current entry point and configuration when the project is available. Treat the questions below as a decision ladder, not an intake checklist. Ask at most one or two unanswered questions that can change the next recommendation: - What command or function performs one run, and which input values change? - What work can fail and be retried independently without corrupting outputs? - Which model, dataset, simulator, or other expensive state could stay loaded across runs? - Then ask how resources are launched, whether stages depend on one another, where outputs live, or how retries should behave only when the described workflow makes that fact relevant to the immediate decision. Treat facts stated by the user or visible in the repository as answered. When risk is low, state a reasonable assumption and give a conditional recommendation instead of waiting for a complete questionnaire. Do not ask about stage dependencies for a single-stage program or ask about machine scope when the launcher already answers it. Do not ask a newcomer “Which Worker type?”, “What route?”, “How many Queues?”, or “Which Labtasker deployment?” Translate the facts into a recommendation: | Project fact | Recommended mapping | | --- | --- | | One independently retryable run | One Task | | Values that change the execution | `args` | | Searchable experiment or batch labels | `metadata` | | Small metrics and artifact references | `result` | | Existing executable with acceptable per-run startup | Command Worker | | Expensive reusable in-process state | Python Worker | | Equivalent executors for the same work | One shared route | | A separately managed body of work | One Queue | | GPU, node, Pod, or process allocation | Existing launcher or scheduler | | Stage dependencies and barriers | Existing workflow controller | | Large or durable outputs | Existing shared or object storage | When a fact leaves a real trade-off, ask about the consequence rather than the Labtasker mechanism. For example, ask whether avoiding repeated model loading is worth a small Python refactor; do not ask the user to choose between a Command Worker and Python Worker. Before editing code, summarize the proposed migration in project terms first: 1. the current entry point and independent work item; 2. what code, launcher, and storage remain unchanged; 3. what Labtasker will submit, distribute, retry, and record; 4. what stays owned by the resource scheduler, workflow system, or artifact store; and 5. the concrete Task, Worker, route, Queue, and deployment mapping, with reasons. ## Map an experiment Submit each independently retryable case as a Task. In AIGC this may be one prompt/seed/checkpoint/ablation combination. In embodied-AI evaluation it may be one benchmark suite or subtask. Avoid fixed GPU shards when runtimes vary: start one Worker process on each already allocated resource, and let each process take another Task when it finishes. Use: - `args` for values the implementation executes; - `metadata` for searchable grouping such as benchmark, checkpoint, or sweep; - `result` for compact JSON metrics and external artifact references; - `priority` to choose urgent pending work first; and - `max_attempts` for charged execution attempts. Use one Queue for one independently managed body of work. Do not create a Queue per GPU, Worker, model, or implementation. ## Design routes explicitly A Worker declares exactly one route. A Task declares one or more routes and is eligible only when the Worker's exact route is in that list. Matching is case-sensitive; `SDXL` and `sdxl` differ. The default route on both sides is `default`. Routes have no wildcard, regular expression, negation, priority, or fallback syntax. They are not registered resources and do not prove that an implementation is online. Prefer human-readable implementation or workload names such as `robotwin`, `clip-openai`, and `clip-openclip`; avoid names consisting only of a hash or random string so people can recognize what work a route accepts. For a rollout, run separate Workers for the old and new routes. A Task may list both when either implementation is acceptable: ```python labtasker.submit_task( {"image": "outputs/001.png"}, routes=["clip-openai", "clip-openclip"], ) ``` Starting a new Worker never changes old Tasks. To let a new implementation help with a pending backlog, explicitly replace the selected Tasks' complete routes list. Running Tasks cannot be updated. ### Record route settings and confirm reuse before submission Keep a durable route record in the experiment project's existing route document, or use `experiments/labtasker-routes.md` when none exists. Reuse that same document across agent sessions. This is a project convention, not a Server-side registry. For each route, record: - its exact name, purpose, and Server/project and Queue scope, without credentials; - the first confirmed submission date, including timezone; leave it unknown for historical routes when evidence is unavailable; - the concrete execution settings: implementation and entry point, fixed Worker configuration, model/checkpoint revision, relevant dependencies or environment, and any resource requirements that affect compatibility; - the code repository and commit, plus any uncommitted changes that affect execution; a branch name alone is not a reproducible version; - the accepted Task args and their allowed variation, expected outputs, and compatibility conditions; link to durable configuration files where useful; - the parameter combinations the Worker needs: record concrete fixed startup arguments and the required per-Task args, their types, defaults, and supported combinations or constraints. Include a reusable command or configuration example, distinguishing fixed values from values that may vary per Task; - dated compatibility decisions or revisions, preserving earlier settings rather than silently overwriting them. Representative Task IDs are not needed. Before submitting a batch: 1. Read the route record and inspect all pages of online Worker observations and running Tasks in the target Server and Queue. Summarize observed Worker routes/activity, Task `routes`, and relevant recorded settings. Read [observations-and-counts.md](observations-and-counts.md) for presence limits; follow the pagination guidance in [operations-and-recovery.md](operations-and-recovery.md). A running Task's routes list describes acceptable implementations, not which route its current Worker actually uses, and is not an inventory of Workers. 2. Compare the proposed workload with those routes and the recorded historical routes. Never decide compatibility-based reuse on the user's behalf. Present the candidate route, matching settings, and any differences, then explicitly ask: "你提交的实验似乎和 `xxx` route 类似,可能是同一组实验,是否复用?" Adapt the wording to the user's language and replace `xxx` with the actual route. Wait for explicit approval of that route for the proposed batch before submitting with it. Similar settings, historical reuse, a route record, or a general request to submit experiments is not consent to reuse. A shared experiment name or broad task category is also insufficient. Ask once for the batch, not per Task; an explicit approval already given for this exact batch and route remains valid while the relevant settings are unchanged. 3. If there are no running Tasks, say so and check the document for reusable routes; absence of running Tasks does not imply a route is obsolete or has no available Worker. If inspection fails or settings are unknown, disclose the gap and ask the user to resolve it before submission rather than assuming a match or treating failure as an empty result. 4. Reuse a route only with the user's explicit approval and when its executors can accept the new Task inputs and produce acceptable outputs under the documented settings. Changes to ordinary per-Task values within that contract do not require a new route. Incompatible implementation, configuration, or output changes need a distinct readable route; never silently redefine an old route that existing Workers still use. 5. Prepare the route entry before submission, marking an unsubmitted entry as planned. After the first confirmed successful submission, record its date. On reuse, preserve the original first-submission date and record any newly confirmed compatible settings. Do not invent missing revisions or dates. Keep this record available to subsequent agent sessions. ## Wrap an existing command Use the required `--` separator followed by one argv template: ```bash CUDA_VISIBLE_DEVICES=0 labtasker loop \ --route robotwin \ --metadata '{"node":"node-a","gpu_ids":["0"]}' \ -- \ python evaluate.py \ --task '%{task}' \ --checkpoint '%{checkpoint}' ``` Labtasker executes argv directly. It does not invoke a shell, join or split arguments, expand `$VARS`, or interpret pipes and redirections. Each `%{path}` resolves to exactly one argv element, even when it contains spaces. A selected JSON string is inserted directly. Other JSON values, including numbers, Booleans, null, arrays, and objects, become compact deterministic JSON inside that one argv element; object keys are sorted. An empty string remains an empty argv element, while NUL cannot be represented and fails binding. Prefer direct argv. If the workload deliberately requires shell syntax, make the shell visible and pass resolved Task values as positional arguments rather than interpolating them into the shell program text: ```bash labtasker loop --route preprocess -- \ bash -lc 'python preprocess.py --input "$1" > "$2"' \ labtasker-shell '%{input}' '%{output}' ``` For `bash -c`, the first argument after the program text supplies `$0`; later arguments supply `$1`, `$2`, and so on. Shell quoting, expansion, pipeline exit behavior, and redirection are then the user's responsibility. Do not add a wrapper merely to reproduce output capture: Labtasker forwards child output live and writes the raw combined output to the run's `run.log`. For a per-Task environment value on POSIX, an explicit external `env` command is simpler than a shell: ```bash labtasker loop --route train -- \ env 'LR=%{lr}' python train.py --seed '%{seed}' ``` Static environment values belong on the Worker process itself. Command placeholder paths traverse JSON objects using dot-separated ASCII identifier segments, such as `%{seed}` or `%{judge.threshold}`. They do not support array indices, hyphenated or Unicode keys, quoted segments, defaults, wildcards, or expressions. Reshape the args or use a Python Worker with `task_info().args` when arbitrary JSON access is required. Use `%{{` when the child must receive a literal `%{` opener. For example, `%{{name}` resolves to the literal text `%{name}` rather than reading a Task argument. Ordinary percent signs and stray closing braces are otherwise literal. Static template syntax errors stop the Worker before it claims anything. A missing key or non-object intermediate belongs to a claimed Task, prevents child startup, and is a normal charged Task failure. Terminal handling is automatic and has no public `--pty` or `--no-pty` option. When the Worker's stdin, stdout, and stderr are attached to an interactive POSIX terminal, Labtasker uses an internal PTY and relays input, output, and terminal size. In a scheduler, pipeline, or redirected run it uses ordinary pipes, drains stdout and stderr concurrently, and connects child stdin to `/dev/null`. Both modes forward output live and append the raw bytes to the current run's `run.log`. Pipe mode preserves separate stdout and stderr destinations for the caller even though the journal contains both. Do not add `tee` merely to obtain the Labtasker run log. Without `finish()`, exit code zero succeeds with `{}` and nonzero is a charged failure. Existing child code may report a structured result: ```python labtasker.finish({"score": 0.94}, skip_if_no_labtasker=True) ``` `finish()` accepts one JSON-compatible object. Convert values such as `Path` to strings and keep NumPy arrays, tensors, and other large data in external storage rather than passing arbitrary Python objects or top-level scalars. Once accepted, `finish()` is stable: later cleanup failure or nonzero process exit cannot rewrite the succeeded Task. Command Workers have no reserved child exit codes or output-text protocol. Every nonzero exit code or signal is the same charged Task failure, while stdout and stderr are only relayed and logged. A child exit code does not become the outer Worker's exit code; after resolving that Task, the Worker normally continues subject to the consecutive-failure guard below. Use a Python Worker when the workload must deliberately choose `TransientError`, `TaskError`, or `FatalWorkerError`. ## Choose Worker lifetime deliberately Both Worker styles wait for newly eligible Tasks after an empty claim. The public `idle_timeout` defaults to 300 seconds, resets after each executable claim, and then ends the Worker normally if no work appears. Set `idle_timeout=0` or CLI `--idle-timeout 0` to exit on the first empty claim; this does not mean “run exactly one Task” when the Queue remains non-empty. Empty polling backs off with jitter to an actual maximum of about 8 through 12 seconds, so a newly submitted Task may not be claimed immediately by an already idle Worker. Temporary claim communication failures do not consume `idle_timeout`; they use an independent fixed five-minute recovery window. There is no infinite-wait value, daemon mode, `once`, `max_tasks`, or automatic Worker restart. Use an external process supervisor when a Worker must be kept available indefinitely or restarted after process failure. Each loop has `max_consecutive_failures=5`; Python `loop()` accepts this keyword and CLI `labtasker loop` exposes `--max-consecutive-failures`. It must be a positive non-Boolean integer, with no disable value or environment setting. Accepted ordinary failures (including binding errors, child startup failures and nonzero exits) and accepted `TransientError` unclaims increment it. Successful completion resets it, including accepted `finish()` followed by ordinary cleanup failure. Empty polls, cancellation and ownership loss leave it unchanged; transport retries and observation failures do not increment it. After reporting and cleanup, reaching the limit raises `FatalWorkerError` before another claim; CLI exits `1`. Explicit `FatalWorkerError` still exits immediately, even after `finish()`. This local guard changes neither Task retry accounting nor fencing. Each invocation starts at zero: configure supervisor restart backoff and frequency limits externally. ## Reuse loaded Python state Use a Python Worker when setup should happen once: ```python import labtasker @labtasker.loop( route="sdxl-diffusers", metadata={"node": "node-a", "gpu_ids": ["0"]}, ) def generate( pipeline, prompt: str = labtasker.TaskArg(), seed: int = labtasker.TaskArg(), steps: int = labtasker.TaskArg(default=30), ) -> None: image = pipeline(prompt, seed=seed, steps=steps) path = labtasker.task_info().run_dir / "image.png" image.save(path) labtasker.finish({"image": str(path)}) generate(load_pipeline_once()) ``` Only parameters whose default is `TaskArg(...)` bind from Task args. Other arguments are fixed when the Worker starts. Binding uses the annotation's strict schema: an `int` does not accept a string, float, or Boolean, and Labtasker does not add casts. A `TaskArg(default=value)` default passes through the same resolver and annotation validation as a submitted value. `TaskArg(path="judge.threshold")` selects a nested object field. A resolver receives that selected raw JSON value, returns a value, and the annotation then validates the return. Extra Task args are ignored by named binding, including when the handler declares `**kwargs`; that parameter receives only ordinary keyword arguments supplied when the Worker starts. Read the complete object through `task_info().args`. Python Worker handlers and `TaskArg` resolvers must be synchronous. An `async def` handler, an object with an asynchronous `__call__`, an `async def` resolver, an object with an asynchronous resolver `__call__`, an otherwise invalid static handler definition, an unusable annotation, or a non-callable resolver fails before the first claim. A particular Task's missing value, type mismatch, or resolver failure happens after claim and is a normal charged Task failure. Argument shape never affects Server eligibility; Queue, pending state, and route decide the claim. A normal return succeeds with `{}`. ## Report current progress separately from the result An active Python program may publish one compact latest snapshot without completing its Task: ```python labtasker.report_progress( { "completed": completed_cases, "total": total_cases, "metrics": {"validation_loss": validation_loss}, }, skip_if_no_labtasker=True, ) ``` This works inside a Python Worker and inside Python launched by a Command Worker. Each report replaces the complete previous `progress` object. It does not merge, complete the Task, renew the lease, or create a history series. Keep the final successful summary in `finish(result)` and keep metric history, artifacts, and checkpoints in their existing external systems. Report at useful evaluation or checkpoint boundaries rather than every inner-loop step. The object has no required business keys. For a determinate Labtasker WebUI indicator, use top-level `completed` and `total` only when both are finite numbers, `0 <= completed <= total`, and `total > 0`. Metrics, best-so-far values, and early-stop diagnostics may use any other JSON keys. The Server records the report time and attempt. It retains the last snapshot after success, failure, unclaim, expiry, or cancellation for diagnosis, then clears it when a new claim starts. Reporting is supplementary and best effort. The Python helper returns `True` when accepted and `False` for an isolated transport or Server rejection; such a failure must not fail the workload. Invalid data or missing execution context is a programming error unless `skip_if_no_labtasker=True` handles the latter. Use Worker telemetry, not Task progress, for a latest resource/load snapshot shared across the Worker's successive Tasks: ```python labtasker.report_worker_telemetry( {"gpu_util_pct": gpu_utilization, "memory_used_gb": memory_used_gb} ) ``` The synchronous call completely replaces the previous Worker telemetry object and returns whether the Server accepted it. It does not renew Worker presence or affect the Task. Labtasker does not detect resource fields, sample periodically, retry, merge, or keep history. If periodic sampling is needed, user code owns the thread or schedule. All distributed ranks inherit one Worker ID and replace the same snapshot. ## Use single-node distributed launchers Keep one Labtasker Command Worker outside a single-node launcher: ```bash labtasker loop --route robotwin -- \ torchrun --nproc-per-node=8 evaluate.py --task '%{task}' ``` The launcher owns its ranks. Only its main rank calls `finish()`. Do not start a Labtasker Worker inside every rank. Multi-node allocation and rendezvous remain the external scheduler's responsibility. Labtasker's documented launcher integration stops at the single-node pattern above; it does not define the ownership topology for one Task spanning several machines. Do not present a custom multi-node outer-Worker arrangement as a supported Labtasker interface.
-
-
SKILL.md 7.9 KB
--- name: labtasker description: Use Labtasker v2 to queue and run independent ML inference, evaluation, or experiment Tasks; migrate pipelines; design routes and Workers; inspect Task demand and Worker activity; and recover Tasks. Do not use it as a GPU allocator, cluster scheduler, workflow DAG, or artifact store. --- # Labtasker Use Labtasker when many independent ML jobs should be distributed across processes the user already controls, with progress, retries, and small structured results kept in one place. Keep GPU allocation, process launching, cluster management, dependent workflows, and artifact storage outside Labtasker. Prefer the documented public path even when a custom workaround is technically possible. Do not add a Server, shell wrapper, Queue, or compatibility mechanism unless the workload needs it. Documentation map: <https://raw.githubusercontent.com/luocfprime/labtasker/refs/heads/main/docs/llms.txt> ## Read the relevant reference - Read [deployment-and-capabilities.md](references/deployment-and-capabilities.md) for installation, local versus shared operation, HTTP authentication, Windows, Unix-socket requests, package selection, version warnings, or “does it support this?” questions. - Read [workers-and-workloads.md](references/workers-and-workloads.md) when submitting Tasks, converting an experiment, choosing routes or Queues, binding Task args, wrapping a command, reusing a loaded model, reporting progress, or using a distributed launcher. - Read [operations-and-recovery.md](references/operations-and-recovery.md) for idempotent submission, priority, filtering, fuzzy name search, pagination, updates, cancellation, external early-stop decisions, retries, interruption, and rerunning work. - Read [observations-and-counts.md](references/observations-and-counts.md) for online Workers, route presence, busy/idle activity, Worker metadata/latest telemetry, grouped counts, and paginated monitoring queries. Read every reference relevant to the request before proposing commands. Before the first Labtasker CLI operation in an environment, run the selected executable's `--version`. Before the first use of each distinct command path, run its exact `--help` (or the narrowest command-group help that exposes the needed subcommand). Use the same launcher and environment for these checks as for the real operation (for example, `uv run labtasker`). Treat the installed help as authoritative when it differs from this Skill; do not guess an option, subcommand, default, or output shape from memory. ## Adapt an existing pipeline on the user's terms When the user asks to migrate, convert, or adapt an existing pipeline, first set an internal working preference from the conversation: either lead from their existing project or review a concrete migration design they already proposed. This preference controls the agent's behavior; never name it, present it as a mode, or ask the user to select it. - If they are new to Labtasker, work project-first. Ask about their existing command or function, what varies between runs, expensive setup, independent failure and retry units, current resource launching, dependencies, and output storage. Do not ask them to choose a Task, Worker, route, Queue, or Labtasker deployment. Make those mappings yourself and explain them after the relevant project facts are known. - If they already propose a concrete Labtasker design, collaborate at that level, correct mistaken mappings, and still recommend a complete design rather than returning the decisions to them. - If neither is clear, ask naturally whether this is their first Labtasker integration or whether they already have a concrete migration design to work from. Never offer “use Labtasker concepts” as a conversation mode. Do not ask a classification question when the context already answers it. Ask only one or two decision-changing project questions at a time, inspect the current pipeline when available, then present the existing flow, what stays unchanged, what Labtasker coordinates, and what remains externally owned. Read [workers-and-workloads.md](references/workers-and-workloads.md) for the detailed migration interview and mapping rules. ## Use the managed-local path first Labtasker requires Python 3.10 or newer. In an ordinary POSIX experiment project, install the complete package: ```bash python -m pip install labtasker # or in a uv project uv add labtasker ``` No MongoDB, configuration file, or TCP port is needed. A default Client selects exact `CWD/.labtasker` and connects to its derived Unix socket, but it does not start a process. Explicitly authorize startup on the first operation: ```bash labtasker --auto-start-local-server queue list ``` Submit one Task: ```bash labtasker task submit \ --name sample-1 \ --args '{"prediction":"cat","reference":"cat"}' \ --route text-eval ``` The flag is invocation-scoped and idempotent. Later commands connect without it while the daemon remains healthy. Use global `--labtasker-root PATH` or `LABTASKER_ROOT` to select another exact root; never search parent or VCS directories. Run an existing program once for every compatible Task: ```bash CUDA_VISIBLE_DEVICES=0 labtasker loop --route text-eval -- \ python evaluate.py \ --prediction '%{prediction}' \ --reference '%{reference}' ``` Start another Worker process on each additional resource already allocated by the user. Each Worker executes one Task at a time and asks for another when it finishes. Labtasker does not select the GPU. Inspect the recorded state and result: ```bash labtasker task list --status succeeded labtasker task get t_ABCDEFGHIJKL ``` With a uv project, run these commands through `uv run`. Use `labtasker config show` to inspect the selected endpoint without starting or contacting a Server. ## Keep the working model small - A **Task** is one independent job plus its JSON inputs, state, retry count, metadata, and small result. - A **Worker** is one user-started process that repeatedly executes compatible Tasks. The Server stores authoritative Tasks and supplementary expiring Worker observations with optional user-defined resource details; it does not discover, allocate, or manage processes or GPU capacity. - A **route** is an exact, case-sensitive compatibility label shared by a Task and the implementation allowed to run it. - A **Queue** is an independently managed body of Tasks, not a Worker, GPU, model, or route. Put executable inputs in `args`, searchable grouping data in `metadata`, and compact JSON outputs in `result`. Save images, videos, checkpoints, trajectories, and detailed reports outside Labtasker and return their paths, URLs, checksums, or summaries. Use `progress` for one replace-only snapshot of provisional work position, metrics, or early-stop diagnostics. Keep `result` for the final successful output. Progress is not a history series and does not decide or complete the Task. When a determinate display is useful, use top-level finite numeric `completed` and `total` values with `0 <= completed <= total` and `total > 0`; all other keys remain workload-defined. Use a command Worker for an existing executable. Use a Python Worker when a model, dataset, simulator, or evaluator should be initialized once and reused. ## Preserve explicit behavior - Do not infer Worker eligibility from Task args; use routes. - Do not invent v1 aliases or implicit coercion. CLI objects are strict JSON. - Inspect before mutating. Use `cancel`, `requeue`, and `delete` rather than patching status. - Do not silently start, stop, or reconfigure an externally managed Server. Confirm its ownership and deployment scope first. - Unix sockets and HTTP are both public Server transports. Require an explicit `labtasker-server serve --connection socket|http` choice; do not infer one. Prefer direct argv; add a wrapper only when the workload itself needs shell or multi-step logic. - Treat the Server as authoritative. Local run journals are diagnostic records, not a second source of Task state.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.