{"slug":"datarobot-workload-api","title":"datarobot-workload-api","summary":"Use when the user wants to create, configure, scale, debug, observe, or roll out container workloads on DataRobot's Workload API. Triggers include: deploying a container as a managed service, listing/starting/stopping workloads, changing replica counts or autoscaling, picking CPU","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-11T17:36:45.967473Z","repo":{"url":"https://github.com/datarobot-oss/datarobot-agent-skills","stars":27,"forks":23,"license":"Apache-2.0","updatedAt":"2026-09-24T02:43:43Z"},"bodyHtml":"<hr>\n<h2>name: datarobot-workload-api\ndescription: &gt;-\nUse when the user wants to create, configure, scale, debug, observe, or roll\nout container workloads on DataRobot's Workload API. Triggers include:\ndeploying a container as a managed service, listing/starting/stopping\nworkloads, changing replica counts or autoscaling, picking CPU/GPU compute\nbundles, injecting DataRobot credentials as env vars, diagnosing workloads\nthat are stuck / errored / crash-looping (CrashLoopBackOff, ImagePullBackOff,\nOOMKilled, probe failures, exec format error), pulling application logs /\nOpenTelemetry traces / metrics / request stats, creating or iterating\ncontainer artifacts, building images server-side, locking artifacts for\nproduction, or doing a zero-downtime rolling artifact replacement.</h2>\n<h1>DataRobot Workload API</h1>\n<p>Run container images as managed, autoscalable services on DataRobot. One skill, four jobs — pick the section by user intent:</p>\n<ol>\n<li><strong>Create / configure / scale</strong> — deploy a container; change replicas, resources, autoscaling, bundle; inject credentials</li>\n<li><strong>Diagnose</strong> — workload is stuck, errored, or crash-looping</li>\n<li><strong>Observe</strong> — logs, traces, metrics, service stats for a running workload</li>\n<li><strong>Artifact lifecycle</strong> — iterate drafts, build images, lock for production, roll out new versions</li>\n</ol>\n<h2>Prerequisites</h2>\n<p><code>DATAROBOT_ENDPOINT</code> (must end in <code>/api/v2</code>) and <code>DATAROBOT_API_TOKEN</code> must be set. Run <code>datarobot-setup</code> if not. Auth header: <code>Authorization: Bearer ${DATAROBOT_API_TOKEN}</code>. The Workload API is not in the <code>datarobot</code> Python SDK — call REST directly.</p>\n<p><strong>Transport.</strong> Examples use Python <code>httpx</code> (<code>pip install httpx</code>). The API is plain HTTP, so equivalent calls work via <code>curl</code> or the <code>pulumi-datarobot</code> Pulumi provider declaratively. The skill teaches the model; transport is interchangeable.</p>\n<h2>Bundled scripts</h2>\n<p>Runnable Python in <code>scripts/</code> (this skill's folder). Each uses <code>httpx</code> and reads <code>DATAROBOT_ENDPOINT</code> + <code>DATAROBOT_API_TOKEN</code>:</p>\n<ul>\n<li><code>wait_for_running.py &lt;workload_id&gt;</code> — poll until <code>running</code>; exit 2 on terminal failure, 3 on timeout</li>\n<li><code>diagnose_workload.py &lt;workload_id&gt;</code> — run the 5-step debug flow, print a structured diagnosis (<code>--json</code> for machine-readable)</li>\n<li><code>wait_for_build.py &lt;artifact_id&gt; &lt;build_id&gt;</code> — poll a server-side image build; dumps last 2KB of logs on <code>FAILED</code></li>\n<li><code>wait_for_replacement.py &lt;workload_id&gt;</code> — poll a rolling replacement; handles the 404-when-cleared case</li>\n<li><code>check_limits.py</code> — print the user's effective org-set scaling limits via <code>/account/info/</code></li>\n</ul>\n<h2>Deeper docs in references/</h2>\n<p>SKILL.md is the operational core; occasional detail lives in <code>references/</code>:</p>\n<ul>\n<li><code>status-vocabulary.md</code> — workload + proton status enums and transitions</li>\n<li><code>common-error-patterns.md</code> — CrashLoopBackOff / ImagePullBackOff / OOMKilled / probe / exec-format / pending</li>\n<li><code>schema-reference.md</code> — schemas to look up, credential-type→key maps, public-spec path quirks</li>\n<li><code>lifecycle-flows.md</code> — artifact draft→lock→prod rules, replacement preconditions, redeploy matrix, <code>imageUri</code> gotchas</li>\n<li><code>code-to-workload.md</code> — deploy from source: <code>dr</code> CLI, <code>codeRef</code>, Execution Environments, iterate-rebuild loop</li>\n<li><code>web-uis-behind-the-edge.md</code> — browser-facing web app through the endpoint: prefix stripping, auth gate, <code>Authorization</code> hijack, shim, CSRF, WebSockets</li>\n</ul>\n<h2>OpenAPI spec is source of truth</h2>\n<p>At <code>${DATAROBOT_ENDPOINT}/openapi.yaml</code>. <strong>~5 MB — never dump it whole.</strong> Save once, then slice with <code>yq</code> (or <code>print()</code> only the specific key in Python):</p>\n<pre><code>curl -sS \"${DATAROBOT_ENDPOINT}/openapi.yaml\" -o /tmp/wapi-spec.yaml\nyq '.components.schemas.CreateWorkloadRequest' /tmp/wapi-spec.yaml\nyq '.components.schemas | keys | .[]' /tmp/wapi-spec.yaml | grep -i workload   # discover\n</code></pre>\n<p>All workload paths are keyed with the <code>/api/v2/</code> prefix — see <code>references/schema-reference.md</code>.</p>\n<hr>\n<h1>1. Create / configure / scale</h1>\n<h2>Run a container as a workload (the 90% case)</h2>\n<pre><code># spec.yaml — JSON also accepted; spec is sent verbatim\nname: my-api-service\nimportance: low\nartifact:\n  name: my-api-service-artifact\n  spec:\n    type: service\n    containerGroups:\n      - name: default\n        containers:\n          - name: main\n            imageUri: ghcr.io/org/my-app:latest\n            port: 8000\n            primary: true\n            readinessProbe: {path: /readyz, port: 8000, initialDelaySeconds: 10}\n            livenessProbe: {path: /healthz, port: 8000, initialDelaySeconds: 30}\nruntime:\n  containerGroups:\n    - name: default          # must match artifact.spec.containerGroups[].name (above)\n      replicaCount: 1\n      containers:\n        - name: main\n          resourceAllocation: {cpu: 1, memory: \"512MB\"}\n</code></pre>\n<pre><code>dr workload create --spec-file spec.yaml         # v0.2.74+; 4xx: 400=schema/limit, 403=cap (run check_limits.py), 409=name conflict\ndr workload get &lt;workload_id&gt;                    # or `dr workload status` — poll until status=running\n</code></pre>\n<p>Lifecycle one-liners (v0.2.74+): <code>dr workload {stop|start|delete|endpoint|list} &lt;id&gt;</code>.</p>\n<p>Raw fallback when CLI unavailable: <code>httpx.post(f\"{base}/workloads/\", headers=headers, json=spec)</code> + <code>r.raise_for_status()</code> + <code>r.json()[\"id\"]</code>. Then <code>python scripts/wait_for_running.py &lt;workload_id&gt;</code>.</p>\n<p><strong>Critical gotchas:</strong></p>\n<ul>\n<li><code>importance</code>: <code>low</code>/<code>moderate</code>/<code>high</code>/<code>critical</code>; <code>type</code>: <code>service</code> (default) or <code>nim</code>. Exactly one container per group has <code>primary: true</code>.</li>\n<li><code>cpu</code> is cores (float OK). <code>memory</code> accepts decimal string (<code>\"512MB\"</code>, units B/KB/MB/GB) or byte integer; Kubernetes binary suffixes (<code>Mi</code>/<code>Gi</code>) NOT supported.</li>\n<li><code>port</code> MUST be <code>&gt;= 1024</code>. The container must actually listen on it (set via image env vars or entrypoint).</li>\n<li>Image must include a <strong>linux/amd64</strong> manifest. Apple Silicon defaults to ARM64 and crash-loops with <code>exec format error</code>. Build with <code>docker buildx build --platform linux/amd64,linux/arm64 -t &lt;ref&gt; --push .</code>.</li>\n<li>Status lifecycle: <code>submitted</code> → <code>provisioning</code> → <code>launching</code> → <code>running</code> (happy path); <code>updating</code> during rolling redeploys; <code>errored</code> recoverable; <code>failed</code>/<code>terminated</code> unrecoverable. Full table in <code>references/status-vocabulary.md</code>.</li>\n</ul>\n<h2>Serving a browser-facing web UI through the endpoint</h2>\n<p>If the container serves a <strong>web app (UI + its own backend/API/WebSocket)</strong> opened in a browser via <code>dr workload endpoint &lt;id&gt;</code> (not a headless service), the DataRobot edge gateway serves it under a path prefix and: <strong>strips the prefix inbound</strong> (no outbound rewrite — the app must be sub-path aware); <strong>is the auth gate</strong> (DataRobot login required) and <strong>hijacks the <code>Authorization</code> header</strong> (→ <code>401 {\"message\":\"Invalid API key\"}</code>, never reaching the container); <strong>passes WebSockets through</strong>. Winning pattern: set the app's base-path to the prefix + re-add it inbound (derive it from the injected <code>WORKLOAD_ID</code>), <strong>disable the app's own auth (trust the edge)</strong>, disable CSRF, probe an unauthenticated path. Full guidance, shim code, and per-symptom diagnostics: <code>references/web-uis-behind-the-edge.md</code>.</p>\n<h2>\"Update the workload\" disambiguation</h2>\n<table>\n<thead>\n<tr>\n<th>User intent</th>\n<th>Endpoint</th>\n<th>Effect</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Rename / redescribe / change importance</td>\n<td><code>PATCH /workloads/{id}/</code></td>\n<td>Metadata only — no restart</td>\n</tr>\n<tr>\n<td>Change replicas / resources / autoscaling on the same artifact</td>\n<td><code>PATCH /workloads/{id}/settings/</code></td>\n<td>Triggers rolling redeploy</td>\n</tr>\n<tr>\n<td>Deploy a different artifact (new image / version)</td>\n<td><code>POST /workloads/{id}/replacement/</code></td>\n<td>Rolling swap — see section 4</td>\n</tr>\n</tbody>\n</table>\n<h2>Replicas, resources, autoscaling</h2>\n<p><code>PATCH /workloads/{wid}/settings/</code> with full body shape — use exactly one of <code>replicaCount</code> or <code>autoscaling</code>. Read settings first via <code>GET /workloads/{wid}/settings/</code>, then PATCH back:</p>\n<pre><code>httpx.patch(\n    f\"{base}/workloads/{wid}/settings/\",\n    headers=headers,\n    json={\n        \"runtime\": {\n            \"containerGroups\": [\n                {\n                    \"name\": \"default\",\n                    \"replicaCount\": 3,\n                    \"containers\": [\n                        {\n                            \"name\": \"main\",\n                            \"resourceAllocation\": {\"cpu\": 2, \"memory\": \"1GB\"},\n                        }\n                    ],\n                    # OR: \"autoscaling\": {\"enabled\": True, \"policies\": [{\n                    #       \"scalingMetric\": \"cpuAverageUtilization\",\n                    #       \"target\": 70, \"minCount\": 1, \"maxCount\": 10}]}\n                }\n            ]\n        }\n    },\n)\n</code></pre>\n<p>Valid <code>scalingMetric</code> values: <code>cpuAverageUtilization</code>, <code>httpRequestsConcurrency</code>, <code>gpuCacheUtilization</code>, <code>gpuRequestQueueDepth</code>, or a custom NIM metric. Settings updates are <strong>rolling</strong>; zero-downtime only with <code>replicaCount &gt;= 2</code> (or autoscaling <code>minCount &gt;= 2</code>).</p>\n<h2>Org-set scaling limits — check before scaling</h2>\n<p>Two admin-set caps: <code>maxConcurrentWorkloads</code> and <code>maxWorkloadReplicas</code>. Value <code>0</code> = unlimited; users can't change them. Read via <strong><code>GET /account/info/</code></strong> — response includes <code>{\"limits\": {\"maxConcurrentWorkloads\": N, \"maxWorkloadReplicas\": M}}</code> (or <code>python scripts/check_limits.py</code>). The spec's <code>/users/{uid}/</code> and <code>/organizations/{id}/</code> paths require Admin API access. Exceeding either limit returns <strong>HTTP 403</strong> with <code>{\"detail\": \"Requested replicas (N) exceeds the maximum allowed (M).\"}</code> — check limits first, then propose the max allowed or flag that admin help is needed.</p>\n<h2>GPU type / VRAM — set via compute bundle, not direct</h2>\n<p><code>resourceAllocation</code> only accepts <code>cpu</code>, <code>memory</code>, <code>gpu</code> (count). There is NO <code>gpuType</code> or <code>gpuMemory</code> field. To target a GPU model / VRAM size: <code>GET /mlops/compute/bundles/</code> lists bundles (<code>cpu.small</code>, <code>gpu.l4.small</code>, <code>gpu.a10g.medium</code>); pass via <code>\"resourceBundles\": [\"gpu.l4.small\"]</code> (a list, but exactly ONE bundle allowed) under the container group. When a bundle is set, CPU/memory in <code>resourceAllocation</code> are ignored — the bundle defines them.</p>\n<h2>Credential injection — never hardcode secrets</h2>\n<p>DataRobot credentials are stored centrally and injected into <code>environmentVars</code> by reference:</p>\n<pre><code>\"environmentVars\": [\n    {\"name\": \"PLAIN_VAR\", \"value\": \"literal-value\"},\n    {\"source\": \"dr-credential\", \"name\": \"AWS_ACCESS_KEY_ID\",\n     \"drCredentialId\": \"&lt;credential-id&gt;\", \"key\": \"awsAccessKeyId\"},\n]\n</code></pre>\n<p>Workflow: <code>GET /credentials/?limit=50</code> → note the credential's <code>credentialType</code> → look up the valid <code>key</code> field names for that type in <code>references/schema-reference.md</code> (covers <code>s3</code>, <code>basic</code>, <code>api_token</code>, <code>bearer</code>, <code>oauth</code>, <code>gcp</code>, <code>azure_*</code>, <code>databricks_*</code>, <code>snowflake_*</code>, …).</p>\n<h2>Create from an existing artifact</h2>\n<p>Provide <code>artifactId</code> instead of the inline <code>artifact</code> block. The <code>containerGroups[].name</code> and <code>containers[].name</code> in <code>runtime</code> must match what the artifact defines.</p>\n<hr>\n<h1>2. Diagnose — workload is stuck, errored, or crash-looping</h1>\n<h2>One command for the full diagnosis</h2>\n<pre><code>python scripts/diagnose_workload.py &lt;workload_id&gt;\n</code></pre>\n<p>Runs all 5 steps below, prints a structured report (status / logTail signals / flagged events / proton K8s detail / evidence / recommended next step / console URL). <code>--json</code> for machine-readable. If <code>Evidence</code> is empty, pull application logs via section 3 — don't guess from status alone.</p>\n<h2>The 5-step flow</h2>\n<p>The script encapsulates this; use the model below for ambiguous output or one-off calls.</p>\n<ol>\n<li><strong><code>GET /workloads/{id}/</code></strong> — <code>status</code>, <code>statusDetails.logTail</code> (~30 lines; scan for <code>error</code>/<code>exception</code>/<code>traceback</code>/<code>killed</code>/<code>permission denied</code>/<code>connection refused</code>), <code>statusDetails.conditions</code>. Guard <code>statusDetails</code> — it's <code>null</code> during <code>submitted</code>/<code>provisioning</code>.</li>\n<li><strong><code>GET /workloads/{id}/events/</code></strong> — flag <code>type: Warning</code> or <code>reason</code> with <code>Failed</code>/<code>Error</code>/<code>Kill</code>/<code>OOM</code>; the last Warning before <code>errored</code> is usually the trigger.</li>\n<li><strong><code>GET /workloads/{id}/protons/</code></strong> — pick <code>role: \"active\"</code> (or the <code>candidate</code> during a rolling replacement; else newest <code>createdAt</code>).</li>\n<li><strong><code>GET /workloads/{id}/protons/{pid}/statusDetails/</code></strong> — <code>204</code> while initializing (not an error). Read <code>replicas[*].containers[*].status</code>+<code>restartCount</code> → <code>replicas[*].conditions[*]</code> (any <code>value:false</code>) → <code>overallStatus.summary</code>.</li>\n<li><strong>Application logs</strong> — section 3.</li>\n</ol>\n<p>Common patterns (<code>CrashLoopBackOff</code>, <code>ImagePullBackOff</code>, <code>OOMKilled</code>, probe/pending, <code>exec format error</code>) and fixes: <code>references/common-error-patterns.md</code>.</p>\n<h2>Reporting findings</h2>\n<pre><code>Workload {id} — Diagnosis\n- Status: {current}\n- Root cause: {one sentence}\n- Evidence: {the specific logTail line, condition, container reason, or event}\n- Recommended fix: {actionable next step — section 1 (settings), section 4 (artifact), or app code}\n- Console: https://app.datarobot.com/console-nextgen/workloads/{id}/overview\n</code></pre>\n<hr>\n<h1>3. Observe — logs, traces, metrics, service stats</h1>\n<table>\n<thead>\n<tr>\n<th>Stream</th>\n<th>Endpoint</th>\n<th>Needs app instrumentation?</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Logs</td>\n<td><code>/otel/workload/{id}/logs/</code></td>\n<td>No — auto from stdout/stderr</td>\n</tr>\n<tr>\n<td>Traces</td>\n<td><code>/otel/workload/{id}/traces/</code></td>\n<td><strong>Yes</strong> (OTEL spans)</td>\n</tr>\n<tr>\n<td>Metrics</td>\n<td><code>/otel/workload/{id}/metrics/autocollectedValues/</code></td>\n<td>Partially</td>\n</tr>\n<tr>\n<td>Service stats</td>\n<td><code>/workloads/{id}/stats/</code></td>\n<td>No — DataRobot edge proxy</td>\n</tr>\n<tr>\n<td>Replacement history</td>\n<td><code>/workloads/{id}/history/</code></td>\n<td>No — platform</td>\n</tr>\n<tr>\n<td>Lifecycle events</td>\n<td><code>/workloads/{id}/events/</code></td>\n<td>No — platform</td>\n</tr>\n</tbody>\n</table>\n<p>Always check <code>r.status_code</code> before <code>.json()</code>: 401 = bad token; 404 = workload not found; 429 = rate limited (exponential backoff). All list endpoints accept <code>limit</code> + <code>offset</code>.</p>\n<h2>Logs</h2>\n<pre><code>dr workload logs &lt;wid&gt; --level error --limit 100   # v0.2.74+; --follow streams; --output-format json\n</code></pre>\n<p><code>--level</code> is an EXACT severity match (not a threshold). For substring filtering on the message body, or proton-scoped logs (find proton IDs in section 2), drop to REST — <code>dr workload logs</code> doesn't expose those filters:</p>\n<pre><code>r = httpx.get(\n    f\"{base}/otel/workload/{wid}/logs/\",\n    headers=headers,\n    params=[\n        (\"searchKeys\", \"proton_id\"),\n        (\"searchValues\", pid),\n        (\"searchKeys\", \"level\"),\n        (\"searchValues\", \"error\"),\n    ],\n)\n</code></pre>\n<p><code>searchKeys</code> / <code>searchValues</code> are positional parallel lists — pass a <strong>list of tuples</strong> to httpx (dict can't repeat keys). <code>includes=&lt;substring&gt;</code> does case-sensitive substring filtering on the message body.</p>\n<h2>Traces</h2>\n<pre><code>traces = httpx.get(f\"{base}/otel/workload/{wid}/traces/\", headers=headers).json()[\n    \"data\"\n]\n# summary: traceId, rootSpanName, rootServiceName, duration (NANOSECONDS), spansCount, errorSpansCount\ntrace_id = next(\n    (t[\"traceId\"] for t in traces if t.get(\"errorSpansCount\", 0) &gt; 0),\n    traces[0][\"traceId\"],\n)\ntrace = httpx.get(\n    f\"{base}/otel/workload/{wid}/traces/{trace_id}/\", headers=headers\n).json()\n</code></pre>\n<blockquote>\n<p><strong><code>duration</code> is NANOSECONDS</strong> on summaries AND spans. Divide by 1,000,000 for ms before display. Empty <code>data</code> = app isn't instrumented; direct the user to wire up OTEL.</p>\n</blockquote>\n<h2>Metrics + service stats</h2>\n<p>Convert before display: <code>bytes</code>→MB (<code>/1024**2</code>), <code>nanocores</code>→cores (<code>/1_000_000</code>), <code>percentage</code> already %.</p>\n<pre><code>stats = httpx.get(f\"{base}/workloads/{wid}/stats/\", headers=headers).json()\n# {\"period\": {...}, \"metrics\": {totalRequests, serverErrors, userErrors, slowRequests,\n#   responseTime, requestsPerMinute, concurrentRequests, *ErrorRate}}. /workloads/stats/ = aggregate.\n</code></pre>\n<blockquote>\n<p><strong>Destructive:</strong> <code>DELETE /workloads/{id}/stats/?metricName=&lt;name&gt;</code> zeroes a metric's history — only on explicit request.</p>\n</blockquote>\n<h2>Presenting results</h2>\n<p>Logs: <code>timestamp | level | message</code>, ERROR/CRITICAL first. Traces: table sorted by errors desc then recency. Metrics: apply unit conversion before display. Service stats one-liner: <em>\"<code>{totalRequests}</code> requests, <code>{totalErrorRate*100:.2f}%</code> errors, <code>{responseTime:.1f}</code> ms avg, <code>{requestsPerMinute}</code> req/min.\"</em> Empty data → say <em>why</em> (not running, not instrumented, empty window), don't just \"no data\".</p>\n<hr>\n<h1>4. Artifact lifecycle</h1>\n<p>An <strong>artifact</strong> is the immutable-after-lock definition of what a workload runs (image, port, env vars, probes). A <strong>workload</strong> is the running instance + its runtime (replicas, resources, autoscaling). Resources do NOT belong on the artifact.</p>\n<h2>Picking the right path</h2>\n<p>Find the running artifact (<code>workload[\"artifactId\"]</code>), check <code>artifact[\"status\"]</code>. A running workload does <strong>not</strong> auto-adopt a rebuild until you redeploy.</p>\n<ul>\n<li><strong>Same draft (the C2W loop) — in-place change or rebuild.</strong> PATCH/rebuild the draft, then roll onto it with <code>PATCH /workloads/{id}/settings/</code>: re-send the runtime body (even unchanged values trigger a rolling <code>202</code> redeploy that re-reads the current spec + latest <code>COMPLETED</code> build). Zero-downtime at ≥2 replicas. (<code>POST /replacement/</code> onto the same draft also works.)</li>\n<li><strong>Different / locked artifact.</strong> <code>POST /replacement/</code> onto the other artifact ID. Locked in-place edit: clone → PATCH clone → lock → replace onto the clone.</li>\n</ul>\n<p><strong>Lock:</strong> <code>dr artifact lock &lt;id&gt;</code> (= <code>PATCH /artifacts/{id}/ {\"status\":\"locked\"}</code>). <strong>Promote</strong> (<code>POST /workloads/{wid}/promote/</code>, 200) locks the running draft in place, no restart. Runtime-only changes (replicas/resources/autoscaling) → <code>PATCH /settings/</code>; a PATCH to the artifact doesn't affect live workloads until you redeploy.</p>\n<p>Preconditions (status-match, same-artifact rule) and the full redeploy matrix: <code>references/lifecycle-flows.md</code>.</p>\n<h2>How does your image get to DataRobot?</h2>\n<p>The artifact's <code>imageUri</code> must point at a registry DataRobot can pull from (image-pull creds aren't accepted at workload creation yet). Two paths:</p>\n<ol>\n<li><strong>Bring your own image</strong> — public registry or one the admin pre-configured. <code>docker buildx ... --platform linux/amd64</code>, push, set <code>imageUri</code>. Default flow.</li>\n<li><strong>Code-to-Workload (C2W)</strong> — no local Docker / no public registry: <code>dr artifact code init</code> + <code>sync</code>, then <code>dr artifact build create</code> builds server-side, pushes to DataRobot's internal registry, and populates <code>imageUri</code>. Full flow in <code>references/code-to-workload.md</code>.</li>\n</ol>\n<p>Poll builds with <code>python scripts/wait_for_build.py &lt;artifact_id&gt; &lt;build_id&gt;</code>; only drafts build. <strong><code>imageUri</code> is build-managed</strong> — never PATCH it by hand (<code>422</code> \"not permitted on this cluster\"), and never PATCH the spec <em>mid-build</em> (a whole-spec write clobbers the pending build image → redeploys the old one). Sequence spec edits before <code>build create</code> or after <code>COMPLETED</code>.</p>\n<blockquote>\n<p><strong>C2W is preview / feature-flagged</strong> — <code>ENABLE_WORKLOAD_API_CONTAINERS=true</code> (org) + <code>DATAROBOT_CLI_FEATURE_WORKLOAD=true</code> (client).</p>\n</blockquote>\n<h2>Rolling artifact replacement</h2>\n<pre><code>httpx.post(\n    f\"{base}/workloads/{wid}/replacement/\",\n    headers=headers,\n    json={\n        \"artifactId\": new_artifact_id,\n        \"strategy\": \"rolling\",  # only \"rolling\" supported\n        \"config\": {\"warmupDurationMinutes\": 2, \"keepOldVersionMinutes\": 5},  # optional\n        # \"runtime\": {...}  # optional; same shape as PATCH /settings/\n    },\n)\n</code></pre>\n<p>Monitor with <code>python scripts/wait_for_replacement.py &lt;workload_id&gt;</code>. Preconditions: status must match (draft↔draft / locked↔locked, else <code>400</code>); same-artifact replacement 422s for locked but works for drafts — to roll the same draft without replacement use <code>PATCH /settings/</code>. <strong>Not idempotent</strong> (a second <code>POST</code> queues another swap); <code>GET .../replacement/</code> <code>404</code> = none in progress; <code>DELETE</code> to cancel. Detail in <code>references/lifecycle-flows.md</code>.</p>\n<hr>\n<h2>Related skills</h2>\n<ul>\n<li><code>datarobot-setup</code> — install SDK, configure auth, set env vars</li>\n<li><code>datarobot-app-framework-cicd</code> — declarative artifact + workload management via Pulumi and CI/CD</li>\n<li><code>datarobot-external-agent-monitoring</code> — instrument arbitrary agent code with OTEL → DataRobot</li>\n</ul>\n","files":[{"path":"references/code-to-workload.md","sizeBytes":14297,"isText":true},{"path":"references/common-error-patterns.md","sizeBytes":7572,"isText":true},{"path":"references/lifecycle-flows.md","sizeBytes":7026,"isText":true},{"path":"references/schema-reference.md","sizeBytes":5082,"isText":true},{"path":"references/status-vocabulary.md","sizeBytes":4408,"isText":true},{"path":"references/web-uis-behind-the-edge.md","sizeBytes":13493,"isText":true},{"path":"scripts/check_limits.py","sizeBytes":2817,"isText":true},{"path":"scripts/diagnose_workload.py","sizeBytes":8872,"isText":true},{"path":"scripts/wait_for_build.py","sizeBytes":4190,"isText":true},{"path":"scripts/wait_for_replacement.py","sizeBytes":4070,"isText":true},{"path":"scripts/wait_for_running.py","sizeBytes":3056,"isText":true},{"path":"SKILL.md","sizeBytes":19950,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"notes-only","suspicious":0,"notes":1,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-16T15:57:15.165357Z","sha256":"25D5B8B8BC3A3CC3C538482762B8DEB698C9A5C511C829A3256BF96B1C84D2C1","sizeBytes":41215},"review":null,"source":{"repositoryUrl":"https://github.com/datarobot-oss/datarobot-agent-skills","path":"skills/datarobot-workload-api","license":"Apache-2.0","commit":"023e5b77fb4c3f651f52d9afc943d87061e22769","subtreeSha":"BC3557E7E5929BF9F808E5D248B4AE09AF55123E2ACE66C8D45C40B250AD1D75","lastSyncedAt":"2026-09-24T06:49:20.424986Z"},"reviewedAt":"2026-09-16T16:05:34.227525Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/datarobot-oss/datarobot-agent-skills/tree/main/skills/datarobot-workload-api"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install datarobot-oss-datarobot-agent-skills@llmmart"},{"target":"git","command":"git clone https://github.com/datarobot-oss/datarobot-agent-skills.git"}]}