Claude Cursor Skill

datarobot-predictions

Tools and guidance for making predictions with DataRobot deployments, including real-time predictions, batch scoring, prediction dataset generation, and prediction explanations (SHAP/XEMP). Use when making predictions, running batch scoring, generating prediction datasets, or exp

LLM Mart · 0 points · 6 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download datarobot-oss-datarobot-agent-skills-skills_datarobot-predictions-e6dddbe.zip · 10 KB
Part of datarobot-oss/datarobot-agent-skills — 14 skills

Install

skills CLI npx skills add https://github.com/datarobot-oss/datarobot-agent-skills/tree/main/skills/datarobot-predictions
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install datarobot-oss-datarobot-agent-skills@llmmart
Git git clone https://github.com/datarobot-oss/datarobot-agent-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole datarobot-oss/datarobot-agent-skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

DataRobot Predictions Skill

This skill provides comprehensive guidance for working with DataRobot predictions, including real-time predictions, batch scoring, and generating prediction datasets.

Quick Start

Most common use case: Generate predictions for a deployment

  1. Get deployment features: get_deployment_features(deployment_id) to understand required columns
  2. Generate template: generate_prediction_data_template(deployment_id, n_rows) to create CSV structure
  3. Make predictions: Use deployment.predict_batch(...) (works for both single-row “real-time” and batch scoring)

Example: "Generate a prediction dataset template for deployment abc123 with 10 rows"

To also explain predictions: pass --max-explanations N to make_prediction.py (or the max_explanations=N kwarg in code). See Prediction Explanations below.

When to use this skill

Use this skill when you need to:

  • Make predictions from deployed DataRobot models
  • Explain individual predictions from a deployment (SHAP or XEMP, per-row)
  • Generate prediction dataset templates
  • Validate prediction data before scoring
  • Understand deployment feature requirements
  • Perform batch predictions on large datasets
  • Get sample training data to understand expected formats

For post-hoc explanations against a training project / leaderboard model (not a deployment), use the datarobot-model-explainability skill instead. This skill covers deployment-time explanations returned alongside scoring.

Key capabilities

1. Understanding Deployment Requirements

Before making predictions, you need to understand what features a deployment requires:

  • Feature names and types: Know which columns are needed (numeric, categorical, text, date)
  • Feature importance: Understand which features matter most
  • Target information: Know what you're predicting
  • Time series configuration: If applicable, understand datetime columns and series IDs

2. Generating Prediction Datasets

Create properly formatted prediction datasets:

  • Generate CSV templates with all required columns
  • Include sample values appropriate for each feature type
  • Add metadata comments explaining the model structure
  • Ensure correct column ordering

3. Validating Prediction Data

Validate datasets before making predictions:

  • Check for missing required features
  • Verify data types match expected types
  • Identify missing low-importance features (warnings)
  • Note extra columns that will be ignored

4. Making Predictions

Execute predictions using various methods:

  • Real-time predictions: Fast, synchronous predictions for individual records
  • Batch predictions: Process large datasets efficiently
  • Time series predictions: Handle forecasting scenarios with proper datetime handling

Workflow examples

Example 1: Generate prediction dataset for a new scenario

User request: "I want to predict sales for next week for store_A with temperatures of 75°F each day and no promotions."

Agent workflow:

  1. Get deployment features to understand required columns
  2. Generate a prediction data template with 7 rows (one week)
  3. Fill in the template with user's specific values:
    • Set temperature = 75 for all rows
    • Set promotion = 0 for all rows
    • Set store_id = "store_A" for all rows
    • Set dates for next 7 days
  4. Validate the data to ensure it's correct
  5. Make predictions using the validated dataset

Example 2: Batch scoring a CSV file

User request: "Score all records in my prediction_data.csv file using deployment abc123."

Agent workflow:

  1. Validate the CSV file structure matches deployment requirements
  2. Upload the file or provide file path
  3. Submit batch prediction job
  4. Monitor job status
  5. Retrieve and return prediction results

Using DataRobot SDK

This skill guides you to use the DataRobot Python SDK directly. Install the SDK if needed:

pip install datarobot datarobot-predict

Key SDK Operations

Use these DataRobot SDK methods to work with predictions:

Deployment Information:

  • dr.Deployment.get(deployment_id) - Get deployment details
  • deployment.get_features() - Get required features (name/type/importance)

Predictions:

  • deployment.predict_batch(source) - Convenience batch prediction API (CSV path, file object, or pandas DataFrame)
  • dr.BatchPredictionJob.score(deployment=deployment, ...) - Advanced batch prediction control
  • job.get_result_when_complete() - Wait for batch scoring to finish and download results

Data Management:

  • dr.Dataset.create_from_file(file_path) - Upload dataset
  • dr.Dataset.get(dataset_id) - Get dataset info

See the Common Patterns section below for complete examples.

Prediction Explanations

Deployments can return per-row explanations (top feature contributions) alongside predictions. Two algorithms are available depending on how the deployment was configured:

  • SHAP (shap): SHapley Additive exPlanations. Available on tree-based models when SHAP was enabled at deployment time. Returns signed contributions in the model's score space.
  • XEMP (xemp): DataRobot's eXplainable AI for the eXact Model Prediction. Default when SHAP is not enabled. Returns top-N strongest features with a qualitative strength (+++, --, etc.).

If you omit explanation_algorithm, the deployment's default is used.

How to request explanations

Pass max_explanations=N (and any optional filters) when calling datarobot_predict.deployment.predict:

import datarobot as dr
import pandas as pd
from datarobot_predict.deployment import predict as dr_predict

dr.Client()
deployment = dr.Deployment.get("abc123")

result = dr_predict(
    deployment=deployment,
    data_frame=pd.DataFrame([{"feature1": 10, "feature2": 20}]),
    max_explanations=3,  # top 3 contributors per row
    explanation_algorithm="shap",  # or "xemp"; omit for deployment default
    # threshold_high=0.8,                # optional: only explain rows scoring > 0.8
    # threshold_low=0.2,                 # optional: only explain rows scoring < 0.2
    # passthrough_columns="all",         # optional: echo input columns through to output
)
print(result.dataframe.to_dict(orient="records"))

The result DataFrame includes columns like EXPLANATION_1_FEATURE_NAME, EXPLANATION_1_ACTUAL_VALUE, EXPLANATION_1_STRENGTH, EXPLANATION_1_QUALITATIVE_STRENGTH for each of the top-N contributors.

Parameter reference

Parameter Purpose
max_explanations Top-N contributors per row. 0 (default) disables explanations.
max_ngram_explanations Text models only: cap text-segment explanations per row.
threshold_high Only explain rows with prediction probability above this (0–1).
threshold_low Only explain rows with prediction probability below this (0–1).
explanation_algorithm "shap" or "xemp"; omit to use deployment default.
passthrough_columns "all" or set of input column names to echo through to output.

CLI shortcut

python scripts/make_prediction.py abc123 '{"feature1": 10, "feature2": 20}' \
    --max-explanations 3 --explanation-algorithm shap

When to use which threshold

  • threshold_high is useful when only positive (high-risk / fraud / churn-likely) predictions need explaining — saves compute on a large batch.
  • threshold_low is the mirror image for low-probability rows.
  • Setting both restricts explanations to rows outside the [low, high] band.

Common errors

  • "Prediction explanations not enabled": the deployment was created without explanations support. Re-deploy the model with explanations enabled, or use a deployment that has them.
  • max_explanations ignored / no explanation columns in output: confirm you're calling datarobot_predict.deployment.predict(...) and that the deployment has explanations enabled. The deployment.predict_batch() convenience wrapper on the SDK is intended for plain scoring; use datarobot_predict.deployment.predict when you need explanation kwargs.

Helper Scripts

This skill includes executable helper scripts that Claude can run directly:

  • scripts/get_deployment_features.py - Get deployment feature requirements
  • scripts/generate_prediction_data_template.py - Generate CSV template
  • scripts/validate_prediction_data.py - Validate prediction data
  • scripts/make_prediction.py - Make real-time predictions

Usage example:

# Get deployment features
python scripts/get_deployment_features.py abc123

# Generate template
python scripts/generate_prediction_data_template.py abc123 10 template.csv

# Validate data
python scripts/validate_prediction_data.py abc123 prediction_data.csv

# Make prediction
python scripts/make_prediction.py abc123 '{"feature1": 10, "feature2": 20}'

# Make prediction with top-3 SHAP explanations
python scripts/make_prediction.py abc123 '{"feature1": 10, "feature2": 20}' \
    --max-explanations 3 --explanation-algorithm shap

Claude can run these scripts directly or use them as reference when writing code.

Best practices

  1. Always validate first: Validate prediction data before submitting predictions to catch errors early
  2. Use templates: Generate templates to ensure correct structure and avoid missing columns
  3. Check feature types: Ensure numeric features are numbers, categorical features match training values
  4. Handle time series: For time series models, ensure datetime columns and series IDs are properly formatted
  5. Monitor batch jobs: For large batch predictions, check job status and handle errors appropriately

Common patterns

Pattern 1: Get deployment features and make single prediction (optionally with explanations)

import datarobot as dr
import pandas as pd
from datarobot_predict.deployment import predict as dr_predict

# Initialize client
dr.Client()

deployment = dr.Deployment.get("abc123")

prediction_data = {
    "feature1": value1,
    "feature2": value2,
    # ... all required features (excluding target)
}

# Score one row. Add max_explanations=N to get top-N explanations per row.
result = dr_predict(
    deployment=deployment,
    data_frame=pd.DataFrame([prediction_data]),
    max_explanations=3,  # optional; 0/omit to disable explanations
    explanation_algorithm="shap",  # optional; omit to use deployment default
)
print(result.dataframe.to_dict(orient="records"))

Pattern 2: Generate template and batch predictions

import datarobot as dr
import pandas as pd

# Initialize client
dr.Client()

# Get deployment features
deployment = dr.Deployment.get("abc123")
model = dr.Model.get(deployment.model["id"])
features = model.get_features()

# Create template DataFrame
prediction_features = [f for f in features if f.name != model.target_name]
template_df = pd.DataFrame(columns=[f.name for f in prediction_features])

# Add sample rows
for i in range(100):
    row = {}
    for feature in prediction_features:
        if feature.feature_type == "Numeric":
            row[feature.name] = 0.0
        elif feature.feature_type == "Categorical":
            row[feature.name] = "sample_value"
        else:
            row[feature.name] = ""
    template_df = pd.concat([template_df, pd.DataFrame([row])], ignore_index=True)

# Save template
template_df.to_csv("prediction_template.csv", index=False)

# Fill template with actual data (modify CSV as needed)
# ...

# Submit batch prediction
job = dr.BatchPredictionJob.score(
    deployment_id=deployment.id,
    intake_settings={"type": "localFile", "file": "prediction_template.csv"},
    output_settings={"type": "localFile", "path": "predictions_output.csv"},
)

# Monitor job
job_status = dr.BatchPredictionJob.get(job.id)
print(f"Job status: {job_status.status}")

# Download results when complete
if job_status.status == "completed":
    results = dr.BatchPredictionJob.download(job.id)

Error handling

Common errors and solutions:

  • Missing required features: Use get_deployment_features to get complete list
  • Wrong data types: Check feature types and convert accordingly
  • Invalid categorical values: Use training data sample to see valid values
  • Time series format errors: Ensure datetime format matches training data

SDK Setup

Install DataRobot SDK

pip install datarobot datarobot-predict

Initialize Client

import datarobot as dr

# Initialize client with API credentials
dr.Client()

Environment Variables

Set these environment variables or pass them directly:

  • DATAROBOT_API_TOKEN - Your DataRobot API token
  • DATAROBOT_ENDPOINT - Your DataRobot endpoint (default: https://app.datarobot.com)

Resources

Files (datarobot-agent-skills)
  • scripts
    • generate_prediction_data_template.py 3 KB
      #!/usr/bin/env python3
      # Copyright (c) 2026 DataRobot, Inc. All rights reserved.
      # SPDX-License-Identifier: Apache-2.0
      
      """
      Generate a CSV template for prediction data.
      
      Usage:
          python generate_prediction_data_template.py <deployment_id> [n_rows] [output_file]
      
      Generates a CSV template with all required columns and sample values.
      """
      
      import csv
      import sys
      
      import datarobot as dr
      
      
      def generate_prediction_data_template(
          deployment_id: str, n_rows: int = 1, output_file: str | None = None
      ) -> str:
          """
          Generate a CSV template for prediction data.
      
          Args:
              deployment_id: The deployment ID
              n_rows: Number of template rows to generate (default: 1)
              output_file: Optional output file path (default: prints to stdout)
      
          Returns:
              CSV template content
          """
          # Initialize client
          dr.Client()
      
          # Get deployment features
          deployment = dr.Deployment.get(deployment_id)
          model = dr.Model.get(deployment.model["id"])
          features = model.get_features()
      
          # Filter out target feature
          prediction_features = [f for f in features if f.name != model.target_name]
      
          # Generate template rows with sample values
          import io
      
          output = io.StringIO()
          writer = csv.DictWriter(output, fieldnames=[f.name for f in prediction_features])
          writer.writeheader()
      
          # Generate sample rows based on feature types
          for i in range(n_rows):
              row = {}
              for feature in prediction_features:
                  if feature.feature_type == "Numeric":
                      row[feature.name] = 0.0
                  elif feature.feature_type == "Categorical":
                      row[feature.name] = "sample_category"
                  elif feature.feature_type == "Text":
                      row[feature.name] = "sample text"
                  elif feature.feature_type == "Date":
                      row[feature.name] = "2024-01-01"
                  else:
                      row[feature.name] = ""
              writer.writerow(row)
      
          csv_content = output.getvalue()
      
          # Add metadata comments
          metadata_comments = f"""# Prediction Data Template for Deployment: {deployment_id}
      # Model: {model.project_name}
      # Target: {model.target_name}
      # Generated: {n_rows} template rows
      # 
      # Instructions:
      # 1. Fill in the values for each feature
      # 2. Ensure data types match feature types
      # 3. Use validate_prediction_data.py to check before submitting
      #
      """
      
          full_content = metadata_comments + csv_content
      
          if output_file:
              with open(output_file, "w") as f:
                  f.write(full_content)
              return f"Template written to {output_file}"
          else:
              return full_content
      
      
      if __name__ == "__main__":
          if len(sys.argv) < 2:
              print(
                  "Usage: python generate_prediction_data_template.py <deployment_id> [n_rows] [output_file]",
                  file=sys.stderr,
              )
              sys.exit(1)
      
          deployment_id = sys.argv[1]
          n_rows = int(sys.argv[2]) if len(sys.argv) > 2 else 1
          output_file = sys.argv[3] if len(sys.argv) > 3 else None
      
          result = generate_prediction_data_template(deployment_id, n_rows, output_file)
          print(result)
      
    • get_deployment_features.py 2.8 KB
      #!/usr/bin/env python3
      # Copyright (c) 2026 DataRobot, Inc. All rights reserved.
      # SPDX-License-Identifier: Apache-2.0
      
      """
      Get comprehensive information about features required by a deployment.
      
      Usage:
          python get_deployment_features.py <deployment_id>
      
      Outputs JSON with feature information, types, importance, and time series config.
      """
      
      import json
      import sys
      
      import datarobot as dr
      
      
      def get_deployment_features(deployment_id: str) -> dict:
          """
          Get comprehensive information about features required by a deployment.
      
          Args:
              deployment_id: The deployment ID
      
          Returns:
              Dictionary with feature information, types, importance, and time series config
          """
          # Initialize client
          dr.Client()
      
          deployment = dr.Deployment.get(deployment_id)
          model = dr.Model.get(deployment.model["id"])
          project = dr.Project.get(model.project_id)
      
          # Get feature information
          features = model.get_features()
          feature_importance = model.get_feature_impact()
      
          # Build feature list
          feature_list = []
          for feature in features:
              importance = 0.0
              for fi in feature_importance:
                  if fi["featureName"] == feature.name:
                      importance = fi.get("impactNormalized", 0.0)
                      break
      
              feature_list.append(
                  {
                      "feature_name": feature.name,
                      "feature_type": feature.feature_type,
                      "importance": importance,
                      "is_target": feature.name == model.target_name,
                  }
              )
      
          # Get time series config if applicable
          time_series_config = None
          if project.use_time_series:
              try:
                  time_series_info = project.get_time_series_info()
                  time_series_config = {
                      "datetime_column": time_series_info.datetime_partition_column,
                      "forecast_window_start": time_series_info.forecast_window_start,
                      "forecast_window_end": time_series_info.forecast_window_end,
                      "series_id_columns": time_series_info.multiseries_id_columns or [],
                  }
              except dr.errors.ClientError as e:
                  print(
                      f"Note: time series info unavailable: {e}",
                      file=sys.stderr,
                  )
      
          return {
              "deployment_id": deployment_id,
              "model_type": model.target_type,
              "target": model.target_name,
              "target_type": model.target_type,
              "features": feature_list,
              "time_series_config": time_series_config,
          }
      
      
      if __name__ == "__main__":
          if len(sys.argv) < 2:
              print(
                  "Usage: python get_deployment_features.py <deployment_id>", file=sys.stderr
              )
              sys.exit(1)
      
          deployment_id = sys.argv[1]
          result = get_deployment_features(deployment_id)
          print(json.dumps(result, indent=2))
      
    • make_prediction.py 4.6 KB
      #!/usr/bin/env python3
      # Copyright (c) 2026 DataRobot, Inc. All rights reserved.
      # SPDX-License-Identifier: Apache-2.0
      
      """
      Make a prediction from a deployment, optionally with prediction explanations.
      
      Usage:
          python make_prediction.py <deployment_id> <data_json> [options]
      
      Where <data_json> is either:
          - a JSON object with feature values (single row), or
          - a JSON array of objects (multiple rows).
      
      Options:
          --max-explanations N        Number of top explanations per row (0 disables, default 0).
          --max-ngram-explanations N  Cap text-segment explanations per row (text models only).
          --threshold-high X          Only explain rows with prediction probability above X (0-1).
          --threshold-low X           Only explain rows with prediction probability below X (0-1).
          --explanation-algorithm A   'shap' or 'xemp' (omit to use deployment default).
          --passthrough-columns COLS  'all' or comma-separated input columns to copy through.
      
      Examples:
          # Plain prediction
          python make_prediction.py abc123 '{"feature1": 10, "feature2": 20}'
      
          # Top-3 SHAP explanations per row
          python make_prediction.py abc123 '{"feature1": 10}' --max-explanations 3 \\
              --explanation-algorithm shap
      
      When --max-explanations > 0, each row in the output includes an `explanations` list
      of {feature, value, strength, qualitative_strength} entries describing why the model
      produced that prediction.
      """
      
      import argparse
      import json
      import sys
      
      import datarobot as dr
      import pandas as pd
      from datarobot_predict.deployment import predict as dr_predict
      
      
      def make_prediction(
          deployment_id: str,
          data,
          *,
          max_explanations: int = 0,
          max_ngram_explanations: int | None = None,
          threshold_high: float | None = None,
          threshold_low: float | None = None,
          explanation_algorithm: str | None = None,
          passthrough_columns: str | None = None,
      ) -> dict:
          """Score `data` against `deployment_id`, optionally returning prediction explanations."""
          dr.Client()
      
          deployment = dr.Deployment.get(deployment_id)
      
          rows = data if isinstance(data, list) else [data]
          df = pd.DataFrame(rows)
      
          predict_kwargs = {"deployment": deployment, "data_frame": df}
          if max_explanations and max_explanations > 0:
              predict_kwargs["max_explanations"] = max_explanations
          if max_ngram_explanations is not None:
              predict_kwargs["max_ngram_explanations"] = max_ngram_explanations
          if threshold_high is not None:
              predict_kwargs["threshold_high"] = threshold_high
          if threshold_low is not None:
              predict_kwargs["threshold_low"] = threshold_low
          if explanation_algorithm is not None:
              predict_kwargs["explanation_algorithm"] = explanation_algorithm
          if passthrough_columns is not None:
              predict_kwargs["passthrough_columns"] = (
                  "all"
                  if passthrough_columns == "all"
                  else {c.strip() for c in passthrough_columns.split(",")}
              )
      
          result = dr_predict(**predict_kwargs)
          predictions_df = result.dataframe
      
          return {
              "deployment_id": deployment_id,
              "row_count": len(predictions_df),
              "predictions": predictions_df.to_dict(orient="records"),
          }
      
      
      def _parse_args() -> argparse.Namespace:
          parser = argparse.ArgumentParser(
              description=__doc__, formatter_class=argparse.RawTextHelpFormatter
          )
          parser.add_argument("deployment_id")
          parser.add_argument(
              "data_json", help="JSON object (single row) or JSON array (multiple rows)"
          )
          parser.add_argument("--max-explanations", type=int, default=0)
          parser.add_argument("--max-ngram-explanations", type=int, default=None)
          parser.add_argument("--threshold-high", type=float, default=None)
          parser.add_argument("--threshold-low", type=float, default=None)
          parser.add_argument(
              "--explanation-algorithm", choices=["shap", "xemp"], default=None
          )
          parser.add_argument("--passthrough-columns", default=None)
          return parser.parse_args()
      
      
      if __name__ == "__main__":
          args = _parse_args()
      
          try:
              data = json.loads(args.data_json)
          except json.JSONDecodeError as e:
              print(f"Error: Invalid JSON: {e}", file=sys.stderr)
              sys.exit(1)
      
          result = make_prediction(
              args.deployment_id,
              data,
              max_explanations=args.max_explanations,
              max_ngram_explanations=args.max_ngram_explanations,
              threshold_high=args.threshold_high,
              threshold_low=args.threshold_low,
              explanation_algorithm=args.explanation_algorithm,
              passthrough_columns=args.passthrough_columns,
          )
          print(json.dumps(result, indent=2, default=str))
      
    • validate_prediction_data.py 4 KB
      #!/usr/bin/env python3
      # Copyright (c) 2026 DataRobot, Inc. All rights reserved.
      # SPDX-License-Identifier: Apache-2.0
      
      """
      Validate prediction data for a deployment.
      
      Usage:
          python validate_prediction_data.py <deployment_id> <file_path>
      
      Returns validation report with errors, warnings, and info messages.
      """
      
      import csv
      import json
      import os
      import sys
      
      import datarobot as dr
      
      
      def validate_prediction_data(deployment_id: str, file_path: str) -> dict:
          """
          Validate prediction data for a deployment.
      
          Args:
              deployment_id: The deployment ID
              file_path: Path to CSV file to validate
      
          Returns:
              Validation report with errors, warnings, and info
          """
          # Initialize client
          dr.Client()
      
          deployment = dr.Deployment.get(deployment_id)
          model = dr.Model.get(deployment.model["id"])
          features = model.get_features()
      
          # Get required features (exclude target)
          required_features = {
              f.name: f.feature_type for f in features if f.name != model.target_name
          }
      
          # Read CSV data
          if not os.path.exists(file_path):
              return {"valid": False, "errors": [f"File not found: {file_path}"]}
      
          with open(file_path, "r") as f:
              reader = csv.DictReader(f)
              rows = list(reader)
      
          if not rows:
              return {"valid": False, "errors": ["CSV file is empty"]}
      
          # Validate
          errors = []
          warnings = []
          info = []
      
          # Check for missing required features
          csv_columns = set(rows[0].keys())
          missing_features = set(required_features.keys()) - csv_columns
      
          if missing_features:
              errors.append(
                  f"Missing required features: {', '.join(sorted(missing_features))}"
              )
      
          # Check for extra columns
          extra_columns = csv_columns - set(required_features.keys())
          if extra_columns:
              info.append(
                  f"Extra columns (will be ignored): {', '.join(sorted(extra_columns))}"
              )
      
          # Check data types (simplified - actual validation would be more thorough)
          for row_num, row in enumerate(rows, start=2):  # Start at 2 (header is row 1)
              for feature_name, feature_type in required_features.items():
                  if feature_name in row:
                      value = row[feature_name]
                      if value == "":
                          warnings.append(f"Row {row_num}, {feature_name}: Empty value")
                      elif feature_type == "Numeric" and value:
                          try:
                              float(value)
                          except ValueError:
                              errors.append(
                                  f"Row {row_num}, {feature_name}: Expected numeric, got '{value}'"
                              )
      
          # Get feature importance for warnings
          try:
              feature_importance = model.get_feature_impact()
              low_importance_features = [
                  fi["featureName"]
                  for fi in feature_importance
                  if fi.get("impactNormalized", 0) < 0.05
              ]
      
              missing_low_importance = missing_features & set(low_importance_features)
              if missing_low_importance:
                  warnings.append(
                      f"Missing low-importance features: {', '.join(sorted(missing_low_importance))}"
                  )
          except dr.errors.ClientError as e:
              print(
                  f"Note: feature impact unavailable, skipping importance warnings: {e}",
                  file=sys.stderr,
              )
      
          return {
              "valid": len(errors) == 0,
              "deployment_id": deployment_id,
              "row_count": len(rows),
              "errors": errors,
              "warnings": warnings,
              "info": info,
              "required_features": sorted(required_features.keys()),
              "provided_features": sorted(csv_columns),
          }
      
      
      if __name__ == "__main__":
          if len(sys.argv) < 3:
              print(
                  "Usage: python validate_prediction_data.py <deployment_id> <file_path>",
                  file=sys.stderr,
              )
              sys.exit(1)
      
          deployment_id = sys.argv[1]
          file_path = sys.argv[2]
      
          result = validate_prediction_data(deployment_id, file_path)
          print(json.dumps(result, indent=2))
      
  • SKILL.md 13.3 KB
    ---
    name: datarobot-predictions
    description: Tools and guidance for making predictions with DataRobot deployments, including real-time predictions, batch scoring, prediction dataset generation, and prediction explanations (SHAP/XEMP). Use when making predictions, running batch scoring, generating prediction datasets, or explaining individual predictions from a deployment.
    ---
    
    # DataRobot Predictions Skill
    
    This skill provides comprehensive guidance for working with DataRobot predictions, including real-time predictions, batch scoring, and generating prediction datasets.
    
    ## Quick Start
    
    **Most common use case**: Generate predictions for a deployment
    
    1. **Get deployment features**: `get_deployment_features(deployment_id)` to understand required columns
    2. **Generate template**: `generate_prediction_data_template(deployment_id, n_rows)` to create CSV structure
    3. **Make predictions**: Use `deployment.predict_batch(...)` (works for both single-row “real-time” and batch scoring)
    
    **Example**: "Generate a prediction dataset template for deployment abc123 with 10 rows"
    
    **To also explain predictions**: pass `--max-explanations N` to `make_prediction.py` (or the
    `max_explanations=N` kwarg in code). See [Prediction Explanations](#prediction-explanations) below.
    
    ## When to use this skill
    
    Use this skill when you need to:
    - Make predictions from deployed DataRobot models
    - Explain individual predictions from a deployment (SHAP or XEMP, per-row)
    - Generate prediction dataset templates
    - Validate prediction data before scoring
    - Understand deployment feature requirements
    - Perform batch predictions on large datasets
    - Get sample training data to understand expected formats
    
    > For post-hoc explanations against a **training project / leaderboard model** (not a deployment),
    > use the `datarobot-model-explainability` skill instead. This skill covers deployment-time
    > explanations returned alongside scoring.
    
    ## Key capabilities
    
    ### 1. Understanding Deployment Requirements
    
    Before making predictions, you need to understand what features a deployment requires:
    
    - **Feature names and types**: Know which columns are needed (numeric, categorical, text, date)
    - **Feature importance**: Understand which features matter most
    - **Target information**: Know what you're predicting
    - **Time series configuration**: If applicable, understand datetime columns and series IDs
    
    ### 2. Generating Prediction Datasets
    
    Create properly formatted prediction datasets:
    
    - Generate CSV templates with all required columns
    - Include sample values appropriate for each feature type
    - Add metadata comments explaining the model structure
    - Ensure correct column ordering
    
    ### 3. Validating Prediction Data
    
    Validate datasets before making predictions:
    
    - Check for missing required features
    - Verify data types match expected types
    - Identify missing low-importance features (warnings)
    - Note extra columns that will be ignored
    
    ### 4. Making Predictions
    
    Execute predictions using various methods:
    
    - **Real-time predictions**: Fast, synchronous predictions for individual records
    - **Batch predictions**: Process large datasets efficiently
    - **Time series predictions**: Handle forecasting scenarios with proper datetime handling
    
    ## Workflow examples
    
    ### Example 1: Generate prediction dataset for a new scenario
    
    **User request**: "I want to predict sales for next week for store_A with temperatures of 75°F each day and no promotions."
    
    **Agent workflow**:
    1. Get deployment features to understand required columns
    2. Generate a prediction data template with 7 rows (one week)
    3. Fill in the template with user's specific values:
       - Set temperature = 75 for all rows
       - Set promotion = 0 for all rows
       - Set store_id = "store_A" for all rows
       - Set dates for next 7 days
    4. Validate the data to ensure it's correct
    5. Make predictions using the validated dataset
    
    ### Example 2: Batch scoring a CSV file
    
    **User request**: "Score all records in my prediction_data.csv file using deployment abc123."
    
    **Agent workflow**:
    1. Validate the CSV file structure matches deployment requirements
    2. Upload the file or provide file path
    3. Submit batch prediction job
    4. Monitor job status
    5. Retrieve and return prediction results
    
    ## Using DataRobot SDK
    
    This skill guides you to use the DataRobot Python SDK directly. Install the SDK if needed:
    
    ```bash
    pip install datarobot datarobot-predict
    ```
    
    ### Key SDK Operations
    
    Use these DataRobot SDK methods to work with predictions:
    
    **Deployment Information**:
    - `dr.Deployment.get(deployment_id)` - Get deployment details
    - `deployment.get_features()` - Get required features (name/type/importance)
    
    **Predictions**:
    - `deployment.predict_batch(source)` - Convenience batch prediction API (CSV path, file object, or pandas DataFrame)
    - `dr.BatchPredictionJob.score(deployment=deployment, ...)` - Advanced batch prediction control
    - `job.get_result_when_complete()` - Wait for batch scoring to finish and download results
    
    **Data Management**:
    - `dr.Dataset.create_from_file(file_path)` - Upload dataset
    - `dr.Dataset.get(dataset_id)` - Get dataset info
    
    See the [Common Patterns](#common-patterns) section below for complete examples.
    
    ## Prediction Explanations
    
    Deployments can return per-row explanations (top feature contributions) alongside predictions.
    Two algorithms are available depending on how the deployment was configured:
    
    - **SHAP** (`shap`): SHapley Additive exPlanations. Available on tree-based models when SHAP was
      enabled at deployment time. Returns signed contributions in the model's score space.
    - **XEMP** (`xemp`): DataRobot's eXplainable AI for the eXact Model Prediction. Default when SHAP
      is not enabled. Returns top-N strongest features with a qualitative strength (`+++`, `--`, etc.).
    
    If you omit `explanation_algorithm`, the deployment's default is used.
    
    ### How to request explanations
    
    Pass `max_explanations=N` (and any optional filters) when calling `datarobot_predict.deployment.predict`:
    
    ```python
    import datarobot as dr
    import pandas as pd
    from datarobot_predict.deployment import predict as dr_predict
    
    dr.Client()
    deployment = dr.Deployment.get("abc123")
    
    result = dr_predict(
        deployment=deployment,
        data_frame=pd.DataFrame([{"feature1": 10, "feature2": 20}]),
        max_explanations=3,  # top 3 contributors per row
        explanation_algorithm="shap",  # or "xemp"; omit for deployment default
        # threshold_high=0.8,                # optional: only explain rows scoring > 0.8
        # threshold_low=0.2,                 # optional: only explain rows scoring < 0.2
        # passthrough_columns="all",         # optional: echo input columns through to output
    )
    print(result.dataframe.to_dict(orient="records"))
    ```
    
    The result DataFrame includes columns like `EXPLANATION_1_FEATURE_NAME`,
    `EXPLANATION_1_ACTUAL_VALUE`, `EXPLANATION_1_STRENGTH`, `EXPLANATION_1_QUALITATIVE_STRENGTH` for
    each of the top-N contributors.
    
    ### Parameter reference
    
    | Parameter | Purpose |
    |-----------|---------|
    | `max_explanations` | Top-N contributors per row. `0` (default) disables explanations. |
    | `max_ngram_explanations` | Text models only: cap text-segment explanations per row. |
    | `threshold_high` | Only explain rows with prediction probability **above** this (0–1). |
    | `threshold_low` | Only explain rows with prediction probability **below** this (0–1). |
    | `explanation_algorithm` | `"shap"` or `"xemp"`; omit to use deployment default. |
    | `passthrough_columns` | `"all"` or set of input column names to echo through to output. |
    
    ### CLI shortcut
    
    ```bash
    python scripts/make_prediction.py abc123 '{"feature1": 10, "feature2": 20}' \
        --max-explanations 3 --explanation-algorithm shap
    ```
    
    ### When to use which threshold
    
    - `threshold_high` is useful when only positive (high-risk / fraud / churn-likely) predictions need
      explaining — saves compute on a large batch.
    - `threshold_low` is the mirror image for low-probability rows.
    - Setting both restricts explanations to rows outside the `[low, high]` band.
    
    ### Common errors
    
    - **"Prediction explanations not enabled"**: the deployment was created without explanations support.
      Re-deploy the model with explanations enabled, or use a deployment that has them.
    - **`max_explanations` ignored / no explanation columns in output**: confirm you're calling
      `datarobot_predict.deployment.predict(...)` and that the deployment has explanations enabled.
      The `deployment.predict_batch()` convenience wrapper on the SDK is intended for plain scoring;
      use `datarobot_predict.deployment.predict` when you need explanation kwargs.
    
    ## Helper Scripts
    
    This skill includes executable helper scripts that Claude can run directly:
    
    - `scripts/get_deployment_features.py` - Get deployment feature requirements
    - `scripts/generate_prediction_data_template.py` - Generate CSV template
    - `scripts/validate_prediction_data.py` - Validate prediction data
    - `scripts/make_prediction.py` - Make real-time predictions
    
    **Usage example**:
    ```bash
    # Get deployment features
    python scripts/get_deployment_features.py abc123
    
    # Generate template
    python scripts/generate_prediction_data_template.py abc123 10 template.csv
    
    # Validate data
    python scripts/validate_prediction_data.py abc123 prediction_data.csv
    
    # Make prediction
    python scripts/make_prediction.py abc123 '{"feature1": 10, "feature2": 20}'
    
    # Make prediction with top-3 SHAP explanations
    python scripts/make_prediction.py abc123 '{"feature1": 10, "feature2": 20}' \
        --max-explanations 3 --explanation-algorithm shap
    ```
    
    Claude can run these scripts directly or use them as reference when writing code.
    
    ## Best practices
    
    1. **Always validate first**: Validate prediction data before submitting predictions to catch errors early
    2. **Use templates**: Generate templates to ensure correct structure and avoid missing columns
    3. **Check feature types**: Ensure numeric features are numbers, categorical features match training values
    4. **Handle time series**: For time series models, ensure datetime columns and series IDs are properly formatted
    5. **Monitor batch jobs**: For large batch predictions, check job status and handle errors appropriately
    
    ## Common patterns
    
    ### Pattern 1: Get deployment features and make single prediction (optionally with explanations)
    ```python
    import datarobot as dr
    import pandas as pd
    from datarobot_predict.deployment import predict as dr_predict
    
    # Initialize client
    dr.Client()
    
    deployment = dr.Deployment.get("abc123")
    
    prediction_data = {
        "feature1": value1,
        "feature2": value2,
        # ... all required features (excluding target)
    }
    
    # Score one row. Add max_explanations=N to get top-N explanations per row.
    result = dr_predict(
        deployment=deployment,
        data_frame=pd.DataFrame([prediction_data]),
        max_explanations=3,  # optional; 0/omit to disable explanations
        explanation_algorithm="shap",  # optional; omit to use deployment default
    )
    print(result.dataframe.to_dict(orient="records"))
    ```
    
    ### Pattern 2: Generate template and batch predictions
    ```python
    import datarobot as dr
    import pandas as pd
    
    # Initialize client
    dr.Client()
    
    # Get deployment features
    deployment = dr.Deployment.get("abc123")
    model = dr.Model.get(deployment.model["id"])
    features = model.get_features()
    
    # Create template DataFrame
    prediction_features = [f for f in features if f.name != model.target_name]
    template_df = pd.DataFrame(columns=[f.name for f in prediction_features])
    
    # Add sample rows
    for i in range(100):
        row = {}
        for feature in prediction_features:
            if feature.feature_type == "Numeric":
                row[feature.name] = 0.0
            elif feature.feature_type == "Categorical":
                row[feature.name] = "sample_value"
            else:
                row[feature.name] = ""
        template_df = pd.concat([template_df, pd.DataFrame([row])], ignore_index=True)
    
    # Save template
    template_df.to_csv("prediction_template.csv", index=False)
    
    # Fill template with actual data (modify CSV as needed)
    # ...
    
    # Submit batch prediction
    job = dr.BatchPredictionJob.score(
        deployment_id=deployment.id,
        intake_settings={"type": "localFile", "file": "prediction_template.csv"},
        output_settings={"type": "localFile", "path": "predictions_output.csv"},
    )
    
    # Monitor job
    job_status = dr.BatchPredictionJob.get(job.id)
    print(f"Job status: {job_status.status}")
    
    # Download results when complete
    if job_status.status == "completed":
        results = dr.BatchPredictionJob.download(job.id)
    ```
    
    ## Error handling
    
    Common errors and solutions:
    
    - **Missing required features**: Use `get_deployment_features` to get complete list
    - **Wrong data types**: Check feature types and convert accordingly
    - **Invalid categorical values**: Use training data sample to see valid values
    - **Time series format errors**: Ensure datetime format matches training data
    
    ## SDK Setup
    
    ### Install DataRobot SDK
    
    ```bash
    pip install datarobot datarobot-predict
    ```
    
    ### Initialize Client
    
    ```python
    import datarobot as dr
    
    # Initialize client with API credentials
    dr.Client()
    ```
    
    ### Environment Variables
    
    Set these environment variables or pass them directly:
    
    - `DATAROBOT_API_TOKEN` - Your DataRobot API token
    - `DATAROBOT_ENDPOINT` - Your DataRobot endpoint (default: https://app.datarobot.com)
    
    ## Resources
    
    - [DataRobot Python SDK Documentation](https://datarobot-public-api-client.readthedocs-hosted.com/)
    - [DataRobot Predictions Documentation](https://docs.datarobot.com/en/docs/predictions/index.html)
    - [DataRobot API Reference](https://docs.datarobot.com/en/docs/api/reference/index.html)
    - [Batch Predictions Guide](https://docs.datarobot.com/en/docs/api/reference/sdk/batch-predictions.html)
    
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related