charting-vega-lite
Create interactive data visualizations using Vega-Lite declarative JSON grammar. Supports 20+ chart types (bar, line, scatter, histogram, boxplot, grouped/stacked variations, etc.) via templates and programmatic builders. Use when users upload data for charting, request specific
Install
npx skills add https://github.com/oaustegard/claude-skills/tree/main/plugins/data-and-visualization/skills/charting-vega-lite
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oaustegard-claude-skills@llmmart
git clone https://github.com/oaustegard/claude-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oaustegard/claude-skills collection as a plugin from our marketplace. Git is the plain clone.
README
charting-vega-lite
Create interactive data visualizations using Vega-Lite declarative JSON grammar. Supports 20+ chart types (bar, line, scatter, histogram, boxplot, grouped/stacked variations, etc.) via templates and programmatic builders. Use when users upload data for charting, request specific chart types, or mention visualizations. Produces portable JSON specs with inline data islands that work in Claude artifacts and can be adapted for production.
Skill manifest
Overview
This skill creates interactive Vega-Lite visualizations from uploaded data. The workflow:
- Analyze data structure and context
- Select 5-10 meaningful chart types based on what the data represents
- Build chart specifications programmatically
- Generate React artifact with embedded visualizations
Critical Technical Constraint: Inline Data Island
Claude artifacts cannot use fetch() for computer:// URLs.
All data must be embedded as an inline JavaScript constant:
const DATA = [ /* embedded data array */ ];
// Later in chart specs:
spec.data = { values: DATA };
DO NOT:
- Use fetch() to load external files
- Reference external data URLs
- Create separate data files
This is the only pattern that works in Claude's artifact environment.
Primary Workflow: Data Upload → Chart Explorer
Execute this sequence when user uploads data without specifying chart type:
1. Analyze Data Structure
python /mnt/skills/user/charting-vega-lite/scripts/analyze_data.py /mnt/user-data/uploads/<filename>
Extract from output:
fields[](with types and statistics)suggested_charts[](suggested chart types with encodings)sample_data(first 10 rows for understanding context)
If script fails: Use manual pandas analysis
import pandas as pd
df = pd.read_csv('/mnt/user-data/uploads/<filename>')
# Classify: numeric→quantitative, datetime→temporal, <20 unique→nominal
2. Understand Data Context
Read sample data and column names to infer what the data represents:
- Biomedical data? → Biomarkers, patient outcomes, clinical relevance
- Financial data? → Trends, comparisons, performance metrics
- Sensor data? → Temporal patterns, anomalies, correlations
- E-commerce? → Sales trends, product comparisons, conversions
Ask: What questions would someone analyzing this data want answered?
Examples:
- Assay data: Which biomarkers strongest? Patterns across samples? Variability?
- Financial: What are trends? How volatile? Seasonal patterns?
- IoT: Temporal patterns? Anomalies? Sensor correlations?
3. Select Meaningful Charts (5-10 suggestions)
Filter analyze_data.py suggestions based on context and readability:
Apply readability filters:
- Pie chart with >7 categories → Skip (unreadable)
- Heatmap with >50 categories per axis → Aggregate first
- Multi-line with >10 series → Consider faceting
Prioritize charts that answer domain questions:
- Comparison needs → Bar, box plot, grouped bar
- Distribution analysis → Histogram, box plot
- Pattern recognition → Heatmap, scatter
- Temporal trends → Line, area
- Part-to-whole → Stacked bar (pie only if <7 categories)
Don't suggest charts just because data types match - choose charts that reveal insights.
4. Generate Chart Specs
Build specs programmatically using analyze_data.py encodings:
For each suggested chart type, construct spec using:
- Templates from
assets/templates/for basic types (bar, line, scatter, pie, heatmap, area) - Builder patterns from
references/spec-builder-patterns.mdfor variations (histogram, boxplot, grouped-bar, etc.) - Vega-Lite examples from
references/vega-lite-examples-inventory.mdfor uncommon types
Structure each chart as:
{"type": "Chart Name", "reason": "Why this chart", "spec": {/* vega-lite spec */}}
5. Create Artifact with Inline Data Island
Load data, read template, replace __DATA__ and __CHART_SPECS__ placeholders, write using bash heredoc.
6. Provide Link
[View chart explorer](computer:///mnt/user-data/outputs/ChartExplorer.jsx)
Created 7 contextually relevant charts for your data.
Secondary Workflow: Specific Chart Request
When user specifies chart type (e.g., "make a bar chart"):
1. Analyze Data
python /mnt/skills/user/charting-vega-lite/scripts/analyze_data.py /mnt/user-data/uploads/<filename>
2. Validate Chart Fits Data
Check requirements:
- Bar: needs 1 nominal + 1 quantitative
- Line: needs 1 temporal + 1 quantitative
- Scatter: needs 2 quantitative
- Heatmap: needs 2 nominal + 1 quantitative
- Pie: needs 1 nominal + 1 quantitative + <7 categories
If data doesn't fit:
- Explain: "Bar chart needs categorical data, but all columns are numeric"
- Suggest 2-3 alternatives
- Use Primary Workflow to create explorer with alternatives
3. Generate Spec
Use templates or programmatic builders based on chart type complexity.
4. Create Artifact
Same pattern as Primary Workflow step 5, but with single chart.
Error Prevention
Common failures:
Using fetch() in artifacts
- Solution: Always use inline data island pattern
- Never create external data files
Chart doesn't render
- Verify scripts load: Vega → Vega-Lite → Vega-Embed
- Check data is injected:
spec.data = {values: DATA} - Confirm field names match data columns
Generic/random chart suggestions
- Solution: Consider data context and meaning
- Filter suggestions for relevance and readability
- Prioritize charts that answer meaningful questions
Resources
Scripts:
scripts/analyze_data.py- analyze structure, suggest 8-12 chart types
Components:
assets/components/ChartExplorer.jsx- multi-chart explorer template
Templates:
assets/templates/*.json- 6 basic chart templates (bar, line, scatter, pie, heatmap, area)
References - Progressive Disclosure:
Read spec-builder-patterns.md when building charts programmatically (histogram, boxplot, grouped/stacked bars, multi-line, etc.)
Read vega-lite-examples-inventory.md when user requests uncommon chart type not in spec-builder-patterns
Read chart-types.md when validating specific chart requirements or user asks "what chart should I use for..."
Read advanced-charts.md for complete specs of specialized charts (sankey, waterfall, violin plots, complex layered compositions)
Read contextual-chart-selection.md for extended domain examples if unfamiliar with data domain (biomedical, financial, IoT, etc.)
Read online-resources.md to fetch Vega-Lite docs for advanced features (custom selections, transforms, conditional encoding)
Complete Workflow Example
User uploads assay data CSV (51 assays, 74 samples)
# 1. Analyze
python /mnt/skills/user/charting-vega-lite/scripts/analyze_data.py /mnt/user-data/uploads/assay_data.csv
# 2. Understand context: Multi-analyte immunoassay
# Questions: Which biomarkers strongest? Patterns across samples? Variability?
# 3. Build contextual charts (5-7 specs)
# Bar: Mean signal by assay
# Heatmap: Sample × Assay
# Box plot: Signal distribution by assay
# Histogram: Overall signal distribution
# etc.
# 4. Load data and template
df = pd.read_csv('/mnt/user-data/uploads/assay_data.csv')
data = df.to_dict(orient='records')
template = open('/mnt/skills/user/charting-vega-lite/assets/components/ChartExplorer.jsx').read()
# 5. Replace placeholders and write
artifact = template.replace('__DATA__', json.dumps(data)).replace('__CHART_SPECS__', json.dumps(charts))
# Use bash heredoc to avoid XML conflicts in tool parameters
# 6. Provide link
Created 7 charts for your assay data - bar charts show biomarker signals, heatmap reveals sample patterns, box plots display variability.
Critical Rules
- ALWAYS use inline data island pattern - No fetch(), no external files
- Consider data context - Choose meaningful charts based on what data represents, not just data types
- Filter by readability - Avoid charts with too many categories
- Use bash heredoc for file creation - Prevents XML conflicts when creating artifacts
- Provide links, not content - Output token efficiency
Files (claude-skills)
-
assets
-
components
-
ChartExplorer.jsx 7 KB · in bundle
-
-
templates
-
area.json 464 B
{ "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": "container", "height": 400, "data": {"values": "__DATA__"}, "mark": {"type": "area", "line": true, "point": true}, "encoding": { "x": {"field": "__X_FIELD__", "type": "__X_TYPE__"}, "y": {"field": "__Y_FIELD__", "type": "__Y_TYPE__"}, "tooltip": [ {"field": "__X_FIELD__", "type": "__X_TYPE__"}, {"field": "__Y_FIELD__", "type": "__Y_TYPE__"} ] } } -
bar.json 606 B
{ "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": "container", "height": 400, "data": {"values": "__DATA__"}, "mark": {"type": "bar", "cornerRadius": 5, "tooltip": true}, "encoding": { "x": {"field": "__X_FIELD__", "type": "__X_TYPE__", "axis": {"labelAngle": 0}}, "y": {"field": "__Y_FIELD__", "type": "__Y_TYPE__"}, "color": {"field": "__Y_FIELD__", "type": "__Y_TYPE__", "scale": {"scheme": "viridis"}, "legend": null}, "tooltip": [ {"field": "__X_FIELD__", "type": "__X_TYPE__"}, {"field": "__Y_FIELD__", "type": "__Y_TYPE__"} ] } } -
heatmap.json 588 B
{ "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": "container", "height": 400, "data": {"values": "__DATA__"}, "mark": "rect", "encoding": { "x": {"field": "__X_FIELD__", "type": "__X_TYPE__"}, "y": {"field": "__Y_FIELD__", "type": "__Y_TYPE__"}, "color": {"field": "__COLOR_FIELD__", "type": "__COLOR_TYPE__", "scale": {"scheme": "viridis"}}, "tooltip": [ {"field": "__X_FIELD__", "type": "__X_TYPE__"}, {"field": "__Y_FIELD__", "type": "__Y_TYPE__"}, {"field": "__COLOR_FIELD__", "type": "__COLOR_TYPE__"} ] } } -
line.json 536 B
{ "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": "container", "height": 400, "data": {"values": "__DATA__"}, "mark": {"type": "line", "point": true, "tooltip": true}, "encoding": { "x": {"field": "__X_FIELD__", "type": "__X_TYPE__"}, "y": {"field": "__Y_FIELD__", "type": "__Y_TYPE__"}, "color": {"field": "__COLOR_FIELD__", "type": "__COLOR_TYPE__"}, "tooltip": [ {"field": "__X_FIELD__", "type": "__X_TYPE__"}, {"field": "__Y_FIELD__", "type": "__Y_TYPE__"} ] } } -
pie.json 510 B
{ "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": "container", "height": 400, "data": {"values": "__DATA__"}, "mark": {"type": "arc", "innerRadius": 50, "tooltip": true}, "encoding": { "theta": {"field": "__THETA_FIELD__", "type": "__THETA_TYPE__"}, "color": {"field": "__COLOR_FIELD__", "type": "__COLOR_TYPE__"}, "tooltip": [ {"field": "__COLOR_FIELD__", "type": "__COLOR_TYPE__"}, {"field": "__THETA_FIELD__", "type": "__THETA_TYPE__"} ] } } -
scatter.json 535 B
{ "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": "container", "height": 400, "data": {"values": "__DATA__"}, "mark": {"type": "point", "size": 100, "tooltip": true}, "encoding": { "x": {"field": "__X_FIELD__", "type": "__X_TYPE__"}, "y": {"field": "__Y_FIELD__", "type": "__Y_TYPE__"}, "color": {"field": "__COLOR_FIELD__", "type": "__COLOR_TYPE__"}, "tooltip": [ {"field": "__X_FIELD__", "type": "__X_TYPE__"}, {"field": "__Y_FIELD__", "type": "__Y_TYPE__"} ] } }
-
-
-
references
-
advanced-charts.md 9.2 KB
# Advanced Chart Patterns Complete spec structures for specialized charts beyond basic templates. ## Statistical Charts ## Heatmap ```javascript { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "data": {"values": data}, "mark": "rect", "encoding": { "x": {"field": "category_x", "type": "nominal"}, "y": {"field": "category_y", "type": "nominal"}, "color": { "field": "value", "type": "quantitative", "scale": {"scheme": "viridis"} }, "tooltip": [ {"field": "category_x", "type": "nominal"}, {"field": "category_y", "type": "nominal"}, {"field": "value", "type": "quantitative"} ] } } ``` ## Box Plot ```javascript { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "data": {"values": data}, "mark": { "type": "boxplot", "extent": "min-max" }, "encoding": { "x": {"field": "category", "type": "nominal"}, "y": {"field": "value", "type": "quantitative"} } } ``` ## Violin Plot ```javascript { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "data": {"values": data}, "transform": [ { "density": "value", "groupby": ["category"], "as": ["value", "density"] } ], "mark": "area", "encoding": { "x": { "field": "density", "type": "quantitative", "stack": "center", "impute": null, "axis": null }, "y": {"field": "value", "type": "quantitative"}, "color": {"field": "category", "type": "nominal"}, "column": {"field": "category", "type": "nominal"} } } ``` ## Waterfall Chart ```javascript { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "data": {"values": data}, "transform": [ {"window": [{"op": "sum", "field": "amount", "as": "sum"}]}, {"window": [{"op": "lead", "field": "label", "as": "lead"}]}, { "calculate": "datum.lead === null ? datum.label : datum.lead", "as": "lead" }, { "calculate": "datum.label === 'Begin' ? 0 : datum.sum - datum.amount", "as": "previous_sum" }, { "calculate": "datum.label === 'Begin' || datum.label === 'End' ? 0 : datum.amount", "as": "amount" }, { "calculate": "(datum.label !== 'Begin' && datum.label !== 'End' && datum.amount > 0 ? '+' : '') + datum.amount", "as": "text_amount" }, {"calculate": "datum.sum + datum.amount", "as": "sum_end"} ], "encoding": {"x": {"field": "label", "type": "nominal", "sort": null}}, "layer": [ { "mark": {"type": "bar", "size": 45}, "encoding": { "y": {"field": "previous_sum", "type": "quantitative"}, "y2": {"field": "sum"}, "color": { "condition": [ {"test": "datum.label === 'Begin' || datum.label === 'End'", "value": "#878d96"}, {"test": "datum.sum < datum.previous_sum", "value": "#d33"} ], "value": "#24a148" } } } ] } ``` ## Sankey Diagram Use layers to approximate Sankey flow: ```javascript { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "data": {"values": data}, // format: [{source, target, value}] "transform": [ { "lookup": "source", "from": { "data": {"values": nodes}, // [{id, order}] "key": "id", "fields": ["order"] }, "as": ["source_order"] }, { "lookup": "target", "from": { "data": {"values": nodes}, "key": "id", "fields": ["order"] }, "as": ["target_order"] } ], "mark": {"type": "bar", "cornerRadiusEnd": 4}, "encoding": { "x": {"field": "source_order", "type": "ordinal", "axis": null}, "x2": {"field": "target_order"}, "y": {"field": "value", "type": "quantitative", "stack": "normalize"}, "color": {"field": "source", "type": "nominal"} } } ``` ## Calendar Heatmap ```javascript { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "data": {"values": data}, // format: [{date, value}] "transform": [ {"calculate": "year(datum.date)", "as": "year"}, {"calculate": "week(datum.date)", "as": "week"}, {"calculate": "day(datum.date)", "as": "day"} ], "mark": "rect", "encoding": { "x": {"field": "week", "type": "ordinal", "title": "Week"}, "y": {"field": "day", "type": "ordinal", "title": "Day"}, "color": { "field": "value", "type": "quantitative", "scale": {"scheme": "blues"} }, "facet": {"field": "year", "type": "nominal", "columns": 1}, "tooltip": [ {"field": "date", "type": "temporal"}, {"field": "value", "type": "quantitative"} ] } } ``` ## Horizon Chart ```javascript { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "data": {"values": data}, "transform": [ {"calculate": "datum.value > 0 ? datum.value : 0", "as": "positive"}, {"calculate": "datum.value < 0 ? -datum.value : 0", "as": "negative"} ], "facet": {"field": "category", "type": "nominal"}, "spec": { "height": 50, "layer": [ { "mark": {"type": "area", "clip": true}, "encoding": { "x": {"field": "date", "type": "temporal"}, "y": {"field": "positive", "type": "quantitative"}, "color": {"value": "#08519c"} } }, { "mark": {"type": "area", "clip": true}, "encoding": { "x": {"field": "date", "type": "temporal"}, "y": {"field": "negative", "type": "quantitative"}, "color": {"value": "#a50f15"} } } ] } } ``` ## Radial Chart (Polar) ```javascript { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "data": {"values": data}, "layer": [ { "mark": {"type": "arc", "innerRadius": 20, "stroke": "#fff"} } ], "encoding": { "theta": {"field": "value", "type": "quantitative", "stack": true}, "radius": {"field": "value", "type": "quantitative", "scale": {"type": "sqrt", "zero": true, "rangeMin": 20}}, "color": {"field": "category", "type": "nominal"} } } ``` ## Bubble Chart with Size Legend ```javascript { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "data": {"values": data}, "mark": "point", "encoding": { "x": {"field": "x", "type": "quantitative"}, "y": {"field": "y", "type": "quantitative"}, "size": { "field": "size", "type": "quantitative", "scale": {"range": [0, 5000]}, "legend": {"title": "Size"} }, "color": {"field": "category", "type": "nominal"}, "tooltip": [ {"field": "name", "type": "nominal"}, {"field": "x", "type": "quantitative"}, {"field": "y", "type": "quantitative"}, {"field": "size", "type": "quantitative"} ] } } ``` ## Stacked Area with Normalized View ```javascript { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "data": {"values": data}, "mark": "area", "encoding": { "x": {"field": "date", "type": "temporal"}, "y": { "field": "value", "type": "quantitative", "stack": "normalize" // or "center" for streamgraph }, "color": {"field": "category", "type": "nominal"} } } ``` ## Error Bars ```javascript { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "data": {"values": data}, "layer": [ { "mark": {"type": "errorbar", "extent": "stdev"}, "encoding": { "x": {"field": "category", "type": "nominal"}, "y": {"field": "value", "type": "quantitative"} } }, { "mark": {"type": "point", "filled": true}, "encoding": { "x": {"field": "category", "type": "nominal"}, "y": {"aggregate": "mean", "field": "value", "type": "quantitative"} } } ] } ``` ## Isotype Grid ```javascript { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "data": {"values": data}, "transform": [ { "window": [{"op": "row_number", "as": "id"}], "groupby": ["category"] }, {"calculate": "ceil(datum.id / 10)", "as": "row"}, {"calculate": "datum.id % 10", "as": "col"} ], "mark": {"type": "point", "filled": true, "size": 100}, "encoding": { "x": {"field": "col", "type": "ordinal", "axis": null}, "y": {"field": "row", "type": "ordinal", "axis": null}, "color": {"field": "category", "type": "nominal"}, "facet": {"field": "category", "type": "nominal"} } } ``` ## Slope Chart ```javascript { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "data": {"values": data}, // format: [{category, period, value}] "mark": "line", "encoding": { "x": {"field": "period", "type": "ordinal"}, "y": {"field": "value", "type": "quantitative"}, "color": {"field": "category", "type": "nominal"}, "detail": {"field": "category", "type": "nominal"} } } ``` ## Sparklines (Small Multiples) ```javascript { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "data": {"values": data}, "facet": { "field": "category", "type": "nominal", "columns": 1 }, "spec": { "width": 300, "height": 30, "mark": "line", "encoding": { "x": { "field": "date", "type": "temporal", "axis": {"title": "", "labels": false} }, "y": { "field": "value", "type": "quantitative", "axis": null, "scale": {"zero": false} } } } } ``` -
chart-types.md 5.3 KB
# Chart Type Selection Reference Comprehensive catalog of chart types with data requirements and use cases. ## Standard Chart Types ### Categorical Comparison (Nominal + Quantitative) **Bar Chart** - Compare values across categories - Data: One categorical field, one numeric field - Use: Rankings, comparisons, distributions - Template: `assets/templates/bar.json` **Grouped Bar** - Compare multiple series across categories - Data: Two categorical fields, one numeric field - Use: Multi-series comparisons, subcategory analysis - Build: Use color encoding for second categorical field **Stacked Bar** - Show part-to-whole with categories - Data: Two categorical fields, one numeric field - Use: Composition within categories - Build: Add `"stack": true` to y encoding ### Time Series (Temporal + Quantitative) **Line Chart** - Show trends over time - Data: Date/time field, one or more numeric fields - Use: Trends, patterns, temporal relationships - Template: `assets/templates/line.json` **Area Chart** - Emphasize magnitude over time - Data: Date/time field, one numeric field - Use: Volume visualization, cumulative values - Template: `assets/templates/area.json` **Stacked Area** - Multiple series composition over time - Data: Date/time field, category field, numeric field - Use: Part-to-whole temporal analysis - Build: Add color encoding by category, stack areas ### Correlation (Two Quantitative) **Scatter Plot** - Reveal relationships between variables - Data: Two numeric fields, optional category for color - Use: Correlation analysis, outlier detection, clustering - Template: `assets/templates/scatter.json` **Bubble Chart** - Three-dimensional quantitative relationships - Data: Three numeric fields (x, y, size), optional category - Use: Multi-variable comparison - Build: Add size encoding to scatter plot ### Part-to-Whole (Nominal + Quantitative) **Pie Chart** - Show proportions of a whole - Data: One categorical field, one numeric field - Use: Simple proportions, limited categories (<7) - Template: `assets/templates/pie.json` **Donut Chart** - Pie with center emphasis - Data: Same as pie chart - Use: Same as pie, with focus metric in center - Build: Set `innerRadius` in arc mark ### Two-Dimensional Categorical (Two Nominal + Quantitative) **Heatmap** - Intensity across two dimensions - Data: Two categorical fields, one numeric field - Use: Patterns in matrix data, correlation matrices - Template: `assets/templates/heatmap.json` ### Distribution (Quantitative) **Histogram** - Frequency distribution - Data: One numeric field - Use: Distribution shape, outlier detection - Build: Use bin transform with bar mark **Box Plot** - Statistical distribution summary - Data: Categorical field (groups), numeric field (values) - Use: Compare distributions, identify outliers - Build: Use boxplot mark type ## Advanced Patterns ### Multi-View Compositions **Layered** - Overlay multiple marks ```javascript "layer": [ {"mark": "line", "encoding": {...}}, {"mark": "point", "encoding": {...}} ] ``` Use: Combined chart types (line + points, area + line) **Faceted** - Small multiples ```javascript "facet": {"field": "category", "type": "nominal"}, "spec": {"mark": "bar", "encoding": {...}} ``` Use: Compare patterns across categories **Concatenated** - Side-by-side views ```javascript "hconcat": [ {"mark": "bar", ...}, {"mark": "line", ...} ] ``` Use: Different aspects of same dataset ### Interactive Patterns **Brush and Link** - Cross-filtering across views - Use: Explore data across multiple dimensions - Pattern: Selection parameter shared between views **Zoom and Filter** - Detail on demand - Use: Large datasets with focus + context - Pattern: Dual-view with interval selection **Tooltips** - Contextual information - Use: Additional details on hover - Pattern: Multi-field tooltip encoding ## Selection Guide ### By Data Structure | X Field | Y Field | Color | Chart Type | |---------|---------|-------|------------| | Nominal | Quantitative | - | Bar | | Temporal | Quantitative | Nominal | Line (colored) | | Quantitative | Quantitative | Nominal | Scatter | | Nominal | Nominal | Quantitative | Heatmap | | - | - | Nominal/Theta | Pie | ### By Question | Question | Chart Type | Why | |----------|------------|-----| | How do categories compare? | Bar | Direct magnitude comparison | | What's the trend over time? | Line | Temporal continuity | | Are these variables related? | Scatter | Correlation visibility | | What's the distribution? | Histogram | Frequency patterns | | How do parts make up the whole? | Pie/Stacked Bar | Proportional relationships | | What patterns exist in 2D categories? | Heatmap | Density/intensity visualization | ### By Data Volume | Rows | Suggested Types | Avoid | |------|-----------------|-------| | <50 | Bar, Pie, Scatter | Aggregated views | | 50-500 | Line, Scatter, Bar | Pie (too many slices) | | 500-5000 | Line, Heatmap, Aggregated bar | Individual point scatter | | 5000+ | Binned/aggregated, Heatmap | Raw scatter, detailed bar | ## Template Usage All templates use placeholder syntax: - `__DATA__`: Will be replaced with actual data array - `__X_FIELD__`, `__Y_FIELD__`, etc.: Field names from data - `__X_TYPE__`, `__Y_TYPE__`, etc.: Vega-Lite type (nominal/quantitative/temporal/ordinal) Load template, replace placeholders, use resulting spec. -
contextual-chart-selection.md 5.8 KB
# Contextual Chart Selection ## Principle **Generic chart suggestions based solely on data types are insufficient.** Chart selection must consider: 1. What the data represents (domain/context) 2. What questions the user likely wants answered 3. Which visualizations will reveal meaningful patterns 4. Practical readability constraints ## Decision Framework ### Step 1: Understand Data Context **Read sample data and column names to infer domain:** ```python # Example: Assay data columns = ['Plate', 'Sample ID', 'Assay', 'Well', 'Signal', 'CV'] sample = df.head(10) # → This is multi-analyte immunoassay data ``` **Common data domains:** - **Biomedical:** Assays, patient records, clinical trials, imaging - **Financial:** Transactions, time series, performance metrics, portfolios - **IoT/Sensor:** Time series, spatial data, environmental monitoring - **E-commerce:** Sales, user behavior, inventory, conversions - **Scientific:** Experiments, measurements, observations, simulations ### Step 2: Map Domain to Analysis Questions **For each domain, anticipate user's analytical goals:** **Biomedical assay data:** - Which biomarkers have strongest signals? - Are there patterns across samples? - What's the variability per assay? - Any outliers or quality issues? - Do sample groups differ? **Financial time series:** - What are the trends over time? - How volatile is performance? - Are there seasonal patterns? - How do different assets compare? - Where are inflection points? **IoT sensor data:** - What are temporal patterns? - Are there anomalies? - How do sensors correlate? - What's the distribution of readings? - Are there spatial patterns? ### Step 3: Select Charts That Answer Questions **Match questions to chart types:** | Question | Chart Type | Why | |----------|-----------|-----| | Compare categories | Bar, Box plot | Clear magnitude comparison | | Show distribution | Histogram, Box plot | Reveals shape, outliers | | Reveal patterns in matrix | Heatmap | Shows 2D relationships | | Track changes over time | Line, Area | Temporal continuity | | Find correlations | Scatter, Bubble | Shows relationships | | Show composition | Stacked bar, Pie (if <7 categories) | Part-to-whole | | Compare groups | Grouped bar, Faceted plots | Multi-series comparison | ### Step 4: Apply Readability Filters **Eliminate impractical suggestions:** ```python # Too many categories for pie chart if chart_type == 'pie' and n_categories > 7: skip # Unreadable # Too many series for multi-line if chart_type == 'multi-line' and n_series > 10: consider_faceting_or_filtering # Heatmap with high cardinality if chart_type == 'heatmap' and (x_unique > 50 or y_unique > 50): consider_aggregation_or_sampling ``` ## Domain-Specific Examples ### Multi-Analyte Assay Data **Context indicators:** - Columns: Plate, Sample ID, Assay, Well, Signal, CV - 51 unique assays - 74 samples - Numeric signal values **Meaningful charts:** 1. **Bar: Mean Signal by Assay** → Which biomarkers are strongest? 2. **Heatmap: Sample × Assay** → Pattern recognition across plate 3. **Box plot: Signal by Assay** → Variability and quality metrics 4. **Histogram: Signal Distribution** → Overall data quality 5. **Scatter: Sample Index vs Signal** → Positional effects? 6. **Bar: Top 20 Assays** → Focus on most responsive markers **Avoid:** - Pie chart with 51 categories (unreadable) - Line chart without temporal axis (inappropriate) - Stacked bar (no meaningful composition question) ### Financial Time Series **Context indicators:** - Columns: Date, Ticker, Price, Volume - Temporal data - Multiple securities **Meaningful charts:** 1. **Line: Price over Time** → Trend analysis 2. **Multi-line: Multiple Securities** → Comparative performance 3. **Stacked area: Volume by Ticker** → Market composition 4. **Scatter: Volume vs Price Change** → Liquidity patterns 5. **Box plot: Returns by Ticker** → Volatility comparison **Avoid:** - Bar chart of daily prices (loses continuity) - Heatmap without meaningful 2D structure - Pie chart of volumes (temporal aspect lost) ### E-commerce Sales **Context indicators:** - Columns: Date, Product, Category, Revenue, Units - Transactional data - Multiple products/categories **Meaningful charts:** 1. **Line: Revenue over Time** → Growth trends 2. **Bar: Revenue by Category** → Category comparison 3. **Stacked bar: Revenue by Category over Time** → Composition changes 4. **Scatter: Units vs Revenue** → Pricing patterns 5. **Heatmap: Product × Month** → Seasonal patterns ## Anti-Patterns **DON'T suggest charts just because data types match:** ❌ "You have nominal and quantitative fields, so here's a bar chart" ✅ "Your assay data has 51 biomarkers - bar chart shows which have strongest signals" ❌ "Here are 10 chart types that technically work with your data" ✅ "Here are 7 charts that answer key questions about your assay results" ❌ "Pie chart available because you have categories" ✅ "Skipping pie chart - 51 categories would be unreadable. Bar chart better for comparison." ## Implementation Pattern ```python # 1. Infer domain domain = infer_domain_from_columns_and_samples(df) # 2. Get analysis questions for domain questions = get_domain_questions(domain) # 3. Map questions to chart types candidates = map_questions_to_charts(questions, df) # 4. Filter by readability viable_charts = [c for c in candidates if is_readable(c, df)] # 5. Prioritize by insight value charts = sorted(viable_charts, key=lambda c: c.insight_score, reverse=True)[:8] ``` ## Key Takeaway **Chart selection is an analytical decision, not a template-matching exercise.** Consider: 1. What does this data represent? 2. What will the user want to learn? 3. Which charts will reveal those insights? 4. Which charts are practically readable? This produces focused, meaningful visualizations rather than exhaustive but generic options. -
customization.md 4.4 KB
# Chart Customization Instructions **When to use this reference:** - User says "change colors", "different theme", "make it prettier" - User requests specific formatting (currency, dates, percentages) - User wants interactivity (hover, click, tooltips) - After creating base chart and user wants refinements **How to apply customizations:** 1. Locate the relevant section below 2. Copy the JSON modification 3. Apply to the `spec` object in your chart 4. Regenerate artifact with modified spec 5. Provide new link ## Color Schemes **Execute when user requests color changes.** Replace in spec: ```javascript "encoding": { "color": { "field": "category", "type": "nominal", "scale": {"scheme": "tableau10"} // Change scheme here } } ``` **Sequential** (quantitative data): - `viridis`, `plasma`, `inferno`, `magma` - `blues`, `greens`, `oranges`, `purples`, `reds` **Diverging** (data with meaningful center): - `redblue`, `purplegreen`, `blueorange` **Categorical** (nominal data): - `tableau10` (recommended), `category20` ## Axis Formatting ### Numbers ```javascript "encoding": { "y": { "field": "value", "type": "quantitative", "axis": { "format": "$,.2f" // Currency with 2 decimals } } } ``` Formats: - `,.0f` - Thousands separator, no decimals - `.2f` - 2 decimal places - `.1%` - Percentage with 1 decimal - `.2e` - Scientific notation ### Dates ```javascript "encoding": { "x": { "field": "date", "type": "temporal", "axis": { "format": "%b %Y" // Jan 2024 } } } ``` Formats: - `%Y-%m-%d` - 2024-01-15 - `%b %Y` - Jan 2024 - `%B %d, %Y` - January 15, 2024 ## Tooltips ### Multi-field ```javascript "encoding": { "tooltip": [ {"field": "category", "type": "nominal", "title": "Product"}, {"field": "sales", "type": "quantitative", "title": "Sales", "format": "$,.0f"}, {"field": "date", "type": "temporal", "format": "%B %Y"} ] } ``` ## Titles and Labels ```javascript { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "title": { "text": "Sales Performance", "subtitle": "Q1 2024", "fontSize": 18, "anchor": "start" }, "data": {...}, "encoding": { "x": { "field": "month", "type": "temporal", "axis": {"title": "Month"} }, "y": { "field": "sales", "type": "quantitative", "axis": {"title": "Revenue (USD)"} } } } ``` ## Sizing ### Responsive ```javascript { "width": "container", // Fills parent "height": 400, "autosize": {"type": "fit", "contains": "padding"} } ``` ### Fixed ```javascript { "width": 800, "height": 400 } ``` ## Mark Styling ### Bars ```javascript "mark": { "type": "bar", "cornerRadius": 5, "opacity": 0.8 } ``` ### Lines ```javascript "mark": { "type": "line", "strokeWidth": 3, "point": true // Add points at data values } ``` ### Points ```javascript "mark": { "type": "point", "size": 100, "filled": true, "opacity": 0.8 } ``` ## Interactive Selections ### Highlight on hover ```javascript { "params": [{ "name": "hover", "select": {"type": "point", "on": "mouseover"} }], "mark": "bar", "encoding": { "opacity": { "condition": {"param": "hover", "value": 1}, "value": 0.5 } } } ``` ### Click to filter ```javascript { "params": [{ "name": "select", "select": {"type": "point"} }], "mark": "point", "encoding": { "color": { "condition": {"param": "select", "value": "steelblue"}, "value": "lightgray" } } } ``` ## Legend Configuration ```javascript "encoding": { "color": { "field": "category", "type": "nominal", "legend": { "title": "Product Type", "orient": "right", // "left", "top", "bottom" "labelFontSize": 12 } } } ``` Hide legend: ```javascript "legend": null ``` ## Grid Lines ```javascript "encoding": { "x": { "field": "date", "type": "temporal", "axis": { "grid": true, "gridColor": "#e0e0e0" } } } ``` ## Conditional Styling Color based on value: ```javascript "encoding": { "color": { "condition": { "test": "datum.value > 50", "value": "green" }, "value": "red" } } ``` ## Quick Modifications After creating base chart, most common customizations: 1. Color scheme (aesthetic preference) 2. Axis formatting (readability) 3. Tooltips (additional context) 4. Title/labels (clarity) 5. Sizing (layout fit) Apply modifications to exported spec, test in artifact, finalize for Angular. -
data-driven-workflow.md 8.8 KB
# Data-Driven Workflow: Edge Cases and Advanced Patterns **IMPORTANT: Standard workflows are in SKILL.md. Only consult this file for edge cases.** This reference covers scenarios NOT handled by Primary/Secondary workflows in SKILL.md: - Multi-file datasets requiring joins - Performance optimization for very large datasets (>100K rows) - Complex validation scenarios - Advanced error recovery patterns - Custom aggregation before visualization **For standard data upload → chart creation, use SKILL.md workflows directly.** ## Edge Case 1: Multi-File Datasets When user uploads multiple related files needing joins: ### Identify Join Scenario ```python # User uploads: sales.csv, products.csv, customers.csv # Asks: "Show sales by product category and customer region" import pandas as pd # Load files sales = pd.read_csv('/mnt/user-data/uploads/sales.csv') products = pd.read_csv('/mnt/user-data/uploads/products.csv') customers = pd.read_csv('/mnt/user-data/uploads/customers.csv') # Identify join keys print("Sales columns:", sales.columns.tolist()) print("Products columns:", products.columns.tolist()) print("Customers columns:", customers.columns.tolist()) # Look for: product_id, customer_id, etc. ``` ### Execute Joins ```python # Join datasets merged = sales.merge(products, on='product_id', how='left') merged = merged.merge(customers, on='customer_id', how='left') # Save merged dataset merged.to_csv('/mnt/user-data/uploads/merged_data.csv', index=False) # NOW proceed with standard workflow: # python prepare_data.py /mnt/user-data/uploads/merged_data.csv ``` ### Edge Case: Missing Join Keys If join keys don't match: ```python # Fuzzy matching or manual mapping required # Example: "product_name" vs "name" merged = sales.merge( products, left_on='product_name', right_on='name', how='left' ) ``` ## Edge Case 2: Performance Optimization (Large Datasets) When dataset >100K rows or file size >100MB: ### Strategy 1: Sampling ```python import pandas as pd df = pd.read_csv('/mnt/user-data/uploads/large_file.csv') print(f"Original: {len(df)} rows") # Random sample (maintains distribution) sampled = df.sample(n=10000, random_state=42) sampled.to_csv('/mnt/user-data/uploads/sampled_data.csv', index=False) # Proceed with sampled data # python prepare_data.py /mnt/user-data/uploads/sampled_data.csv # NOTE in chart title: "Based on 10K sample from 500K records" ``` ### Strategy 2: Aggregation ```python # If time series with high frequency df['date'] = pd.to_datetime(df['timestamp']) # Aggregate to hourly/daily instead of per-second aggregated = df.groupby(df['date'].dt.floor('H')).agg({ 'value': 'mean', 'count': 'sum' }).reset_index() aggregated.to_csv('/mnt/user-data/uploads/aggregated_data.csv', index=False) ``` ### Strategy 3: Filtering ```python # Focus on relevant subset filtered = df[df['category'].isin(['A', 'B', 'C'])] # Top 3 categories filtered = filtered[filtered['date'] >= '2024-01-01'] # Recent data only filtered.to_csv('/mnt/user-data/uploads/filtered_data.csv', index=False) ``` ## Edge Case 3: Data Quality Issues ### Missing Values ```python df = pd.read_csv('/mnt/user-data/uploads/file.csv') # Check missing values print(df.isnull().sum()) # Strategy: Drop rows with missing key fields df_clean = df.dropna(subset=['key_field1', 'key_field2']) # OR: Fill with meaningful defaults df['numeric_col'].fillna(df['numeric_col'].median(), inplace=True) df['category_col'].fillna('Unknown', inplace=True) df_clean.to_csv('/mnt/user-data/uploads/cleaned_data.csv', index=False) ``` ### Outliers ```python # Remove outliers (>3 standard deviations) from scipy import stats z_scores = np.abs(stats.zscore(df['value'])) df_no_outliers = df[z_scores < 3] df_no_outliers.to_csv('/mnt/user-data/uploads/no_outliers.csv', index=False) ``` ### Duplicate Records ```python # Check for duplicates duplicates = df.duplicated(subset=['id']).sum() print(f"Found {duplicates} duplicates") # Remove duplicates (keep first occurrence) df_dedup = df.drop_duplicates(subset=['id'], keep='first') df_dedup.to_csv('/mnt/user-data/uploads/deduped_data.csv', index=False) ``` ## Edge Case 4: Complex Validation When data structure is ambiguous: ### Validate Chart Requirements ```python def validate_bar_chart(df, x_field, y_field): """Validate data fits bar chart requirements.""" # Check x is categorical if df[x_field].nunique() > 50: return False, f"{x_field} has too many categories (>50)" # Check y is numeric if not pd.api.types.is_numeric_dtype(df[y_field]): return False, f"{y_field} is not numeric" # Check for nulls null_count = df[[x_field, y_field]].isnull().sum().sum() if null_count > 0: return False, f"Found {null_count} null values" return True, "Valid for bar chart" # Use validation valid, message = validate_bar_chart(df, 'category', 'value') if not valid: # Inform user: message # Suggest alternatives or data cleaning ``` ### Suggest Data Transformations ```python # Too many categories? Suggest grouping if df['category'].nunique() > 20: # Group by top N top_categories = df['category'].value_counts().head(10).index df['category_grouped'] = df['category'].apply( lambda x: x if x in top_categories else 'Other' ) # Use 'category_grouped' for charting ``` ## Edge Case 5: Advanced Error Recovery ### analyze_data.py Fails on Specific Data Types ```python # Example: Mixed type columns causing issues df = pd.read_csv('/mnt/user-data/uploads/problematic.csv') # Fix mixed types for col in df.columns: # Try numeric conversion try: df[col] = pd.to_numeric(df[col], errors='coerce') except: pass # Try datetime conversion try: df[col] = pd.to_datetime(df[col], errors='coerce') except: pass # Save fixed version df.to_csv('/mnt/user-data/uploads/fixed_data.csv', index=False) # Retry analyze_data.py ``` ### prepare_data.py Memory Issues ```python # For very large files, process in chunks chunk_size = 10000 chunks = [] for chunk in pd.read_csv('/mnt/user-data/uploads/huge_file.csv', chunksize=chunk_size): # Process chunk (filter, transform) processed = chunk[chunk['value'] > 0] chunks.append(processed) # Combine processed chunks result = pd.concat(chunks, ignore_index=True) result.to_csv('/mnt/user-data/uploads/processed_data.csv', index=False) ``` ## Edge Case 6: Custom Pre-Aggregation When visualization requires pre-computed metrics: ### Example: Category Totals for Bar Chart ```python # Raw data: transaction-level # Needed: category totals df = pd.read_csv('/mnt/user-data/uploads/transactions.csv') # Aggregate by category category_totals = df.groupby('category').agg({ 'amount': 'sum', 'count': 'count' }).reset_index() category_totals.columns = ['category', 'total_amount', 'transaction_count'] # Save aggregated data category_totals.to_csv('/mnt/user-data/uploads/category_summary.csv', index=False) # Now use standard workflow # python prepare_data.py /mnt/user-data/uploads/category_summary.csv ``` ### Example: Time Series Resampling ```python df = pd.read_csv('/mnt/user-data/uploads/timeseries.csv') df['timestamp'] = pd.to_datetime(df['timestamp']) df.set_index('timestamp', inplace=True) # Resample to daily (from minutely data) daily = df.resample('D').agg({ 'value': 'mean', 'count': 'sum' }).reset_index() daily.to_csv('/mnt/user-data/uploads/daily_data.csv', index=False) ``` ## Performance Guidelines **When to optimize:** - File size >50MB → Consider sampling or aggregation - Row count >100K → Sample or aggregate before visualization - Column count >50 → Select relevant columns only - Categories >50 → Group into top N + "Other" **Optimization decision tree:** ``` IF rows > 100K AND user wants overview: → Sample to 10K-20K rows ELSE IF rows > 100K AND user wants time trend: → Aggregate to lower frequency (hourly→daily, daily→weekly) ELSE IF categories > 50: → Keep top 20, group rest as "Other" ELSE IF file_size > 100MB: → Select relevant columns only ELSE: → Use full dataset ``` ## Critical Reminders 1. **These are EDGE CASES** - Standard workflows handle 95% of requests 2. **Simplify first** - Always try standard workflow before complex transformations 3. **Document transformations** - Tell user what was changed (sampled, aggregated, etc.) 4. **Preserve original data** - Save transformed data as new file 5. **Test standard workflow first** - Only use edge case patterns when standard fails ## When NOT to Use This Reference **DO NOT consult this file for:** - Standard data upload → chart creation (use SKILL.md Primary Workflow) - Single file, reasonable size (<50MB, <100K rows) - Clean data with clear types - User requests standard chart types (bar, line, scatter, pie) - Data already in good structure **These scenarios use SKILL.md workflows without reading this reference.** -
online-resources.md 9.3 KB
# Online Documentation Resources **Progressive disclosure strategy:** Fetch these URLs only when user's request requires features beyond basic templates. ## When to Fetch Documentation **DO NOT fetch preemptively.** Only fetch when: - User requests feature not in skill templates (e.g., "add brush selection", "use conditional formatting") - Error occurs that requires spec validation - User asks "how do I..." for advanced feature - Need to understand parameter syntax for complex interaction **DO NOT fetch for:** - Basic charts covered in templates (bar, line, scatter, pie, heatmap, area) - Standard customizations in customization.md (colors, tooltips, axis formatting) - Data structure questions (use analyze_data.py instead) ## Core Reference Pages ### Spec Structure **URL:** https://vega.github.io/vega-lite/docs/spec.html **When to fetch:** - User asks about overall spec structure - Need to understand top-level properties - Validating spec completeness **Contains:** - Complete spec schema - Required vs optional properties - Top-level configuration options ### Mark Types **URL:** https://vega.github.io/vega-lite/docs/mark.html **When to fetch:** - User requests mark type not in templates - Need mark-specific properties (cornerRadius, interpolate, etc.) - Understanding mark styling options **Contains:** - All mark types (bar, line, point, area, rect, text, tick, circle, square, rule, geoshape, boxplot, errorbar, errorband) - Mark property reference - Mark-specific configuration ### Encoding Channels **URL:** https://vega.github.io/vega-lite/docs/encoding.html **When to fetch:** - User wants encoding channel not used in templates (angle, radius, shape, strokeWidth, etc.) - Complex multi-field encoding needed - Understanding encoding precedence **Contains:** - All encoding channels (x, y, color, size, shape, opacity, etc.) - Channel-specific properties - Type compatibility rules ## Interactive Features ### Selections (Parameters) **URL:** https://vega.github.io/vega-lite/docs/selection.html **When to fetch:** - User says "interactive", "clickable", "filter by clicking" - Need brush, interval, or point selection - Linking multiple views with shared selection **Contains:** - Selection types (point, interval, brush) - Selection configuration - Cross-view selection binding - Selection predicates ### Conditions **URL:** https://vega.github.io/vega-lite/docs/condition.html **When to fetch:** - User wants conditional styling ("color positive values green") - Selection-based styling needed - Data-driven visual encoding **Contains:** - Conditional encoding syntax - Test expressions - Selection-based conditions - Value-based conditions ### Tooltips **URL:** https://vega.github.io/vega-lite/docs/tooltip.html **When to fetch:** - User wants custom tooltip formatting beyond examples in customization.md - Need to disable/customize default tooltips - Multi-field tooltip with complex formatting **Contains:** - Tooltip encoding options - Formatting specifications - Disabling tooltips - Custom tooltip content ## Data Transformations ### Transform Overview **URL:** https://vega.github.io/vega-lite/docs/transform.html **When to fetch:** - User needs data transformation not in basic templates - Request for filtering, aggregation, calculation, binning - Need to understand transform pipeline **Contains:** - Transform types overview - Transform ordering - Common patterns ### Aggregate **URL:** https://vega.github.io/vega-lite/docs/aggregate.html **When to fetch:** - User wants "sum by category", "average per month", etc. - Grouping and aggregation needed - Statistical operations required **Contains:** - Aggregate operations (count, sum, mean, median, min, max, etc.) - Grouping syntax - Multiple aggregations ### Filter **URL:** https://vega.github.io/vega-lite/docs/filter.html **When to fetch:** - User wants to "show only", "exclude", "filter data" - Predicate expressions needed - Time-based filtering **Contains:** - Filter predicate syntax - Comparison operators - Logical operators (and, or, not) - Field predicates ### Calculate **URL:** https://vega.github.io/vega-lite/docs/calculate.html **When to fetch:** - User needs computed fields - Mathematical operations on existing fields - Derived values **Contains:** - Expression syntax - Available functions - Field references ### Bin **URL:** https://vega.github.io/vega-lite/docs/bin.html **When to fetch:** - User wants histogram - Need to create value ranges - Binning continuous data **Contains:** - Bin parameters (maxbins, step, extent) - Binning strategies - Custom bin specification ## Layout and Composition ### Faceting **URL:** https://vega.github.io/vega-lite/docs/facet.html **When to fetch:** - User wants "small multiples", "one chart per category" - Trellis plots needed - Grid layouts of charts **Contains:** - Facet encoding - Row and column facets - Facet configuration ### Layer **URL:** https://vega.github.io/vega-lite/docs/layer.html **When to fetch:** - User wants multiple marks on same chart (e.g., line + points) - Overlay visualizations needed - Combining different mark types **Contains:** - Layer specification - Shared encodings - Layer-specific encodings ### Concat **URL:** https://vega.github.io/vega-lite/docs/concat.html **When to fetch:** - User wants multiple independent charts side-by-side - Dashboard-style layouts - Horizontal/vertical concatenation **Contains:** - Concat specification - Horizontal and vertical concat - Flexible composition ### Repeat **URL:** https://vega.github.io/vega-lite/docs/repeat.html **When to fetch:** - User wants same chart template for multiple fields - Scatterplot matrix (SPLOM) - Repeated specifications **Contains:** - Repeat specification - Row and column repeat - Field substitution ## Styling and Configuration ### Scale **URL:** https://vega.github.io/vega-lite/docs/scale.html **When to fetch:** - User needs custom scale configuration beyond color schemes - Domain/range customization - Scale type questions (linear, log, sqrt, etc.) **Contains:** - Scale types - Domain and range - Scale properties (clamp, padding, nice, etc.) - Color schemes ### Axis **URL:** https://vega.github.io/vega-lite/docs/axis.html **When to fetch:** - User needs axis customization beyond format strings - Custom tick placement - Axis styling details **Contains:** - Axis properties - Tick configuration - Grid lines - Axis orientation ### Legend **URL:** https://vega.github.io/vega-lite/docs/legend.html **When to fetch:** - User needs legend customization beyond position - Custom legend formatting - Legend styling **Contains:** - Legend properties - Symbol configuration - Label formatting - Legend layout ### Title **URL:** https://vega.github.io/vega-lite/docs/title.html **When to fetch:** - User needs complex title configuration - Subtitle, anchor positioning - Title styling details **Contains:** - Title properties - Subtitle support - Positioning options - Text styling ## Time Series Specific ### Time Unit **URL:** https://vega.github.io/vega-lite/docs/timeunit.html **When to fetch:** - User has temporal data needing aggregation by time unit - "Group by month", "show by year" requests - Time-based binning **Contains:** - Time unit types (year, quarter, month, week, day, hour, etc.) - Time unit transformations - Temporal binning ## Examples Gallery ### Example Gallery **URL:** https://vega.github.io/vega-lite/examples/ **When to fetch:** - User's request matches complex pattern not in templates - Need inspiration for advanced visualization - Looking for specific example type **Contains:** - Categorized examples - Interactive specs - Copy-paste ready code **Browse by category:** - Single view: https://vega.github.io/vega-lite/examples/#single-view-plots - Composite views: https://vega.github.io/vega-lite/examples/#composite-marks - Interactive: https://vega.github.io/vega-lite/examples/#interactive - Geo: https://vega.github.io/vega-lite/examples/#geographic ## Fetching Strategy **Step 1: Identify need** ``` IF user_request requires [feature]: IDENTIFY most specific documentation page for [feature] ELSE: USE templates and existing references ``` **Step 2: Fetch documentation** ```bash # Use web_search tool to fetch specific URL # Extract relevant section from page # Apply pattern to user's data ``` **Step 3: Synthesize and apply** ``` EXTRACT relevant syntax from fetched docs MODIFY user's spec with new feature TEST in artifact PROVIDE updated link ``` ## URL Structure Pattern All Vega-Lite docs follow this pattern: ``` https://vega.github.io/vega-lite/docs/[TOPIC].html ``` **Common topics:** - mark.html, encoding.html, transform.html - [specific-transform].html (aggregate.html, filter.html, etc.) - [specific-encoding].html (color.html, size.html, etc.) - config.html (global configuration) - data.html (data loading options) **To find specific feature documentation:** 1. Check if topic exists in inventory above 2. If not, construct URL: `https://vega.github.io/vega-lite/docs/[topic-name].html` 3. Fetch and validate URL works 4. Extract relevant information ## Critical Rules 1. **Only fetch when necessary** - Don't preload documentation 2. **Be specific** - Fetch exact page needed, not entire doc site 3. **Extract and apply** - Don't just link to docs, implement the solution 4. **Cache knowledge** - If fetched once in conversation, reuse that knowledge 5. **Verify applicability** - Ensure fetched pattern works with user's data structure -
spec-builder-patterns.md 12.4 KB
# Spec Builder Patterns: Build Charts Programmatically **PURPOSE:** Build Vega-Lite specs from scratch without templates. Use when analyze_data.py suggests chart types beyond the 6 basic templates. **PRINCIPLE:** Vega-Lite specs are just JSON. Build them programmatically based on data patterns. ## Core Spec Structure Every Vega-Lite spec follows this pattern: ```json { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": "container", "height": 400, "data": {}, "mark": "MARK_TYPE", "encoding": { "CHANNEL": {"field": "FIELD_NAME", "type": "FIELD_TYPE"} } } ``` ## Building Specs by Pattern ### Pattern: Simple Bar Chart ```python def build_bar_chart(x_field, x_type, y_field, y_type): return { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": "container", "mark": "bar", "encoding": { "x": {"field": x_field, "type": x_type}, "y": {"field": y_field, "type": y_type} } } ``` ### Pattern: Grouped Bar Chart ```python def build_grouped_bar(x_field, y_field, color_field): return { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": "container", "mark": "bar", "encoding": { "x": {"field": x_field, "type": "nominal"}, "y": {"field": y_field, "type": "quantitative"}, "color": {"field": color_field, "type": "nominal"}, "xOffset": {"field": color_field} } } ``` ### Pattern: Stacked Bar Chart ```python def build_stacked_bar(x_field, y_field, color_field): return { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": "container", "mark": "bar", "encoding": { "x": {"field": x_field, "type": "nominal"}, "y": {"field": y_field, "type": "quantitative", "stack": True}, "color": {"field": color_field, "type": "nominal"} } } ``` ### Pattern: Histogram ```python def build_histogram(field): return { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": "container", "mark": "bar", "encoding": { "x": {"field": field, "type": "quantitative", "bin": True}, "y": {"aggregate": "count", "type": "quantitative"} } } ``` ### Pattern: Box Plot ```python def build_box_plot(x_field, y_field): return { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": "container", "mark": {"type": "boxplot", "extent": "min-max"}, "encoding": { "x": {"field": x_field, "type": "nominal"}, "y": {"field": y_field, "type": "quantitative"} } } ``` ### Pattern: Multi-Series Line Chart ```python def build_multi_line(x_field, y_field, color_field): return { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": "container", "mark": "line", "encoding": { "x": {"field": x_field, "type": "temporal"}, "y": {"field": y_field, "type": "quantitative"}, "color": {"field": color_field, "type": "nominal"} } } ``` ### Pattern: Stacked Area Chart ```python def build_stacked_area(x_field, y_field, color_field): return { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": "container", "mark": "area", "encoding": { "x": {"field": x_field, "type": "temporal"}, "y": {"field": y_field, "type": "quantitative", "stack": True}, "color": {"field": color_field, "type": "nominal"} } } ``` ### Pattern: Strip Plot (1D Scatter) ```python def build_strip_plot(x_field, y_field): return { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": "container", "mark": {"type": "tick", "thickness": 2}, "encoding": { "x": {"field": x_field, "type": "nominal"}, "y": {"field": y_field, "type": "quantitative"} } } ``` ### Pattern: Bubble Chart ```python def build_bubble_chart(x_field, y_field, size_field, color_field=None): encoding = { "x": {"field": x_field, "type": "quantitative"}, "y": {"field": y_field, "type": "quantitative"}, "size": {"field": size_field, "type": "quantitative"} } if color_field: encoding["color"] = {"field": color_field, "type": "nominal"} return { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": "container", "mark": "circle", "encoding": encoding } ``` ### Pattern: Error Bars ```python def build_error_bars(x_field, y_field, y_error_field): return { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": "container", "layer": [ { "mark": {"type": "errorbar", "extent": "stdev"}, "encoding": { "x": {"field": x_field, "type": "nominal"}, "y": {"field": y_field, "type": "quantitative"} } }, { "mark": {"type": "point", "filled": True}, "encoding": { "x": {"field": x_field, "type": "nominal"}, "y": {"field": y_field, "type": "quantitative", "aggregate": "mean"} } } ] } ``` ### Pattern: Normalized Stacked Bar ```python def build_normalized_bar(x_field, y_field, color_field): return { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": "container", "mark": "bar", "encoding": { "x": {"field": x_field, "type": "nominal"}, "y": { "field": y_field, "type": "quantitative", "stack": "normalize", "axis": {"format": ".0%"} }, "color": {"field": color_field, "type": "nominal"} } } ``` ## Data Pattern → Chart Builder Decision Tree ```python def select_chart_builder(fields): """Select appropriate chart builders based on data structure.""" quant = [f for f in fields if f["type"] == "quantitative"] temp = [f for f in fields if f["type"] == "temporal"] nom = [f for f in fields if f["type"] == "nominal"] builders = [] # Distribution patterns if len(quant) == 1 and len(nom) == 0: builders.append(("Histogram", build_histogram, [quant[0]["name"]])) # Comparison patterns if len(nom) >= 1 and len(quant) >= 1: builders.append(("Bar Chart", build_bar_chart, [nom[0]["name"], "nominal", quant[0]["name"], "quantitative"])) if len(nom) >= 2: builders.append(("Grouped Bar", build_grouped_bar, [nom[0]["name"], quant[0]["name"], nom[1]["name"]])) builders.append(("Stacked Bar", build_stacked_bar, [nom[0]["name"], quant[0]["name"], nom[1]["name"]])) # Distribution comparison if len(nom) >= 1 and len(quant) >= 1: builders.append(("Box Plot", build_box_plot, [nom[0]["name"], quant[0]["name"]])) builders.append(("Strip Plot", build_strip_plot, [nom[0]["name"], quant[0]["name"]])) # Time series patterns if len(temp) >= 1 and len(quant) >= 1: if len(nom) >= 1: builders.append(("Multi-Line", build_multi_line, [temp[0]["name"], quant[0]["name"], nom[0]["name"]])) builders.append(("Stacked Area", build_stacked_area, [temp[0]["name"], quant[0]["name"], nom[0]["name"]])) # Correlation patterns if len(quant) >= 2: if len(quant) >= 3: builders.append(("Bubble Chart", build_bubble_chart, [quant[0]["name"], quant[1]["name"], quant[2]["name"]])) if len(nom) >= 1: builders.append(("Colored Scatter", build_bubble_chart, [quant[0]["name"], quant[1]["name"], None, nom[0]["name"]])) # Proportion patterns if len(nom) >= 2 and len(quant) >= 1: builders.append(("Normalized Bar", build_normalized_bar, [nom[0]["name"], quant[0]["name"], nom[1]["name"]])) return builders ``` ## Usage in Workflow **Instead of loading templates:** ```python # Old way (template-constrained) with open('templates/bar.json') as f: spec = json.load(f) # New way (pattern-driven) spec = build_bar_chart( x_field="category", x_type="nominal", y_field="value", y_type="quantitative" ) ``` **Build multiple chart variations:** ```python import json # Get data analysis fields = analyze_result["fields"] # Select chart builders builders = select_chart_builder(fields) # Generate specs chart_objects = [] for name, builder_func, args in builders: spec = builder_func(*args) chart_objects.append({ "type": name, "reason": f"Generated from pattern: {name.lower()}", "spec": spec }) # Now have 8-12 chart options instead of just 3-5 ``` ## Advanced Patterns ### Layered Charts (Multiple Marks) ```python def build_line_with_points(x_field, y_field): return { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "width": "container", "layer": [ { "mark": "line", "encoding": { "x": {"field": x_field, "type": "temporal"}, "y": {"field": y_field, "type": "quantitative"} } }, { "mark": {"type": "point", "filled": True, "size": 50}, "encoding": { "x": {"field": x_field, "type": "temporal"}, "y": {"field": y_field, "type": "quantitative"} } } ] } ``` ### Faceted Charts (Small Multiples) ```python def build_faceted_bar(x_field, y_field, facet_field): return { "$schema": "https://vega.github.io/schema/vega-lite/v5.json", "facet": {"field": facet_field, "type": "nominal"}, "spec": { "width": 200, "mark": "bar", "encoding": { "x": {"field": x_field, "type": "nominal"}, "y": {"field": y_field, "type": "quantitative"} } } } ``` ## Exploring Vega-Lite Examples Gallery **When user requests uncommon chart type:** 1. Search examples gallery: https://vega.github.io/vega-lite/examples/ 2. Find relevant example 3. Fetch the spec JSON 4. Adapt to user's data structure **Example process:** ```python # User requests: "violin plot" # 1. Not in basic templates # 2. Check if in spec-builder-patterns.md → No # 3. Fetch from examples gallery: # Use web_search to find: # https://vega.github.io/vega-lite/examples/[violin-plot-example].html # Extract spec structure # Adapt field names to user's data # Generate spec ``` ## Chart Type Categories from Vega-Lite **Bar Charts (10+ variations):** - Simple, grouped, stacked, normalized, horizontal, ranged **Line Charts (8+ variations):** - Simple, multi-series, stepped, monotone, area, trail **Scatter & Strip (6+ variations):** - Scatter, bubble, strip, jitter, connected scatter **Distribution (6+ variations):** - Histogram, box plot, violin plot, density plot, QQ plot **Part-to-Whole (5+ variations):** - Pie, donut, stacked bar, normalized bar, treemap **Two-Dimensional (4+ variations):** - Heatmap, density heatmap, hexbin, rect **Advanced (10+ variations):** - Error bars/bands, box plots, waterfall, parallel coordinates, sankey ## Critical Rules 1. **Don't be template-constrained** - Build specs programmatically 2. **Suggest 8-12 chart options** - Not just 3-5 3. **Use builders for variations** - Grouped, stacked, normalized 4. **Reference examples gallery** - For uncommon chart types 5. **Explain pattern reasoning** - Why this chart fits the data ## Integration with analyze_data.py **Enhance suggestions beyond basic templates:** ```python # In analyze_data.py, expand suggestions: suggestions = [] # Always check for: - Histograms (1 quantitative) - Box plots (1 nominal + 1 quantitative) - Grouped/stacked variations (2+ nominal + 1 quantitative) - Multi-series (1 temporal + 1 quantitative + 1 nominal) - Bubble charts (3+ quantitative) - Strip plots (1 nominal + 1 quantitative) - Normalized bars (2+ nominal + 1 quantitative) ``` -
vega-lite-examples-inventory.md 5.6 KB
# Vega-Lite Examples Inventory This is a reference inventory of chart types from the official Vega-Lite examples gallery. Use this to identify appropriate chart patterns when analyze_data.py suggestions are insufficient or when user requests specific uncommon chart types. ## Bar Charts (16 examples) - Simple Bar Chart - Aggregate Bar Chart (Sorted) - Grouped Bar Chart - Stacked Bar Chart (standard, rounded corners, horizontal) - Normalized (Percentage) Stacked Bar Chart - Gantt Chart (Ranged Bar Marks) - Layered Bar Chart - Diverging Stacked Bar Chart (Population Pyramid, with Neutral Parts) - Bar Chart with Labels/Overlays - Bar Chart with Negative Values - Heat Lane Chart ## Histograms & Distributions (11 examples) - Histogram (standard, binned, log-scaled, non-linear) - Relative Frequency Histogram - Density Plot - Stacked Density Estimates - 2D Histogram Scatterplot/Heatmap - Cumulative Frequency Distribution - Wilkinson Dot Plot - Isotype Dot Plot (standard, with Emoji) ## Scatter & Strip Plots (10 examples) - Scatterplot (standard, colored, with null values, filled circles) - 1D Strip Plot / Strip Plot - Bubble Plot (standard, Gapminder, Natural Disasters) - Scatter Plot with Text Marks - Image-based Scatter Plot - Strip plot with custom axis tick labels - Dot Plot with Jittering ## Line Charts (15 examples) - Line Chart (standard, with point markers, stroked markers) - Multi Series Line Chart (standard, with repeat, with halo stroke) - Slope Graph - Step Chart - Line Chart with Monotone Interpolation - Connected Scatterplot (custom paths) - Bump Chart - Line Chart with Varying Size (trail mark) - Comet Chart - Line Chart with Markers and Invalid Values - Line Charts Showing Ranks Over Time - Sine/Cosine Curves (sequence generator) - Line chart with varying stroke dash ## Area Charts & Streamgraphs (6 examples) - Area Chart (standard, with gradient, with overlaying lines) - Stacked Area Chart - Normalized Stacked Area Chart - Streamgraph - Horizon Graph ## Table-based Plots (7 examples) - Table Heatmap - Annual Weather Heatmap - 2D Histogram Heatmap - Table Bubble Plot (Github Punch Card) - Heatmap with Labels - Lasagna Plot (Dense Time-Series Heatmap) - Mosaic Chart with Labels - Wind Vector Map ## Circular Plots (6 examples) - Pie Chart (standard, with percentage tooltip, with labels) - Donut Chart - Radial Plot - Pyramid Pie Chart ## Advanced Calculations (14 examples) - Calculate Difference from Average/Annual Average - Calculate Residuals - Waterfall Chart - Filtering Top-K Items / Top-K Plot with "Others" - Lookup transform to combine data - Parallel Coordinate Plot - Bar Chart Showing Argmax Value - Layering Averages over Raw Values - Layering Rolling Averages over Raw Values - Quantile-Quantile Plot (QQ Plot) - Linear/Loess Regression - Window transform for imputation ## Error Bars & Error Bands (4 examples) - Error Bars (confidence interval, standard deviation) - Line Chart with Confidence Interval Band - Scatterplot with Mean and Standard Deviation Overlay ## Box Plots (3 examples) - Box Plot with Min/Max Whiskers - Tukey Box Plot (1.5 IQR) - Box Plot with Pre-Calculated Summaries ## Labeling & Annotation (10 examples) - Bar Chart with Labels (standard, with emojis) - Layering text over heatmap - Bar Chart Highlighting Values beyond Threshold - Mean overlay over precipitation chart - Histogram with Global Mean Overlay - Line Chart with Highlighted Rectangles - Distributions and Medians of Likert Scale Ratings - Comparative Likert Scale Ratings ## Other Layered Plots (6 examples) - Candlestick Chart - Ranged Dot Plot - Bullet Chart - Layered Plot with Dual-Axis - Weekly Weather Plot - Wheat and Wages Example ## Faceting (9 examples) - Trellis Bar/Stacked Bar Chart - Trellis Scatter Plot (wrapped, Anscombe's Quartet) - Trellis Histograms - Becker's Barley Trellis Plot - Trellis Area (standard, annual temperatures) - Faceted Density Plot - Compact Trellis Grid of Bar Charts ## Repeat & Concatenation (9 examples) - Repeat and Layer for Different Measures - Vertical/Horizontal Concatenation - Interactive Scatterplot Matrix - Marginal Histograms - Discretizing scales - Nested View Concatenation - Population Pyramid ## Geographic/Maps (7 examples) - Choropleth of Unemployment Rate - One Dot per Zipcode/Airport in US - Rules Connecting Airports - Three Choropleths (disjoint data) - US State Capitals on Map - Line between Airports - Income by State, Faceted - London Tube Lines - Projection explorer - Earthquakes ## Interactive Charts (30+ examples) - Bar Chart with Highlighting/Selection - Histogram with Full-Height Hover Targets - Interactive Legend - Scatterplot with External Links and Tooltips - Rectangular Brush / Area Chart with Brush - Paintbrush Highlight - Scatterplot Pan & Zoom - Query Widgets - Interactive Average - Multi Series with Interactive Highlight (line, point, labels, tooltip) - Isotype Grid - Brushing Scatter to show table - Selectable Heatmap - Bar Chart with Minimap - Interactive Index Chart - Focus + Context (Smooth Histogram Zooming) - Dynamic Color Legend - Search Input - Change zorder on hover - Overview and Detail - Crossfilter (Filter/Highlight) - Interactive Scatterplot Matrix - Interactive Dashboard with Cross Highlight - Seattle Weather Exploration - Connections among Airports - Interactive scatter of global health statistics ## Usage Pattern When user requests uncommon chart or analyze_data.py insufficient: 1. Search this inventory for relevant pattern 2. If found, construct spec using pattern from spec-builder-patterns.md or by fetching Vega-Lite example 3. If not found in inventory, fall back to programmatic builder Priority order: spec-builder-patterns.md > this inventory > web_fetch Vega-Lite docs
-
-
scripts
-
analyze_data.py 9.9 KB
#!/usr/bin/env python3 """ Analyzes tabular data and suggests appropriate Vega-Lite chart types. Usage: python analyze_data.py <filepath> Output: JSON with data summary and chart recommendations """ import json import sys import numpy as np import pandas as pd class NumpyEncoder(json.JSONEncoder): """JSON encoder that handles numpy types and NaN values.""" def default(self, obj): if isinstance(obj, np.integer): return int(obj) elif isinstance(obj, np.floating): if np.isnan(obj) or np.isinf(obj): return None return float(obj) elif isinstance(obj, np.ndarray): return obj.tolist() elif isinstance(obj, (np.bool_, bool)): return bool(obj) elif pd.isna(obj): return None return super().default(obj) def infer_field_type(series): """Infer Vega-Lite field type from pandas series.""" if pd.api.types.is_numeric_dtype(series): # Check if it's actually ordinal (few unique integer values) if pd.api.types.is_integer_dtype(series) and series.nunique() < 10: return "ordinal" return "quantitative" elif pd.api.types.is_datetime64_any_dtype(series): return "temporal" else: # Nominal vs ordinal heuristic unique_count = series.nunique() if unique_count < 20 and unique_count < len(series) / 2: return "nominal" return "nominal" # Default to nominal for text def analyze_field(series): """Analyze a single field and return its characteristics.""" field_info = { "name": series.name, "type": infer_field_type(series), "unique_count": series.nunique(), "null_count": series.isnull().sum(), "sample_values": series.dropna().head(5).tolist() } if pd.api.types.is_numeric_dtype(series): field_info["stats"] = { "min": float(series.min()), "max": float(series.max()), "mean": float(series.mean()), "median": float(series.median()) } return field_info def suggest_charts(fields): """Suggest chart types based on field characteristics - EXPANDED.""" suggestions = [] # Count field types quantitative = [f for f in fields if f["type"] == "quantitative"] temporal = [f for f in fields if f["type"] == "temporal"] nominal = [f for f in fields if f["type"] == "nominal"] # DISTRIBUTION PATTERNS # Pattern: Single quantitative = histogram if len(quantitative) >= 1: suggestions.append({ "type": "histogram", "reason": "Distribution of values", "encoding": { "x": quantitative[0]["name"], "bin": True, "y": "count" }, "priority": "medium" }) # COMPARISON PATTERNS # Pattern: Nominal + Quantitative = bar chart + variations if len(quantitative) >= 1 and len(nominal) >= 1: suggestions.append({ "type": "bar", "reason": "Categorical comparison", "encoding": { "x": nominal[0]["name"], "y": quantitative[0]["name"] }, "priority": "high" }) # Box plot for distribution comparison suggestions.append({ "type": "boxplot", "reason": "Distribution comparison across categories", "encoding": { "x": nominal[0]["name"], "y": quantitative[0]["name"] }, "priority": "medium" }) # Strip plot alternative suggestions.append({ "type": "stripplot", "reason": "Individual data points by category", "encoding": { "x": nominal[0]["name"], "y": quantitative[0]["name"] }, "priority": "low" }) # Grouped bar if 2+ nominal if len(nominal) >= 2: suggestions.append({ "type": "grouped-bar", "reason": "Multi-series comparison", "encoding": { "x": nominal[0]["name"], "y": quantitative[0]["name"], "color": nominal[1]["name"] }, "priority": "high" }) suggestions.append({ "type": "stacked-bar", "reason": "Part-to-whole by category", "encoding": { "x": nominal[0]["name"], "y": quantitative[0]["name"], "color": nominal[1]["name"] }, "priority": "medium" }) suggestions.append({ "type": "normalized-bar", "reason": "Percentage composition by category", "encoding": { "x": nominal[0]["name"], "y": quantitative[0]["name"], "color": nominal[1]["name"] }, "priority": "medium" }) # TIME SERIES PATTERNS if len(temporal) >= 1 and len(quantitative) >= 1: suggestions.append({ "type": "line", "reason": "Time series visualization", "encoding": { "x": temporal[0]["name"], "y": quantitative[0]["name"] }, "priority": "high" }) suggestions.append({ "type": "area", "reason": "Time series with magnitude emphasis", "encoding": { "x": temporal[0]["name"], "y": quantitative[0]["name"] }, "priority": "medium" }) # Multi-series if nominal available if len(nominal) >= 1: suggestions.append({ "type": "multi-line", "reason": "Multiple time series comparison", "encoding": { "x": temporal[0]["name"], "y": quantitative[0]["name"], "color": nominal[0]["name"] }, "priority": "high" }) suggestions.append({ "type": "stacked-area", "reason": "Cumulative time series", "encoding": { "x": temporal[0]["name"], "y": quantitative[0]["name"], "color": nominal[0]["name"] }, "priority": "medium" }) # CORRELATION PATTERNS if len(quantitative) >= 2: suggestions.append({ "type": "scatter", "reason": "Correlation analysis", "encoding": { "x": quantitative[0]["name"], "y": quantitative[1]["name"], "color": nominal[0]["name"] if nominal else None }, "priority": "high" }) # Bubble chart if 3+ quantitative if len(quantitative) >= 3: suggestions.append({ "type": "bubble", "reason": "Three-variable correlation", "encoding": { "x": quantitative[0]["name"], "y": quantitative[1]["name"], "size": quantitative[2]["name"], "color": nominal[0]["name"] if nominal else None }, "priority": "medium" }) # PART-TO-WHOLE PATTERNS if len(nominal) >= 1 and len(quantitative) >= 1: # Only suggest pie if reasonable category count if nominal[0].get("unique_count", 999) < 7: suggestions.append({ "type": "pie", "reason": "Part-to-whole relationship", "encoding": { "theta": quantitative[0]["name"], "color": nominal[0]["name"] }, "priority": "low" }) # TWO-DIMENSIONAL CATEGORICAL if len(nominal) >= 2 and len(quantitative) >= 1: suggestions.append({ "type": "heatmap", "reason": "Two-dimensional categorical comparison", "encoding": { "x": nominal[0]["name"], "y": nominal[1]["name"], "color": quantitative[0]["name"] }, "priority": "medium" }) # Sort by priority priority_order = {"high": 0, "medium": 1, "low": 2} suggestions.sort(key=lambda x: priority_order.get(x["priority"], 3)) return suggestions def main(): if len(sys.argv) < 2: print(json.dumps({"error": "Usage: python analyze_data.py <filepath>"}, cls=NumpyEncoder)) sys.exit(1) filepath = sys.argv[1] try: # Load data if filepath.endswith('.csv'): df = pd.read_csv(filepath) elif filepath.endswith('.json'): df = pd.read_json(filepath) elif filepath.endswith(('.xlsx', '.xls')): df = pd.read_excel(filepath) else: print(json.dumps({"error": f"Unsupported file type: {filepath}"}, cls=NumpyEncoder)) sys.exit(1) # Analyze fields fields = [analyze_field(df[col]) for col in df.columns] # Get chart suggestions suggestions = suggest_charts(fields) # Prepare output result = { "row_count": len(df), "column_count": len(df.columns), "fields": fields, "suggested_charts": suggestions, "sample_data": df.head(10).to_dict(orient='records') } print(json.dumps(result, indent=2, cls=NumpyEncoder)) except Exception as e: print(json.dumps({"error": str(e)}, cls=NumpyEncoder)) sys.exit(1) sys.exit(1) if __name__ == "__main__": main() -
prepare_data.py 5.1 KB
#!/usr/bin/env python3 """ Prepares data for external reference by Vega-Lite specs. Creates optimized data files and returns reference information. Usage: python prepare_data.py <input_filepath> [--output-format json|csv] Output: JSON with data file path, format, and reference metadata """ import argparse import json import sys from pathlib import Path import numpy as np import pandas as pd class NumpyEncoder(json.JSONEncoder): """JSON encoder that handles numpy types.""" def default(self, obj): if isinstance(obj, np.integer): return int(obj) elif isinstance(obj, np.floating): if np.isnan(obj) or np.isinf(obj): return None return float(obj) elif isinstance(obj, np.ndarray): return obj.tolist() elif isinstance(obj, (np.bool_, bool)): return bool(obj) elif pd.isna(obj): return None return super().default(obj) def optimize_dtypes(df): """Optimize DataFrame dtypes for efficient serialization.""" for col in df.columns: # Convert object columns with few unique values to category if df[col].dtype == 'object': unique_ratio = df[col].nunique() / len(df) if unique_ratio < 0.5: # More than 50% repetition df[col] = df[col].astype('category') # Downcast numeric types elif pd.api.types.is_integer_dtype(df[col]): df[col] = pd.to_numeric(df[col], downcast='integer') elif pd.api.types.is_float_dtype(df[col]): df[col] = pd.to_numeric(df[col], downcast='float') return df def prepare_data(input_path, output_format='json', output_dir='/mnt/user-data/outputs'): """ Prepare data for external reference. Returns: dict with: - data_path: Path to prepared data file - format: 'json' or 'csv' - rows: Number of rows - columns: List of column names - size_bytes: File size - reference_pattern: How to reference in spec """ # Load data if input_path.endswith('.csv'): df = pd.read_csv(input_path) elif input_path.endswith('.json'): df = pd.read_json(input_path) elif input_path.endswith(('.xlsx', '.xls')): df = pd.read_excel(input_path) else: raise ValueError(f"Unsupported file type: {input_path}") # Optimize dtypes df = optimize_dtypes(df) # Determine output filename input_name = Path(input_path).stem output_path = Path(output_dir) / f"{input_name}_data.{output_format}" # Save optimized data if output_format == 'json': # Use records format for Vega-Lite compatibility data_records = df.to_dict(orient='records') with open(output_path, 'w') as f: json.dump(data_records, f, cls=NumpyEncoder) elif output_format == 'csv': df.to_csv(output_path, index=False) else: raise ValueError(f"Unsupported output format: {output_format}") # Get file size size_bytes = output_path.stat().st_size # Prepare result result = { "data_path": str(output_path), "format": output_format, "rows": len(df), "columns": df.columns.tolist(), "size_bytes": size_bytes, "size_human": f"{size_bytes / 1024:.1f} KB" if size_bytes > 1024 else f"{size_bytes} bytes", "reference_patterns": { "vega_lite_url": { "description": "Reference via URL (for web-hosted data)", "spec": {"data": {"url": f"computer://{output_path}"}} }, "react_fetch": { "description": "Load in React component", "code": f""" // Load data separately in React component const [data, setData] = useState(null); useEffect(() => {{ fetch('computer://{output_path}') .then(r => r.json()) .then(d => setData(d)); }}, []); // Use in spec const spec = {{ ... data: {{values: data}}, ... }}; """ }, "angular_http": { "description": "Load via Angular HttpClient", "code": f""" // In component this.http.get<any[]>('assets/{output_path.name}') .subscribe(data => {{ this.spec.data.values = data; this.renderChart(); }}); """ } } } return result def main(): parser = argparse.ArgumentParser(description='Prepare data for Vega-Lite external reference') parser.add_argument('input_path', help='Path to input data file') parser.add_argument('--output-format', choices=['json', 'csv'], default='json', help='Output format (default: json)') parser.add_argument('--output-dir', default='/mnt/user-data/outputs', help='Output directory (default: /mnt/user-data/outputs)') args = parser.parse_args() try: result = prepare_data(args.input_path, args.output_format, args.output_dir) print(json.dumps(result, indent=2, cls=NumpyEncoder)) except Exception as e: print(json.dumps({"error": str(e)}, cls=NumpyEncoder)) sys.exit(1) if __name__ == "__main__": main()
-
-
README.md 461 B
# charting-vega-lite Create interactive data visualizations using Vega-Lite declarative JSON grammar. Supports 20+ chart types (bar, line, scatter, histogram, boxplot, grouped/stacked variations, etc.) via templates and programmatic builders. Use when users upload data for charting, request specific chart types, or mention visualizations. Produces portable JSON specs with inline data islands that work in Claude artifacts and can be adapted for production. -
SKILL.md 8.2 KB
--- name: charting-vega-lite description: Create interactive data visualizations using Vega-Lite declarative JSON grammar. Supports 20+ chart types (bar, line, scatter, histogram, boxplot, grouped/stacked variations, etc.) via templates and programmatic builders. Use when users upload data for charting, request specific chart types, or mention visualizations. Produces portable JSON specs with inline data islands that work in Claude artifacts and can be adapted for production. metadata: version: 0.1.0 --- ## Overview This skill creates interactive Vega-Lite visualizations from uploaded data. The workflow: 1. Analyze data structure and context 2. Select 5-10 meaningful chart types based on what the data represents 3. Build chart specifications programmatically 4. Generate React artifact with embedded visualizations ## Critical Technical Constraint: Inline Data Island **Claude artifacts cannot use fetch() for computer:// URLs.** All data must be embedded as an inline JavaScript constant: ```javascript const DATA = [ /* embedded data array */ ]; // Later in chart specs: spec.data = { values: DATA }; ``` **DO NOT:** - Use fetch() to load external files - Reference external data URLs - Create separate data files This is the only pattern that works in Claude's artifact environment. ## Primary Workflow: Data Upload → Chart Explorer Execute this sequence when user uploads data without specifying chart type: ### 1. Analyze Data Structure ```bash python /mnt/skills/user/charting-vega-lite/scripts/analyze_data.py /mnt/user-data/uploads/<filename> ``` **Extract from output:** - `fields[]` (with types and statistics) - `suggested_charts[]` (suggested chart types with encodings) - `sample_data` (first 10 rows for understanding context) **If script fails:** Use manual pandas analysis ```python import pandas as pd df = pd.read_csv('/mnt/user-data/uploads/<filename>') # Classify: numeric→quantitative, datetime→temporal, <20 unique→nominal ``` ### 2. Understand Data Context **Read sample data and column names to infer what the data represents:** - **Biomedical data?** → Biomarkers, patient outcomes, clinical relevance - **Financial data?** → Trends, comparisons, performance metrics - **Sensor data?** → Temporal patterns, anomalies, correlations - **E-commerce?** → Sales trends, product comparisons, conversions **Ask:** What questions would someone analyzing this data want answered? Examples: - Assay data: Which biomarkers strongest? Patterns across samples? Variability? - Financial: What are trends? How volatile? Seasonal patterns? - IoT: Temporal patterns? Anomalies? Sensor correlations? ### 3. Select Meaningful Charts (5-10 suggestions) **Filter analyze_data.py suggestions based on context and readability:** **Apply readability filters:** - Pie chart with >7 categories → Skip (unreadable) - Heatmap with >50 categories per axis → Aggregate first - Multi-line with >10 series → Consider faceting **Prioritize charts that answer domain questions:** - Comparison needs → Bar, box plot, grouped bar - Distribution analysis → Histogram, box plot - Pattern recognition → Heatmap, scatter - Temporal trends → Line, area - Part-to-whole → Stacked bar (pie only if <7 categories) **Don't suggest charts just because data types match - choose charts that reveal insights.** ### 4. Generate Chart Specs **Build specs programmatically using analyze_data.py encodings:** For each suggested chart type, construct spec using: - Templates from `assets/templates/` for basic types (bar, line, scatter, pie, heatmap, area) - Builder patterns from `references/spec-builder-patterns.md` for variations (histogram, boxplot, grouped-bar, etc.) - Vega-Lite examples from `references/vega-lite-examples-inventory.md` for uncommon types Structure each chart as: ```python {"type": "Chart Name", "reason": "Why this chart", "spec": {/* vega-lite spec */}} ``` ### 5. Create Artifact with Inline Data Island Load data, read template, replace `__DATA__` and `__CHART_SPECS__` placeholders, write using bash heredoc. ### 6. Provide Link ``` [View chart explorer](computer:///mnt/user-data/outputs/ChartExplorer.jsx) Created 7 contextually relevant charts for your data. ``` ## Secondary Workflow: Specific Chart Request When user specifies chart type (e.g., "make a bar chart"): ### 1. Analyze Data ```bash python /mnt/skills/user/charting-vega-lite/scripts/analyze_data.py /mnt/user-data/uploads/<filename> ``` ### 2. Validate Chart Fits Data **Check requirements:** - Bar: needs 1 nominal + 1 quantitative - Line: needs 1 temporal + 1 quantitative - Scatter: needs 2 quantitative - Heatmap: needs 2 nominal + 1 quantitative - Pie: needs 1 nominal + 1 quantitative + <7 categories **If data doesn't fit:** - Explain: "Bar chart needs categorical data, but all columns are numeric" - Suggest 2-3 alternatives - Use Primary Workflow to create explorer with alternatives ### 3. Generate Spec Use templates or programmatic builders based on chart type complexity. ### 4. Create Artifact Same pattern as Primary Workflow step 5, but with single chart. ## Error Prevention **Common failures:** 1. **Using fetch() in artifacts** - Solution: Always use inline data island pattern - Never create external data files 2. **Chart doesn't render** - Verify scripts load: Vega → Vega-Lite → Vega-Embed - Check data is injected: `spec.data = {values: DATA}` - Confirm field names match data columns 3. **Generic/random chart suggestions** - Solution: Consider data context and meaning - Filter suggestions for relevance and readability - Prioritize charts that answer meaningful questions ## Resources **Scripts:** - `scripts/analyze_data.py` - analyze structure, suggest 8-12 chart types **Components:** - `assets/components/ChartExplorer.jsx` - multi-chart explorer template **Templates:** - `assets/templates/*.json` - 6 basic chart templates (bar, line, scatter, pie, heatmap, area) **References - Progressive Disclosure:** Read `spec-builder-patterns.md` when building charts programmatically (histogram, boxplot, grouped/stacked bars, multi-line, etc.) Read `vega-lite-examples-inventory.md` when user requests uncommon chart type not in spec-builder-patterns Read `chart-types.md` when validating specific chart requirements or user asks "what chart should I use for..." Read `advanced-charts.md` for complete specs of specialized charts (sankey, waterfall, violin plots, complex layered compositions) Read `contextual-chart-selection.md` for extended domain examples if unfamiliar with data domain (biomedical, financial, IoT, etc.) Read `online-resources.md` to fetch Vega-Lite docs for advanced features (custom selections, transforms, conditional encoding) ## Complete Workflow Example **User uploads assay data CSV (51 assays, 74 samples)** ```bash # 1. Analyze python /mnt/skills/user/charting-vega-lite/scripts/analyze_data.py /mnt/user-data/uploads/assay_data.csv # 2. Understand context: Multi-analyte immunoassay # Questions: Which biomarkers strongest? Patterns across samples? Variability? # 3. Build contextual charts (5-7 specs) # Bar: Mean signal by assay # Heatmap: Sample × Assay # Box plot: Signal distribution by assay # Histogram: Overall signal distribution # etc. # 4. Load data and template df = pd.read_csv('/mnt/user-data/uploads/assay_data.csv') data = df.to_dict(orient='records') template = open('/mnt/skills/user/charting-vega-lite/assets/components/ChartExplorer.jsx').read() # 5. Replace placeholders and write artifact = template.replace('__DATA__', json.dumps(data)).replace('__CHART_SPECS__', json.dumps(charts)) # Use bash heredoc to avoid XML conflicts in tool parameters # 6. Provide link ``` [View chart explorer](computer:///mnt/user-data/outputs/ChartExplorer.jsx) Created 7 charts for your assay data - bar charts show biomarker signals, heatmap reveals sample patterns, box plots display variability. ## Critical Rules 1. **ALWAYS use inline data island pattern** - No fetch(), no external files 2. **Consider data context** - Choose meaningful charts based on what data represents, not just data types 3. **Filter by readability** - Avoid charts with too many categories 4. **Use bash heredoc for file creation** - Prevents XML conflicts when creating artifacts 5. **Provide links, not content** - Output token efficiency
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.