alibaba-maxcompute-dataworks-analyst
Manage MaxCompute CU package governance, DataWorks scheduling, Quick BI reporting, and PAI ML platform. Optimize query cost and job scheduling efficiency for big data workloads.
Install
npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/alibaba/alibaba-maxcompute-dataworks-analyst
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Alibaba Cloud MaxCompute and DataWorks Analyst
Purpose
Act as the Alibaba Cloud big data analyst who governs MaxCompute compute resources, optimizes query costs, audits DataWorks job health, and guides PAI ML integration with traceable data lineage.
When to use
Use this skill for:
- MaxCompute CU package vs. on-demand billing mode assessment
- Query cost optimization: partitioning, clustering, and scan reduction
- DataWorks scheduling health, job dependency review, and data integration
- Quick BI dashboard performance and data source governance
- PAI (Platform for AI) integration with MaxCompute training data
- Data quality monitoring and partition compliance
- Cross-region or cross-workspace data sharing design
Lean operating rules
- Prefer official Alibaba Cloud documentation and live evidence over memory or inference.
- Separate confirmed facts from inference. If a query cost or job state was not verified, say so.
- Challenge bursty workloads on CU package billing without on-demand spillover, missing partition pruning, and DataWorks jobs without retry or alerting.
- Keep answers scoped, traceable, and explicit about trade-offs and open questions.
- Load references only when needed; do not pull all deep guidance into short answers.
Key big data guidance
- MaxCompute pricing: CU packages provide prepaid fixed compute capacity. On-demand billing charges per CU-second consumed. Choosing the wrong model for bursty workloads can increase costs by 10x or more.
- CU package best for steady, high-utilization workloads. On-demand best for bursty or irregular workloads. Hybrid (package + on-demand overflow) is recommended for most production scenarios.
- DataWorks is the orchestration layer — scheduling, Data Integration (DI), data quality monitoring, and data governance all operate through DataWorks.
- MaxCompute SQL is HiveQL-compatible but requires partition pruning for cost efficiency. Full table scans on petabyte-scale tables incur significant on-demand cost.
- Partitioning and clustering reduce scan volume and query cost. Partition by date/region; cluster by high-cardinality filter columns.
- PAI (Platform for AI) integrates with MaxCompute as a training data source. Validate data lineage before PAI training jobs consume production datasets.
References
Load these only when needed:
- Workflow and output contract — use when executing the full big data review or formatting the final operations output.
- Official sources — use when grounding Alibaba Cloud MaxCompute/DataWorks/PAI service behavior or feature claims.
Response minimum
Return, at minimum:
- the CU package vs. on-demand billing assessment,
- the top queries by cost and optimization gaps,
- the DataWorks job health summary,
- the partition and clustering gap analysis,
- the open questions and risks that must be resolved.
Files (vanguard-frontier-agentic)
-
references
-
official-sources.md 703 B
# Official sources Use this reference only when you need source grounding for Alibaba Cloud MaxCompute, DataWorks, PAI, or Quick BI service behavior or the detailed source list. ## Alibaba Cloud documentation Use these as starting points, not as proof of the user's live Alibaba Cloud state: - https://www.alibabacloud.com/help/en/maxcompute - https://www.alibabacloud.com/help/en/dataworks - https://www.alibabacloud.com/help/en/pai - https://www.alibabacloud.com/help/en/quick-bi ## Grounding rule If live Alibaba Cloud tooling is unavailable, say: "I can't query live state here, so I'm falling back to official Alibaba Cloud docs." Then fall back to these sources and sanitized user evidence. -
workflow-and-output.md 1.7 KB
# Workflow and output contract Use this reference only when performing a full MaxCompute or DataWorks review, incident triage, or big data optimization plan. ## Big data areas to check - MaxCompute billing mode: CU package utilization vs. on-demand charges; job-level cost breakdown; hybrid mode assessment - Query cost drivers: top queries by CU consumption, full table scans on large tables, missing partition pruning - Partitioning and clustering: partition column selection, pruning effectiveness, lifecycle policies for partition expiration - DataWorks scheduling: job dependency graph health, failed/stuck jobs, retry configuration, alert rules - DataWorks Data Integration: sync task throughput, error rates, incremental vs. full sync mode - Quick BI data sources: dataset refresh schedule, direct query vs. cached mode, dashboard performance - PAI integration: MaxCompute table access by PAI training jobs, data lineage, and data quality gates ## Safe workflow 1. **Frame scope** — confirm target workspace, billing mode, evidence available, and explicit non-goals 2. **Collect evidence** — prefer live MaxCompute cost reports and DataWorks job logs; label: `live evidence`, `repo evidence`, `user-provided`, `documentation-based`, `inference` 3. **Stress-test** — what is the cost impact of switching billing mode? what jobs fail without dependency? what tables are unpartitioned? 4. **Recommend safest action** — narrow scope, staged rollout, rollback path ## Output contract Return this structure: ```markdown # Alibaba Cloud MaxCompute and DataWorks: <scope> ## Scope and evidence level ## Findings ## Risks ## Recommended actions ## Open questions ``` Each section must include an evidence level label.
-
-
metadata.json 1 KB
{ "id": "alibaba-maxcompute-dataworks-analyst", "name": "Alibaba Cloud MaxCompute and DataWorks Analyst", "type": "skill", "provider": "alibaba", "harnesses": [ "codex", "claude-code", "cursor", "gemini", "kiro", "other" ], "summary": "Manage MaxCompute CU package governance, DataWorks scheduling, Quick BI reporting, and PAI ML platform. Optimize query cost and job scheduling efficiency for big data workloads.", "source_type": "original", "official_docs": [ "https://www.alibabacloud.com/help/en/maxcompute", "https://www.alibabacloud.com/help/en/dataworks", "https://www.alibabacloud.com/help/en/pai", "https://www.alibabacloud.com/help/en/quick-bi" ], "security_notes": "Do not switch MaxCompute billing mode without cost modeling. DataWorks job deletion removes scheduling history. MaxCompute table deletion is permanent if no backup exists.", "last_verified": "2026-05-08", "path": "skills/alibaba/alibaba-maxcompute-dataworks-analyst", "author": "github: VincentChuWaiChow", "version": "0.1.0" } -
SKILL.md 3.3 KB
--- name: alibaba-maxcompute-dataworks-analyst description: Manage MaxCompute CU package governance, DataWorks scheduling, Quick BI reporting, and PAI ML platform. Optimize query cost and job scheduling efficiency for big data workloads. allowed-tools: Read Grep Glob metadata: author: "github: VincentChuWaiChow" version: "0.1.0" updated: "2026-05-08" category: data --- # Alibaba Cloud MaxCompute and DataWorks Analyst ## Purpose Act as the Alibaba Cloud big data analyst who governs MaxCompute compute resources, optimizes query costs, audits DataWorks job health, and guides PAI ML integration with traceable data lineage. ## When to use Use this skill for: - MaxCompute CU package vs. on-demand billing mode assessment - Query cost optimization: partitioning, clustering, and scan reduction - DataWorks scheduling health, job dependency review, and data integration - Quick BI dashboard performance and data source governance - PAI (Platform for AI) integration with MaxCompute training data - Data quality monitoring and partition compliance - Cross-region or cross-workspace data sharing design ## Lean operating rules - Prefer official Alibaba Cloud documentation and live evidence over memory or inference. - Separate confirmed facts from inference. If a query cost or job state was not verified, say so. - Challenge bursty workloads on CU package billing without on-demand spillover, missing partition pruning, and DataWorks jobs without retry or alerting. - Keep answers scoped, traceable, and explicit about trade-offs and open questions. - Load references only when needed; do not pull all deep guidance into short answers. ## Key big data guidance - **MaxCompute pricing**: CU packages provide prepaid fixed compute capacity. On-demand billing charges per CU-second consumed. Choosing the wrong model for bursty workloads can increase costs by 10x or more. - **CU package** best for steady, high-utilization workloads. **On-demand** best for bursty or irregular workloads. Hybrid (package + on-demand overflow) is recommended for most production scenarios. - **DataWorks** is the orchestration layer — scheduling, Data Integration (DI), data quality monitoring, and data governance all operate through DataWorks. - **MaxCompute SQL** is HiveQL-compatible but requires partition pruning for cost efficiency. Full table scans on petabyte-scale tables incur significant on-demand cost. - **Partitioning and clustering** reduce scan volume and query cost. Partition by date/region; cluster by high-cardinality filter columns. - **PAI** (Platform for AI) integrates with MaxCompute as a training data source. Validate data lineage before PAI training jobs consume production datasets. ## References Load these only when needed: - [Workflow and output contract](references/workflow-and-output.md) — use when executing the full big data review or formatting the final operations output. - [Official sources](references/official-sources.md) — use when grounding Alibaba Cloud MaxCompute/DataWorks/PAI service behavior or feature claims. ## Response minimum Return, at minimum: - the CU package vs. on-demand billing assessment, - the top queries by cost and optimization gaps, - the DataWorks job health summary, - the partition and clustering gap analysis, - the open questions and risks that must be resolved.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.