fetch-foundation-model-pricing
Fetch live per-token, per-image, and per-GPU-hour prices for foundation models across Anthropic, OpenAI, Google, AWS Bedrock, Azure OpenAI, OCI Generative AI, and Vertex AI. Supports single-model lookup and comparative multi-provider tables. Every price is labeled with source URL
Install
npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/finops/fetch-foundation-model-pricing
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.
README
Fetch Foundation Model Pricing
A FinOps skill that retrieves live public pricing for foundation models across major AI and cloud providers, returning structured tables with mandatory provenance labels and source timestamps.
Purpose
Fetch current per-token, per-image, and per-GPU-hour prices from Anthropic, OpenAI, Google (Vertex AI), AWS Bedrock, Azure OpenAI Service, and OCI Generative AI. Supports single-model lookups and side-by-side comparative tables.
Allowed tools
Read Grep Glob WebFetch
Usage
Single-model lookup: Provide a model name and deployment target (e.g., "What does Claude Sonnet 4.5 cost per million input tokens on Bedrock?"). The skill fetches the live price, labels it with source URL and ISO 8601 timestamp, and returns a single-row price table.
Comparative table: Provide two or more models or providers and a task type (e.g., "Compare GPT-4o, Claude Sonnet, and Gemini Pro on input and output token cost"). The skill fetches each price independently and builds a labeled multi-row comparison with cheapest/most-expensive summary rows.
Trust posture
Read-only. No cloud credentials, billing account IDs, or tenant data accepted. All pricing pages are public and unauthenticated. Every price value carries a provenance label (live-price, documentation-based, assumed, or excluded) with source URL and fetch timestamp.
FOCUS v1.2 column mapping is included for any cost estimate produced (BilledCost, EffectiveCost, ServiceCategory, ChargeCategory, SkuId, SkuPriceId).
See SKILL.md for the full operating protocol, pricing dimensions, and response shape.
Skill manifest
Fetch Foundation Model Pricing
Purpose
Retrieve current public pricing for foundation models across the major AI/cloud providers and return structured, provenance-labeled output. Supports two modes:
- Single-model lookup: fetch the current price for a specific model and deployment target (e.g., Claude Sonnet 4.5 on Anthropic direct, or on Bedrock).
- Comparative table: build a side-by-side price comparison across two or more models or providers for the same task type (text, image, embedding, or GPU-hour).
When to use
Use this skill when:
- The user asks "how much does model X cost per token / per image / per GPU-hour"
- The user wants to compare inference costs across two or more providers for the same model family or equivalent capability tier
- The user needs to estimate monthly AI inference spend given a volume projection
- The user wants to understand how context caching or batch pricing changes the effective cost curve
- The user wants a FOCUS-aware cost breakdown (BilledCost, EffectiveCost) for AI inference line items
Operating rules
- Fetch live prices first. Use WebFetch to retrieve prices from each provider's public pricing page before relying on any internal knowledge. AI model pricing changes frequently; stale numbers mislead.
- Label every price. Each price value must carry exactly one provenance label:
live-price— fetched from a provider's public pricing page or API within this session; include source URL and ISO 8601 timestamp.documentation-based— sourced from official documentation when a live fetch was not possible; note the documentation URL and its visible publication date.assumed— derived from an analogous model or tier when no direct published price exists; state the assumption explicitly.excluded— pricing that exists but was intentionally omitted from the output; state why.
- Include source URL and timestamp. For every live-price value, state the exact URL fetched and the UTC timestamp of the fetch to the minute (e.g.,
2026-05-13T14:32Z). - On-demand pricing only unless told otherwise. Do not apply reserved capacity, committed use, or enterprise negotiated pricing unless the user explicitly requests it.
- No credentials required or accepted. All provider pricing pages are public and unauthenticated. Never ask for API keys, billing account IDs, or tenant-specific data.
- FOCUS column mapping. Where a cost estimate is produced, note the corresponding FOCUS v1.2 columns:
BilledCost(what the provider charges),EffectiveCost(after credits or discounts),ServiceCategory(AI and Machine Learning),ChargeCategory(Usage),SkuId(model ID + deployment tier),SkuPriceId(price dimension: input-token / output-token / cached-token / image / gpu-hour). - Load references only when needed.
Pricing dimensions
| Dimension | Unit | Applies to |
|---|---|---|
| Input tokens | per 1M tokens | All text models |
| Output tokens | per 1M tokens | All text models |
| Cached input tokens | per 1M tokens | Models with context caching (Anthropic, Gemini, Bedrock) |
| Batch input tokens | per 1M tokens | Models with async batch mode (Anthropic, OpenAI, Bedrock) |
| Batch output tokens | per 1M tokens | Models with async batch mode |
| Images (input) | per image or per 1K images | Multimodal models |
| GPU-hour | per hour per GPU type | Self-hosted / dedicated endpoints (Vertex AI, Bedrock provisioned throughput, Azure PTU) |
Providers in scope
| Provider | Deployment target | Reference |
|---|---|---|
| Anthropic | Direct API | references/providers.md |
| OpenAI | Direct API | references/providers.md |
| Vertex AI | references/providers.md | |
| AWS | Bedrock | references/providers.md |
| Azure | Azure OpenAI Service | references/providers.md |
| OCI | Generative AI Service | references/providers.md |
Response minimum
Return, at minimum:
- confirmed model name(s), provider(s), and deployment target(s)
- pricing dimensions covered (input, output, cached, batch, image, GPU-hour)
- line-item price table: model | provider | dimension | unit price | provenance label | source URL | fetch timestamp
- for comparative tables: a summary row noting the cheapest and most expensive option per dimension
- key assumptions (on-demand, no reserved/committed pricing, USD unless stated)
- FOCUS column mapping for any cost estimate produced
References
Load these only when needed:
- Provider pricing URLs — public pricing page URLs per provider for live WebFetch.
- Token economics — input/output/cached/batch pricing model, $/M token math, and context caching cost curve.
Files (vanguard-frontier-agentic)
-
references
-
providers.md 4.2 KB
# Provider Pricing URLs Use these URLs with WebFetch to retrieve live pricing. All pages are public and require no authentication. ## Anthropic | Resource | URL | |---|---| | Pricing page | https://docs.anthropic.com/en/docs/about-claude/pricing | | Model overview | https://docs.anthropic.com/en/docs/about-claude/models/overview | Pricing dimensions available: input tokens ($/1M), output tokens ($/1M), prompt caching write ($/1M), prompt caching read ($/1M), batch input ($/1M), batch output ($/1M). Models published as of last verification: Claude Opus 4, Claude Sonnet 4.5, Claude Haiku 3.5, and prior generation models. ## OpenAI | Resource | URL | |---|---| | Pricing page | https://platform.openai.com/docs/pricing | | Models list | https://platform.openai.com/docs/models | Pricing dimensions available: input tokens ($/1M), cached input tokens ($/1M), output tokens ($/1M). Batch API discounts published separately on the same page. Models published: GPT-4o, GPT-4o mini, o1, o3, o4-mini, text-embedding-3-large/small, DALL-E (per image). ## AWS Bedrock | Resource | URL | |---|---| | Pricing page | https://aws.amazon.com/bedrock/pricing/ | | On-demand pricing table | https://aws.amazon.com/bedrock/pricing/#On-demand | | Batch inference | https://aws.amazon.com/bedrock/pricing/#Batch_inference | | Provisioned throughput | https://aws.amazon.com/bedrock/pricing/#Provisioned_throughput | Pricing dimensions available: on-demand input/output tokens ($/1000 tokens or $/1M tokens depending on model), batch input/output tokens, provisioned throughput model units ($/hour). Models published: Anthropic Claude family, Amazon Titan, Amazon Nova, Meta Llama, Mistral, Cohere, AI21 Jurassic, Stability AI. Note: Bedrock token prices are sometimes published per 1,000 tokens rather than per 1M. Normalize to $/1M for comparison by multiplying by 1,000. ## Azure OpenAI Service | Resource | URL | |---|---| | Pricing page | https://azure.microsoft.com/en-us/pricing/details/cognitive-services/openai-service/ | | Model catalog | https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/models | Pricing dimensions available: pay-as-you-go input/output tokens ($/1K tokens), provisioned throughput units (PTU, $/hour/PTU). Prices vary by Azure region. Models published: GPT-4o, GPT-4o mini, o1, o3, o4-mini, text-embedding, DALL-E 3. Note: Azure OpenAI prices are often published per 1,000 tokens. Normalize to $/1M by multiplying by 1,000. ## Google Vertex AI | Resource | URL | |---|---| | Generative AI pricing | https://cloud.google.com/vertex-ai/generative-ai/pricing | | Gemini model pricing | https://cloud.google.com/vertex-ai/generative-ai/pricing#gemini-models | | Embedding pricing | https://cloud.google.com/vertex-ai/generative-ai/pricing#embedding-models | Pricing dimensions available: input tokens ($/1M), output tokens ($/1M), context caching ($/1M per hour storage + discounted input tokens), grounding ($/1K queries), image input (per image or per 1K images for multimodal). Models published: Gemini 2.5 Pro, Gemini 2.5 Flash, Gemini 1.5 Pro/Flash, text-embedding-004, Imagen 3. ## OCI Generative AI | Resource | URL | |---|---| | Service overview | https://www.oracle.com/cloud/ai/generative-ai/ | | Pricing page | https://www.oracle.com/cloud/ai/generative-ai/pricing/ | Pricing dimensions available: on-demand token pricing (input and output per 1M tokens), dedicated AI cluster pricing (unit/hour). Models published: Cohere Command R/R+, Meta Llama 3 family. Note: OCI Generative AI pricing may not be listed on the main service overview page. If the pricing page URL returns no results, also try https://www.oracle.com/cloud/price-list/ and search for "Generative AI". ## Fetch strategy When using WebFetch against a pricing page: 1. Fetch the canonical pricing URL listed above. 2. Record the response timestamp in UTC (ISO 8601, e.g., `2026-05-13T14:32Z`). 3. Extract the relevant model row(s) for the requested model name or family. 4. If the page uses JavaScript to render a pricing table, it may return partial or placeholder content — in that case, label the result `documentation-based` and note the URL fetched. 5. If the fetch fails entirely, fall back to the most recent documentation-based price known and label it `documentation-based` with the documentation URL. -
token-economics.md 4.9 KB
# Token Economics ## Pricing model overview Foundation model API pricing is charged on consumption of compute resources expressed as token counts or equivalent units. The billing dimensions are: | Dimension | Definition | |---|---| | Input tokens | Tokens in the prompt sent to the model, including system prompt, user turn, tool definitions, and injected context | | Output tokens | Tokens generated by the model in its response, including reasoning tokens where applicable | | Cached input tokens | Input tokens served from the provider's prompt cache rather than re-computed; priced at a discount to standard input | | Batch input tokens | Input tokens submitted through an asynchronous batch API; priced at a discount to real-time input | | Batch output tokens | Output tokens returned from an asynchronous batch job | | Image input | Image tiles or full images submitted as multimodal input; some providers charge per tile, others per image | | GPU-hour | Compute time consumed by a dedicated model endpoint (provisioned throughput), independent of token count | ## $/M token calculation To express a per-token price as $/1M tokens: - If the provider publishes $/token: multiply by 1,000,000. - If the provider publishes per 1K tokens: multiply by 1,000. - If the provider publishes per 1M tokens: use as-is. Example: AWS Bedrock publishes Claude Haiku at $0.00025 per 1K input tokens. Normalized: $0.00025 × 1,000 = $0.25 per 1M input tokens. ## Monthly cost estimate formula ``` monthly_cost = (monthly_input_tokens / 1_000_000 × input_price_per_M) + (monthly_output_tokens / 1_000_000 × output_price_per_M) ``` Where cached or batch tokens replace standard tokens at the relevant discount rate: ``` effective_input_cost = (uncached_input_tokens / 1_000_000 × standard_input_price) + (cached_input_tokens / 1_000_000 × cached_input_price) ``` ## How context caching changes the cost curve Context caching allows the provider to store the KV (key-value) state of a fixed prefix — typically a long system prompt or retrieval context — and reuse it across multiple requests without recomputing it from scratch. Effect on pricing: - The cached prefix is charged at the **cache write price** (typically equal to or slightly above the standard input price) on first ingestion. - Subsequent requests that hit the cache are charged at the **cached read price**, which is typically 80–90% cheaper than the standard input price. - Storage of cached context may also incur a time-based charge (e.g., Vertex AI charges per 1M cached tokens per hour of storage). Break-even analysis: context caching becomes cost-effective when the same prefix is reused across enough requests that the cumulative cache read savings exceed the cache write cost plus storage cost. Example: a 100K-token system prompt sent to 10 requests per hour. Without caching: 100K tokens × 10 req/hr × $3.00/1M = $3.00/hr input cost. With caching (write once at $3.75/1M, read at $0.30/1M + $0.48/1M/hr storage for 100K tokens): - Write: 100K × $3.75/1M = $0.375 (one-time per cache TTL) - Reads: 100K × 10 × $0.30/1M = $0.30/hr - Storage: 100K × $0.48/1M/hr = $0.048/hr - Total: $0.348/hr after initial write — 88% cheaper than uncached. The specific prices in this example are illustrative; always fetch live prices before computing. ## Batch pricing trade-offs Batch APIs (Anthropic Batch API, OpenAI Batch API, AWS Bedrock Batch Inference) offer 50% discounts on input and output tokens in exchange for: - Asynchronous delivery (results returned within a fixed SLA window, typically 24 hours) - No real-time latency guarantee - Minimum batch size requirements on some providers Batch pricing is appropriate for: offline evaluation, document processing, embedding generation at scale, non-interactive summarization pipelines. Batch pricing is not appropriate for: real-time user-facing inference, streaming responses, interactive agents. ## FOCUS column mapping When expressing foundation model costs as FOCUS v1.2 line items: | FOCUS column | Value | |---|---| | `ServiceCategory` | AI and Machine Learning | | `ServiceName` | Provider-specific (e.g., Amazon Bedrock, Azure OpenAI Service, Vertex AI) | | `ChargeCategory` | Usage | | `ChargeFrequency` | Usage-based | | `BilledCost` | Actual amount charged (post any negotiated discounts) | | `EffectiveCost` | BilledCost minus any credits or committed-use amortization | | `ListCost` | Published on-demand price before discounts | | `SkuId` | Model identifier (e.g., `anthropic.claude-sonnet-4-5-v1:0`) | | `SkuPriceId` | Price dimension (e.g., `input-token`, `output-token`, `cached-input-token`, `batch-input-token`) | | `ResourceId` | API endpoint or model deployment ARN/resource ID if available | | `UsageType` | Token dimension (e.g., `InputTokens`, `OutputTokens`) | | `UsageQuantity` | Number of tokens in the pricing unit (e.g., count of 1M-token increments) |
-
-
metadata.json 1.3 KB
{ "id": "fetch-foundation-model-pricing", "name": "Fetch Foundation Model Pricing", "type": "skill", "provider": "multi-cloud", "harnesses": [ "codex", "claude-code", "cursor", "gemini", "kiro", "other" ], "summary": "Fetch live per-token, per-image, and per-GPU-hour prices for foundation models across Anthropic, OpenAI, Google, AWS Bedrock, Azure OpenAI, OCI Generative AI, and Vertex AI. Supports single-model lookup and comparative multi-provider tables with provenance labels.", "source_type": "original", "official_docs": [ "https://docs.anthropic.com/en/docs/about-claude/pricing", "https://platform.openai.com/docs/pricing", "https://aws.amazon.com/bedrock/pricing/", "https://azure.microsoft.com/en-us/pricing/details/cognitive-services/openai-service/", "https://cloud.google.com/vertex-ai/generative-ai/pricing", "https://www.oracle.com/cloud/ai/generative-ai/" ], "security_notes": "All provider pricing pages are public and unauthenticated. Never accept or request API keys, billing account IDs, cost export access, tenant IDs, or any cloud credentials to fetch list prices.", "last_verified": "2026-05-13", "path": "skills/finops/fetch-foundation-model-pricing", "author": "github: VincentChuWaiChow", "version": "0.1.1", "lifecycle": "experimental" } -
README.md 1.6 KB
# Fetch Foundation Model Pricing A FinOps skill that retrieves live public pricing for foundation models across major AI and cloud providers, returning structured tables with mandatory provenance labels and source timestamps. ## Purpose Fetch current per-token, per-image, and per-GPU-hour prices from Anthropic, OpenAI, Google (Vertex AI), AWS Bedrock, Azure OpenAI Service, and OCI Generative AI. Supports single-model lookups and side-by-side comparative tables. ## Allowed tools `Read` `Grep` `Glob` `WebFetch` ## Usage **Single-model lookup:** Provide a model name and deployment target (e.g., "What does Claude Sonnet 4.5 cost per million input tokens on Bedrock?"). The skill fetches the live price, labels it with source URL and ISO 8601 timestamp, and returns a single-row price table. **Comparative table:** Provide two or more models or providers and a task type (e.g., "Compare GPT-4o, Claude Sonnet, and Gemini Pro on input and output token cost"). The skill fetches each price independently and builds a labeled multi-row comparison with cheapest/most-expensive summary rows. ## Trust posture Read-only. No cloud credentials, billing account IDs, or tenant data accepted. All pricing pages are public and unauthenticated. Every price value carries a provenance label (`live-price`, `documentation-based`, `assumed`, or `excluded`) with source URL and fetch timestamp. FOCUS v1.2 column mapping is included for any cost estimate produced (BilledCost, EffectiveCost, ServiceCategory, ChargeCategory, SkuId, SkuPriceId). See [SKILL.md](SKILL.md) for the full operating protocol, pricing dimensions, and response shape. -
SKILL.md 5.2 KB
--- name: fetch-foundation-model-pricing description: Fetch live per-token, per-image, and per-GPU-hour prices for foundation models across Anthropic, OpenAI, Google, AWS Bedrock, Azure OpenAI, OCI Generative AI, and Vertex AI. Supports single-model lookup and comparative multi-provider tables. Every price is labeled with source URL and ISO 8601 fetch timestamp. No credentials accepted. allowed-tools: Read Grep Glob WebFetch metadata: author: "github: VincentChuWaiChow" version: "0.1.1" updated: "2026-05-13" category: finops lifecycle: experimental --- # Fetch Foundation Model Pricing ## Purpose Retrieve current public pricing for foundation models across the major AI/cloud providers and return structured, provenance-labeled output. Supports two modes: - **Single-model lookup**: fetch the current price for a specific model and deployment target (e.g., Claude Sonnet 4.5 on Anthropic direct, or on Bedrock). - **Comparative table**: build a side-by-side price comparison across two or more models or providers for the same task type (text, image, embedding, or GPU-hour). ## When to use Use this skill when: - The user asks "how much does model X cost per token / per image / per GPU-hour" - The user wants to compare inference costs across two or more providers for the same model family or equivalent capability tier - The user needs to estimate monthly AI inference spend given a volume projection - The user wants to understand how context caching or batch pricing changes the effective cost curve - The user wants a FOCUS-aware cost breakdown (BilledCost, EffectiveCost) for AI inference line items ## Operating rules - **Fetch live prices first.** Use WebFetch to retrieve prices from each provider's public pricing page before relying on any internal knowledge. AI model pricing changes frequently; stale numbers mislead. - **Label every price.** Each price value must carry exactly one provenance label: - `live-price` — fetched from a provider's public pricing page or API within this session; include source URL and ISO 8601 timestamp. - `documentation-based` — sourced from official documentation when a live fetch was not possible; note the documentation URL and its visible publication date. - `assumed` — derived from an analogous model or tier when no direct published price exists; state the assumption explicitly. - `excluded` — pricing that exists but was intentionally omitted from the output; state why. - **Include source URL and timestamp.** For every live-price value, state the exact URL fetched and the UTC timestamp of the fetch to the minute (e.g., `2026-05-13T14:32Z`). - **On-demand pricing only unless told otherwise.** Do not apply reserved capacity, committed use, or enterprise negotiated pricing unless the user explicitly requests it. - **No credentials required or accepted.** All provider pricing pages are public and unauthenticated. Never ask for API keys, billing account IDs, or tenant-specific data. - **FOCUS column mapping.** Where a cost estimate is produced, note the corresponding FOCUS v1.2 columns: `BilledCost` (what the provider charges), `EffectiveCost` (after credits or discounts), `ServiceCategory` (AI and Machine Learning), `ChargeCategory` (Usage), `SkuId` (model ID + deployment tier), `SkuPriceId` (price dimension: input-token / output-token / cached-token / image / gpu-hour). - **Load references only when needed.** ## Pricing dimensions | Dimension | Unit | Applies to | |---|---|---| | Input tokens | per 1M tokens | All text models | | Output tokens | per 1M tokens | All text models | | Cached input tokens | per 1M tokens | Models with context caching (Anthropic, Gemini, Bedrock) | | Batch input tokens | per 1M tokens | Models with async batch mode (Anthropic, OpenAI, Bedrock) | | Batch output tokens | per 1M tokens | Models with async batch mode | | Images (input) | per image or per 1K images | Multimodal models | | GPU-hour | per hour per GPU type | Self-hosted / dedicated endpoints (Vertex AI, Bedrock provisioned throughput, Azure PTU) | ## Providers in scope | Provider | Deployment target | Reference | |---|---|---| | Anthropic | Direct API | references/providers.md | | OpenAI | Direct API | references/providers.md | | Google | Vertex AI | references/providers.md | | AWS | Bedrock | references/providers.md | | Azure | Azure OpenAI Service | references/providers.md | | OCI | Generative AI Service | references/providers.md | ## Response minimum Return, at minimum: - confirmed model name(s), provider(s), and deployment target(s) - pricing dimensions covered (input, output, cached, batch, image, GPU-hour) - line-item price table: model | provider | dimension | unit price | provenance label | source URL | fetch timestamp - for comparative tables: a summary row noting the cheapest and most expensive option per dimension - key assumptions (on-demand, no reserved/committed pricing, USD unless stated) - FOCUS column mapping for any cost estimate produced ## References Load these only when needed: - [Provider pricing URLs](references/providers.md) — public pricing page URLs per provider for live WebFetch. - [Token economics](references/token-economics.md) — input/output/cached/batch pricing model, $/M token math, and context caching cost curve.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.