frontend-finops-cost-to-serve-review
Build a cost-to-serve model covering CDN egress, SSR/edge compute, image transformation, and CI build-minute spend for a frontend surface, and rank remediation options by dollar savings weighed against Core Web Vitals and security impact, without treating cost-cutting and securit
Install
npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/frontend/frontend-finops-cost-to-serve-review
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Frontend FinOps Cost-to-Serve Review
Purpose
Frontend architecture choices — SSR vs. static generation, ISR revalidation cadence, image-pipeline design, third-party script sprawl, CI build-minute consumption — are cloud-spend decisions as much as they are UX decisions, but they are almost never modeled that way: performance is reviewed by one team, cloud spend by another, and the causal link between them goes unmeasured. This skill exists to build a defensible cost-to-serve model for a frontend surface (CDN egress, SSR/edge compute, image transform, build-minutes), tie every dollar figure to an explicit evidence level, and rank remediation options by savings-to-risk ratio — never presenting a modeled estimate as an audited invoice, and never treating cost reduction as separable from performance and security posture.
When to use
Use this skill when the user asks to:
- estimate the CDN egress, SSR/edge-compute, or image-transform cost of a frontend surface,
- evaluate whether an SSR/ISR/edge-function architecture choice is cost-appropriate at current or projected traffic,
- identify which third-party scripts or dependencies are the largest cost/performance line items,
- rank cost-reduction options by dollar impact versus Core Web Vitals or security trade-off,
- explain why a frontend surface's cloud bill grew disproportionately to its traffic.
Context7 Documentation Protocol
Before making any framework-specific claim about caching, revalidation, rendering mode, or image-optimization behavior (Next.js, or any other framework named in the task), resolve the library via Context7 (resolve-library-id) and query current docs (query-docs) rather than relying on training-data memory. Frameworks change caching/revalidation defaults across major versions (for example, Next.js has changed default fetch caching behavior between major versions), and a cost model built on a stale caching assumption will misstate invocation counts by an order of magnitude. Label every framework-behavior claim as context7-grounded (as of <library-id>/<version if resolved>), documentation-based (unverified against Context7), or inference — never state framework caching/billing behavior as fact without one of these labels. If Context7 has no coverage for a named library or version, say so explicitly and fall back to official vendor docs, marking the claim uncertain.
Lean operating rules
- Always state the evidence level of every dollar figure:
billing-data-verified(user supplied actual invoice/usage export),modeled-from-public-pricing(calculated from published rate cards and estimated volume), orinference(no volume data, rough order of magnitude only). Never present a modeled estimate as an audited number. - Ground SSR/ISR/edge-invocation-count assumptions in the actual framework's documented caching/revalidation behavior (via Context7/official docs) before estimating compute cost — invocation-shape assumptions are the single biggest source of cost-model error. Time-based revalidation (stale-while-revalidate) and on-demand revalidation (tag/path invalidation) produce fundamentally different invocation curves; do not conflate them.
- Do not recommend removing a script, feature, or rendering mode purely because it is expensive without checking whether it drives revenue (checkout, support chat, personalization) — cost-to-serve is a trade-off model, not a cost-minimization mandate.
- Never recommend removing a named security control (CSP, WAF rule, image-pipeline malware/content scan, TLS termination tier, bot-mitigation layer) to cut cost without flagging it as requiring explicit security-owner approval — cost review is not a backdoor to loosen the security posture reviewed elsewhere.
- Distinguish traffic-linear cost growth (predictable, budgetable) from non-linear cost growth (e.g., uncached per-request SSR, unbounded on-demand image-transform variants, retry storms) and flag non-linear growth as an urgent architectural risk regardless of current dollar total — a small bill growing 3x per traffic-doubling is a bigger red flag than a large flat bill.
- Do not treat a public cloud/CDN price sheet as guaranteed pricing for the user's account: committed-use discounts, negotiated enterprise rates, and regional price variance can change real cost by 30-70%. Label public-rate-card math accordingly and ask for a billing export when precision matters.
- Load
references/ssr-isr-invocation-cost-modeling.mdonly when the SSR/ISR/edge-function invocation shape is the primary cost driver being modeled. - Load
references/third-party-script-cost-attribution.mdonly when ranking third-party scripts/dependencies by cost and performance impact against business value.
References
Load these only when needed:
- SSR/ISR invocation cost modeling — use when modeling edge/SSR compute cost, grounding invocation-count assumptions in the framework's actual caching/revalidation behavior, and distinguishing linear from non-linear cost-growth patterns.
- Third-party script cost attribution — use when ranking third-party scripts/dependencies by their bundle-weight, request-count, and CDN-egress contribution against their measured business value.
Response minimum
Return, at minimum:
- the cost-to-serve breakdown by category (CDN egress, SSR/edge compute, image transform, CI build-minutes) with an explicit evidence level per figure,
- dollar cost per 1,000 pageviews (or per 1,000 requests) at current or stated traffic,
- a ranked remediation list, each item with estimated dollar savings and its stated Core Web Vitals or security trade-off,
- an explicit flag on any non-linear cost-growth risk found, independent of current dollar total,
- an explicit flag on any recommendation that touches a named security control, stating it requires named-owner sign-off before action.
Files (vanguard-frontier-agentic)
-
references
-
ssr-isr-invocation-cost-modeling.md 7.6 KB
# SSR/ISR Invocation Cost Modeling Use this reference when the primary cost driver being modeled is server-side rendering, incremental static regeneration (ISR), or edge-function invocation — i.e., compute that runs per-request or per-revalidation rather than being served as a static, fully-cached asset. ## What people get wrong The common bad assumption is: > "We use SSR/ISR, so it's basically static — the cache handles it." That is incomplete, and it is the single most common cause of frontend cloud-cost surprises. Caching mode is not binary; it determines *how often* compute runs, and that invocation rate is the entire cost driver. Two surfaces that both say "we use ISR" can have 100x different compute bills depending on revalidation cadence, tag-invalidation frequency, and whether requests bypass cache due to query-string or header variance. ## Officially grounded shape (Context7-verified, Next.js docs) Per current Next.js documentation (`/vercel/next.js`): - **Time-based revalidation** (`fetch(url, { next: { revalidate: N } })`) uses a stale-while-revalidate pattern: cached content is served immediately, and regeneration happens in the background once the content's age exceeds `N` seconds. This bounds worst-case regeneration frequency to roughly `traffic-during-window / N`, not per-request. - **On-demand revalidation** (`revalidateTag()`, `revalidatePath()`) explicitly invalidates cached content from a server action or route handler, causing the *next* request after invalidation to trigger a fresh render. This decouples regeneration from a timer entirely — invocation count is driven by how often the invalidation trigger fires (e.g., a CMS webhook), not by a fixed interval. - **Image optimization** (legacy `next/image` pipeline): images are optimized dynamically on first request and cached in `<distDir>/cache/images`; on expiration, a stale image is served immediately while regeneration happens in the background and the new result is cached. This means image-transform compute cost is driven by the *cardinality of distinct requested variants* (width/quality/format combinations), not by pageview count — a large `deviceSizes`/`imageSizes` matrix with the same source image multiplies cache-miss compute even for one logical image. > Version note: caching defaults have changed across Next.js major versions (`fetch` caching behavior, `dynamicIO`/cache-components experiments). Verify default caching behavior against the installed version via Context7 or official docs before assuming a specific revalidation model — do not assume App Router defaults carry over from Pages Router or from an older major version. ## Non-negotiable design rules ### 1. Classify the invocation shape before pricing anything Do not price "SSR" as a single line item. Classify each route/surface into one of: - **Static / fully cached** — served from CDN edge cache, effectively zero marginal compute per request. - **Time-based ISR** — bounded regeneration rate; cost scales with `unique-pages / revalidate-interval`, not with traffic. - **On-demand revalidated** — regeneration rate scales with invalidation-trigger frequency (e.g., CMS publish events), which can spike independently of traffic. - **Per-request SSR (no cache)** — regeneration rate equals request rate; this is the only shape where cost is strictly traffic-linear in compute, and the most expensive per-pageview. Mixing these into one blended "SSR cost" figure hides which routes are the actual problem. ### 2. Treat cache-key fragmentation as a cost multiplier Query parameters, cookies, or headers included in the cache key (e.g., per-locale, per-experiment-variant, per-logged-in-state rendering) multiply the effective number of "unique pages" that each need independent regeneration. A page with 3 locales x 4 A/B variants x authenticated/anonymous is 24 cache entries, not 1 — model compute cost against that multiplied cardinality, not the logical route count. ### 3. Distinguish linear from non-linear cost growth explicitly - **Linear**: cost per pageview is roughly constant as traffic grows (static assets, well-bounded time-based ISR, edge-cached SSR responses). - **Non-linear**: cost grows faster than traffic (per-request SSR with no cache, unbounded image-variant cardinality, cache-key fragmentation that scales with user/session count, retry storms from a slow origin). Flag any non-linear pattern as an architectural risk even if the current bill is small — non-linear cost curves are the ones that blow budgets during a traffic spike or a viral moment, precisely when the business can least afford a surprise, and precisely when engineering has the least slack to fix it under pressure. ### 4. Do not price edge/serverless invocations from memory Provider pricing for edge-function invocations, SSR compute duration, and image-transform requests changes across pricing tiers and providers, and differs materially between "included in platform plan" vs. "metered pay-per-use" billing. Ground any specific per-invocation or per-GB-second rate in the provider's current published pricing page or a user-supplied billing export — do not assert a specific dollar rate from training-data memory, since these rates are revised frequently and vary by committed-use tier. ## Minimal safe modeling flow 1. Classify each route/surface by invocation shape (static / time-based ISR / on-demand / per-request SSR). 2. For each non-static shape, identify the cache-key dimensions (locale, auth state, experiment variant, query params) and compute effective cardinality. 3. Get actual or estimated traffic volume per surface (analytics export, CDN log sample, or user-stated estimate — label whichever it is). 4. Compute expected invocation count: time-based ISR ≈ `cardinality / revalidate_interval_seconds × window_seconds`; on-demand ≈ trigger frequency; per-request SSR ≈ request volume. 5. Apply current provider pricing (billing export preferred; public rate card as fallback, labeled `modeled-from-public-pricing`) to invocation count and compute duration/memory profile. 6. Flag any surface where invocation count scales faster than traffic as non-linear risk, independent of the resulting dollar figure. ## Adversarial checklist Before presenting an SSR/ISR/edge cost figure as reliable, answer these: - Did I classify by actual invocation shape, or did I average everything into one "SSR" bucket? - Does the cache key include anything that fragments cardinality (locale, auth, experiment, query string)? - Is the revalidation interval time-based, on-demand, or both layered together? - Did I verify current caching-default behavior against the installed framework version via Context7/official docs, or am I assuming defaults from memory? - Is the pricing rate billing-data-verified, or modeled from a public rate card that may not reflect the account's actual committed-use discount? - Have I flagged any non-linear growth pattern regardless of whether the current dollar total looks small? If any answer is "I don't know," say so explicitly rather than presenting a confident number. ## When to push back Push back if the user asks to: - disable a cache layer, revalidation window, or CDN tier purely to "make the number smaller" without checking Core Web Vitals impact, - assume a flat per-request cost figure without classifying invocation shape first, - accept a public price-sheet estimate as final without flagging that real committed-use pricing may differ substantially, - treat a non-linear cost-growth pattern as acceptable because "the bill is still small today." Those shortcuts produce a number that looks precise and is not, and they defer an architectural risk instead of surfacing it. -
third-party-script-cost-attribution.md 6.7 KB
# Third-Party Script Cost Attribution Use this reference when ranking third-party scripts, tags, or dependencies by their cost and performance contribution against their measured business value — not just their bundle size. ## What people get wrong The common bad assumption is: > "Third-party scripts are free — they're hosted by the vendor, not by us." That is wrong in two separate ways. First, the vendor's hosting cost is irrelevant; what matters is the cost *the frontend surface incurs on behalf of loading and running it*: the CDN egress/request count if self-hosted or proxied, the CI build-minute cost if bundled at build time, and — the largest and most commonly ignored cost — the Core Web Vitals and conversion-rate impact of the script's execution, which translates directly into revenue at the margin. Second, "third party" bundles wildly different cost profiles under one label: a 2 KB feature flag SDK and a 400 KB chat-widget bundle with its own CDN calls are not comparable line items. ## Officially grounded shape (web.dev / Lighthouse performance budgets) Per current web.dev guidance: - Performance budgets can be defined per resource type (script, image, font, stylesheet, third-party, total) and per resource *count*, not just size — a `resourceCounts` budget on `third-party` requests is a first-class Lighthouse budget primitive (`budget.json`), independent from a size-based budget. - Quantity-based metrics (page weight, request count, third-party request count) are explicitly recommended as an early-stage, easy-to-communicate proxy — useful for a cost/business conversation with non-engineering stakeholders even before deeper timing-based metrics are modeled. - A commonly cited critical-path budget is under ~170 KB of compressed/minified critical-path resources for a fast experience on constrained devices/networks — useful as a reference point when arguing that a given script's weight materially eats the surface's total budget, not just "feels heavy." ## Non-negotiable design rules ### 1. Attribute cost per script, not per "third-party" bucket For each script, separately account for: - **Bundle weight** — transferred bytes (compressed), which drives CDN egress cost if proxied/self-hosted, or drives page-load time regardless of hosting. - **Request count** — each script may fan out to its own CDN/API calls (analytics beacons, ad-tech auctions, chat-widget polling) that are invisible in a simple bundle-size measurement. - **Execution cost** — main-thread blocking time and its measured effect on Core Web Vitals (particularly INP and LCP), which is a performance cost, not a cloud-billing cost, but converts to a revenue/conversion cost. - **Build-time cost** — if the script or its wrapper is bundled/transpiled/type-checked in CI, attribute a share of CI build-minute spend to it. ### 2. Require a stated business justification before ranking a script as "remove" A script's cost is only meaningful relative to its value. Do not rank a script for removal purely by weight or execution cost — pair every cost figure with the script's stated business function (conversion tracking, support chat, personalization, fraud detection, ad revenue) and, where available, a measured lift/attribution figure. A high-cost script tied to a verified revenue driver ranks differently than an equally expensive script with no attributed owner or unclear purpose. ### 3. Separate "safe to remove," "safe to defer/lazy-load," and "requires owner sign-off" - **Safe to remove**: no attributed business owner, duplicate functionality with another already-loaded script, or confirmed dead/unused (verify via network trace or tag-manager audit, not assumption). - **Safe to defer/lazy-load**: legitimate business function but not required for above-the-fold interaction (e.g., load after first user interaction or via `next/script`'s lazy-loading strategies where the framework supports it) — verify the framework's specific lazy-loading API via Context7/official docs before recommending a specific implementation, since loading-strategy APIs vary by framework and version. - **Requires owner sign-off**: anything tied to security (bot mitigation, fraud detection), compliance (consent management, accessibility overlays with a legal mandate), or a verified revenue driver — never recommend removing or deferring these unilaterally on cost grounds alone; escalate to the named business/security owner. ### 4. Do not conflate script cost with hosting cost A script's own CDN bill is the vendor's problem. The frontend surface's cost exposure is: its own egress if self-hosting/proxying the script, its own CI cost if bundling it, and its own performance-to-revenue cost from execution. Keep these separate in the report so a reader does not think "removing this script saves us their hosting bill." ## Minimal safe attribution flow 1. Enumerate all third-party scripts/tags currently loaded (network trace, tag manager audit, or `package.json`/bundle-analyzer output for bundled third-party code). 2. For each script, capture: transferred bytes, request count/fan-out, measured main-thread/CWV impact if available (lab data from Lighthouse/WebPageTest, or field data from CrUX/RUM), and CI build-time share if bundled. 3. Attribute a stated business function and, if available, a measured value signal (conversion lift, support-ticket deflection, ad revenue) — mark as `owner-confirmed` or `unattributed` if no owner responds. 4. Rank by cost-to-value ratio, bucketing into safe-to-remove, safe-to-defer, and requires-sign-off. 5. Present dollar cost (CDN egress + build-minute share, evidence-labeled) alongside the performance/CWV cost and the stated business value for each item. ## High-risk assumptions to kill - "It's third-party, so it doesn't cost us anything." - "Nobody's complained about this script, so it must be fine to keep loading it eagerly." - "This script is small, so its execution cost doesn't matter." - "We can just lazy-load everything without checking what breaks." - "Unattributed scripts are automatically safe to remove" — unattributed does not mean unused; verify before removing. ## When to push back Push back if the user asks to: - remove a script tied to fraud detection, bot mitigation, consent management, or an accessibility legal mandate purely for cost savings, without named security/compliance owner sign-off, - rank scripts by bundle size alone without checking request fan-out or CWV execution impact, - treat "no one knows what this does" as justification for silent removal in a live production surface without a verified owner check or a monitored rollback path. Those are cost-cutting shortcuts that trade an unmeasured business or security risk for a measured dollar saving — surface the trade-off explicitly instead of making the call silently.
-
-
metadata.json 1.2 KB
{ "id": "frontend-finops-cost-to-serve-review", "name": "Frontend FinOps Cost-to-Serve Review", "type": "skill", "provider": "frontend", "harnesses": [ "codex", "copilot", "claude-code", "cursor", "gemini", "kiro" ], "summary": "Builds a defensible cost-to-serve model (CDN egress, SSR/edge compute, image transform, build-minutes) for a frontend surface and ranks remediation options by dollar savings versus performance/security risk, keeping performance and cloud spend tied together instead of reviewed separately.", "source_type": "adapted", "official_docs": [ "https://web.dev/articles/total-byte-weight", "https://nextjs.org/docs/app/guides/self-hosting", "https://www.finops.org/framework/principles/", "https://developer.chrome.com/docs/crux" ], "security_notes": "Never recommend disabling CSP, WAF, image-pipeline security scanning, or TLS termination purely to cut cost; any such trade-off requires explicit security-owner sign-off, not a default recommendation.", "last_verified": "2026-07-02", "path": "skills/frontend/frontend-finops-cost-to-serve-review", "author": "github: VincentChuWaiChow", "version": "0.1.0" } -
SKILL.md 6.4 KB
--- name: frontend-finops-cost-to-serve-review description: Build a cost-to-serve model covering CDN egress, SSR/edge compute, image transformation, and CI build-minute spend for a frontend surface, and rank remediation options by dollar savings weighed against Core Web Vitals and security impact, without treating cost-cutting and security/performance as unrelated trade-offs. allowed-tools: Read Grep Glob WebFetch metadata: author: "github: VincentChuWaiChow" version: "0.1.0" updated: "2026-07-02" category: finops --- # Frontend FinOps Cost-to-Serve Review ## Purpose Frontend architecture choices — SSR vs. static generation, ISR revalidation cadence, image-pipeline design, third-party script sprawl, CI build-minute consumption — are cloud-spend decisions as much as they are UX decisions, but they are almost never modeled that way: performance is reviewed by one team, cloud spend by another, and the causal link between them goes unmeasured. This skill exists to build a defensible cost-to-serve model for a frontend surface (CDN egress, SSR/edge compute, image transform, build-minutes), tie every dollar figure to an explicit evidence level, and rank remediation options by savings-to-risk ratio — never presenting a modeled estimate as an audited invoice, and never treating cost reduction as separable from performance and security posture. ## When to use Use this skill when the user asks to: - estimate the CDN egress, SSR/edge-compute, or image-transform cost of a frontend surface, - evaluate whether an SSR/ISR/edge-function architecture choice is cost-appropriate at current or projected traffic, - identify which third-party scripts or dependencies are the largest cost/performance line items, - rank cost-reduction options by dollar impact versus Core Web Vitals or security trade-off, - explain why a frontend surface's cloud bill grew disproportionately to its traffic. ## Context7 Documentation Protocol Before making any framework-specific claim about caching, revalidation, rendering mode, or image-optimization behavior (Next.js, or any other framework named in the task), resolve the library via Context7 (`resolve-library-id`) and query current docs (`query-docs`) rather than relying on training-data memory. Frameworks change caching/revalidation defaults across major versions (for example, Next.js has changed default `fetch` caching behavior between major versions), and a cost model built on a stale caching assumption will misstate invocation counts by an order of magnitude. Label every framework-behavior claim as `context7-grounded (as of <library-id>/<version if resolved>)`, `documentation-based (unverified against Context7)`, or `inference` — never state framework caching/billing behavior as fact without one of these labels. If Context7 has no coverage for a named library or version, say so explicitly and fall back to official vendor docs, marking the claim uncertain. ## Lean operating rules - Always state the evidence level of every dollar figure: `billing-data-verified` (user supplied actual invoice/usage export), `modeled-from-public-pricing` (calculated from published rate cards and estimated volume), or `inference` (no volume data, rough order of magnitude only). Never present a modeled estimate as an audited number. - Ground SSR/ISR/edge-invocation-count assumptions in the actual framework's documented caching/revalidation behavior (via Context7/official docs) before estimating compute cost — invocation-shape assumptions are the single biggest source of cost-model error. Time-based revalidation (stale-while-revalidate) and on-demand revalidation (tag/path invalidation) produce fundamentally different invocation curves; do not conflate them. - Do not recommend removing a script, feature, or rendering mode purely because it is expensive without checking whether it drives revenue (checkout, support chat, personalization) — cost-to-serve is a trade-off model, not a cost-minimization mandate. - Never recommend removing a named security control (CSP, WAF rule, image-pipeline malware/content scan, TLS termination tier, bot-mitigation layer) to cut cost without flagging it as requiring explicit security-owner approval — cost review is not a backdoor to loosen the security posture reviewed elsewhere. - Distinguish traffic-linear cost growth (predictable, budgetable) from non-linear cost growth (e.g., uncached per-request SSR, unbounded on-demand image-transform variants, retry storms) and flag non-linear growth as an urgent architectural risk regardless of current dollar total — a small bill growing 3x per traffic-doubling is a bigger red flag than a large flat bill. - Do not treat a public cloud/CDN price sheet as guaranteed pricing for the user's account: committed-use discounts, negotiated enterprise rates, and regional price variance can change real cost by 30-70%. Label public-rate-card math accordingly and ask for a billing export when precision matters. - Load `references/ssr-isr-invocation-cost-modeling.md` only when the SSR/ISR/edge-function invocation shape is the primary cost driver being modeled. - Load `references/third-party-script-cost-attribution.md` only when ranking third-party scripts/dependencies by cost and performance impact against business value. ## References Load these only when needed: - [SSR/ISR invocation cost modeling](references/ssr-isr-invocation-cost-modeling.md) — use when modeling edge/SSR compute cost, grounding invocation-count assumptions in the framework's actual caching/revalidation behavior, and distinguishing linear from non-linear cost-growth patterns. - [Third-party script cost attribution](references/third-party-script-cost-attribution.md) — use when ranking third-party scripts/dependencies by their bundle-weight, request-count, and CDN-egress contribution against their measured business value. ## Response minimum Return, at minimum: - the cost-to-serve breakdown by category (CDN egress, SSR/edge compute, image transform, CI build-minutes) with an explicit evidence level per figure, - dollar cost per 1,000 pageviews (or per 1,000 requests) at current or stated traffic, - a ranked remediation list, each item with estimated dollar savings and its stated Core Web Vitals or security trade-off, - an explicit flag on any non-linear cost-growth risk found, independent of current dollar total, - an explicit flag on any recommendation that touches a named security control, stating it requires named-owner sign-off before action.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.