frontend-observability-rum-instrumentation
Design or review browser-side Real User Monitoring instrumentation for Core Web Vitals (LCP, INP, CLS) using the web-vitals attribution build and distributed tracing via OpenTelemetry Web, enforcing lab-vs-field evidence labeling, sampling/cardinality sizing, and PII-in-telemetry
Install
npx skills add https://github.com/VincentChuWaiChow/vanguard-frontier-agentic/tree/master/skills/frontend/frontend-observability-rum-instrumentation
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install vincentchuwaichow-vanguard-frontier-agentic@llmmart
git clone https://github.com/VincentChuWaiChow/vanguard-frontier-agentic.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole vincentchuwaichow/vanguard-frontier-agentic collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Frontend Observability RUM Instrumentation
Purpose
Lab-only performance testing (Lighthouse, synthetic CI runs) systematically misses the device/network diversity of real users, and ad-hoc telemetry instrumentation routinely leaks PII into span attributes or blows up observability cost through unsampled high-cardinality export. This skill wires and reviews field RUM (web-vitals + OpenTelemetry Web) with explicit evidence labeling and privacy/cost guardrails baked in from the start.
When to use
Use this skill when the user asks to:
- instrument Core Web Vitals (LCP, INP, CLS) field measurement in a web app,
- set up or review OpenTelemetry Web browser tracing (document load, fetch/XHR spans),
- review an existing RUM/telemetry pipeline for PII exposure or sampling/cost issues,
- interpret or report on Core Web Vitals data and needs a lab-vs-field distinction made explicit,
- set or validate a performance budget against field p75 data.
When NOT to use
- Diagnosing why a specific LCP/INP/CLS number is high (phase decomposition of an existing regression) — hand off to
core-web-vitals-triage; this skill instruments the pipeline that produces the data, it does not decompose an already-captured regression. - Bundle size, code-splitting, or caching remediation once a cause is known — hand off to
bundle-budget-code-splitting-revieworservice-worker-cache-strategy-review. - Server-side or backend-service tracing with no browser-origin span — this skill is scoped to browser-side RUM and OpenTelemetry Web only.
Context7 Documentation Protocol
web-vitals and OpenTelemetry Web packages ship breaking API changes across major versions (attribution-build option shapes, sampler class names, exporter constructor options), and memorized snippets go stale. Before writing or reviewing instrumentation code:
- Call
ToolSearchwith query"context7"(or"select:mcp__Context7__resolve-library-id,mcp__Context7__query-docs") to load the Context7 tools if not already loaded this session. - Call
mcp__Context7__resolve-library-idforweb-vitalsand/oropentelemetry-jsbefore making any claim about their current API shape. Verified IDs as of this skill'supdateddate:/googlechrome/web-vitalsand/open-telemetry/opentelemetry-js. - Call
mcp__Context7__query-docsfor the specific mechanism in scope — e.g. "attribution build generateTarget option", "INPAttributionReportOpts durationThreshold", "WebTracerProvider spanProcessors config", "TraceIdRatioBasedSampler" — before ruling on it. Do this per review; do not reuse a prior session's memory of these APIs. - Known facts verified via Context7 as of this skill's
updateddate: the attribution build (web-vitals/attribution) acceptsAttributionReportOpts(reportAllChanges,generateTarget) ononCLS/onFCP/onLCP/onTTFB, andINPAttributionReportOptsadditionally exposesdurationThreshold(default40) andincludeProcessedEventEntries(defaulttrue) ononINP.WebTracerProvider(from@opentelemetry/sdk-trace-web) takes aspanProcessorsarray in its constructor — do not tell a user to call a separateaddSpanProcessormethod as the primary pattern without checking the installed SDK version.TraceIdRatioBasedSamplertakes a single ratio argument (0–1) and is normally wrapped inParentBasedSampler({ root: ... })so that downstream/parent sampling decisions are respected. - If Context7 is unavailable or returns no relevant match, fall back to
official_docs/references/*.mdand mark the claimdocumentation-based (Context7 unavailable)rather than presenting it as freshly verified. - Never invent a
web-vitalsmetric field, attribution property, OpenTelemetry exporter option, or sampler class that no queried source confirms.
Lean operating rules
- Always label every performance number as lab evidence (Lighthouse/CI/synthetic) or field evidence (RUM, real users — state the percentile; p75 is the CWV standard) — never let the two blur into an unlabeled "the site is fast" claim.
- Default to the standard
web-vitalsbuild (onLCP/onINP/onCLS) for production reporting; only use the/attributionbuild's richer diagnostic payload when actively debugging a specific regression, and only send that extra payload to a dev/debug destination, not blanket production export. - Do not enable
reportAllChanges: truein production by default — it multiplies event volume for a debugging-only benefit. - Before adding any custom span/metric attribute, run it through a PII check: no raw URLs with query strings, no user IDs, no free-text form input, no precise geolocation, unless explicitly scrubbed and justified.
- Size the OpenTelemetry sampling rate (e.g.
TraceIdRatioBasedSamplerwrapped inParentBasedSampler) against the stated/estimated traffic volume and the backend's cost/cardinality limits; never recommend 100% unsampled export for a high-traffic app without an explicit cost review. - Use OpenTelemetry Semantic Conventions attribute names instead of inventing ad-hoc span/attribute naming, so telemetry stays queryable/comparable across services.
- Treat any recommended performance budget threshold as grounded in web.dev's documented "good" ranges (LCP ≤2.5s, INP ≤200ms, CLS ≤0.1 at p75) unless the org has an explicitly stricter documented SLO.
- This skill performs static review and instrumentation-code authoring only; it does not deploy telemetry configuration to production or flip sampling/export settings on a live collector without explicit human sign-off logged outside this skill.
References
Load these only when needed:
- web-vitals attribution build wiring — use when writing or reviewing the actual
onLCP/onINP/onCLSinstrumentation call, choosing between standard and attribution builds, or settinggenerateTarget/durationThreshold/reportAllChanges. - OpenTelemetry Web tracing wiring — use only when wiring or reviewing
WebTracerProvider,DocumentLoadInstrumentation,FetchInstrumentation, exporter, or sampler configuration. - Sampling, cardinality, and PII controls — use when sizing sampling rates against traffic/cost, naming attributes via Semantic Conventions, or auditing an existing pipeline for PII leakage.
Response minimum
Return, at minimum:
- the metric/trace component in scope and whether guidance concerns lab or field measurement,
- evidence level for any performance claim (lab-evidence, field-evidence with percentile, or documentation-based),
- instrumentation code or review findings with exact API options used,
- PII-in-telemetry check result for every attribute touched,
- sampling-rate rationale tied to stated traffic volume and cost constraints.
Files (vanguard-frontier-agentic)
-
references
-
opentelemetry-web-tracing-wiring.md 7.5 KB
# OpenTelemetry Web Tracing Wiring Use this reference only when wiring or reviewing `WebTracerProvider`, `DocumentLoadInstrumentation`/`DocumentLoad`, `FetchInstrumentation`, or exporter configuration for browser-origin traces. > Version note: `@opentelemetry/sdk-trace-web` and instrumentation package APIs (constructor option shapes, plugin package names) are version-sensitive. Verify against Context7-fetched docs (`/open-telemetry/opentelemetry-js`) and the installed package versions before prescribing exact constructor syntax. ## What people get wrong The naive story is: > "I'll create a `WebTracerProvider`, register it, and traces just work." Wrong. Official OpenTelemetry JS docs imply at least four separate concerns that must each be configured, not assumed: 1. **Context propagation across async boundaries** — the browser's default context manager does not automatically follow promises/timers/event callbacks; `ZoneContextManager` (from `@opentelemetry/context-zone`) is the documented pattern for supporting asynchronous operations in `provider.register()`. 2. **Instrumentation registration is separate from provider creation** — creating a `WebTracerProvider` does not, by itself, produce document-load or fetch/XHR spans. Instrumentations (`DocumentLoad`/`DocumentLoadInstrumentation`, `FetchInstrumentation`, `XMLHttpRequestInstrumentation`) must be explicitly registered via `registerInstrumentations({instrumentations: [...]})`. 3. **Exporter and span-processor choice affects both delivery reliability and payload timing** — `SimpleSpanProcessor` exports each span as it ends (useful for local `ConsoleSpanExporter` debugging); `BatchSpanProcessor` batches and is the documented production pattern for network exporters like OTLP. 4. **Trace-propagation format must match the backend** — the B3 propagator vs. W3C Trace Context are not interchangeable; using the example combination that doesn't match your collector/backend silently breaks distributed trace stitching. ## Officially grounded wiring shape Minimal document-load + console-export setup (from `@opentelemetry/sdk-trace-web` docs): ```javascript import { ConsoleSpanExporter, SimpleSpanProcessor, WebTracerProvider, } from '@opentelemetry/sdk-trace-web'; import { DocumentLoad } from '@opentelemetry/plugin-document-load'; import { ZoneContextManager } from '@opentelemetry/context-zone'; import { registerInstrumentations } from '@opentelemetry/instrumentation'; const provider = new WebTracerProvider({ spanProcessors: [new SimpleSpanProcessor(new ConsoleSpanExporter())], }); provider.register({ contextManager: new ZoneContextManager(), }); registerInstrumentations({ instrumentations: [new DocumentLoad()], }); ``` Production-shaped fetch instrumentation + OTLP export (batched): ```javascript import { BatchSpanProcessor, WebTracerProvider, } from '@opentelemetry/sdk-trace-web'; import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-http'; import { FetchInstrumentation } from '@opentelemetry/instrumentation-fetch'; import { ZoneContextManager } from '@opentelemetry/context-zone'; import { registerInstrumentations } from '@opentelemetry/instrumentation'; const exporter = new OTLPTraceExporter({ url: '<opentelemetry-collector-url>/v1/traces', headers: {}, concurrencyLimit: 10, }); const provider = new WebTracerProvider({ spanProcessors: [ new BatchSpanProcessor(exporter, { maxQueueSize: 100, maxExportBatchSize: 10, scheduledDelayMillis: 500, exportTimeoutMillis: 30000, }), ], }); provider.register({ contextManager: new ZoneContextManager(), }); registerInstrumentations({ instrumentations: [new FetchInstrumentation()], }); ``` Note the OTLP endpoint must end in `/v1/traces` per the exporter's documented contract. ## Non-negotiable design rules ### 1. Never ship `ConsoleSpanExporter`/`SimpleSpanProcessor` as the production pipeline These are debugging tools. Production browser tracing needs a network exporter (e.g., `OTLPTraceExporter`) paired with `BatchSpanProcessor` so span export is batched and rate-limited instead of firing one network request per span. ### 2. Register a context manager that supports async work Without `ZoneContextManager` (or an equivalent async-aware context manager), span parent/child relationships across `await`, `setTimeout`, and event-handler boundaries can silently break, producing disconnected traces that look correct individually but don't stitch into one distributed trace. ### 3. Treat sampling as a first-class config, not an afterthought A browser SDK with no sampler configured effectively samples at 100% by default in many setups — that is a cost and cardinality risk at scale. See `sampling-cardinality-pii-controls.md` for sizing `TraceIdRatioBasedSampler` against traffic volume. ### 4. Do not put secrets or bearer tokens directly in exporter headers as the default pattern If the collector endpoint requires auth, treat exporter `headers` the same way you would treat any credential-bearing config: source it from a managed secret/config mechanism appropriate to the deployment target, not hardcoded inline in client-shipped JS (client-shipped JS is fully visible to end users — never place a durable secret there regardless of source). ### 5. Match the propagation format to what the backend actually consumes Verify (via Context7 or the collector/backend's own docs) whether the destination expects W3C Trace Context or B3 headers before wiring a propagator; a mismatch silently drops trace correlation without an error. ## Minimal safe implementation flow 1. Confirm the target collector/backend endpoint and its expected trace-propagation header format. 2. Create `WebTracerProvider` with a `BatchSpanProcessor` wrapping the production exporter (OTLP or equivalent) — not `SimpleSpanProcessor`/`ConsoleSpanExporter`. 3. Register with `ZoneContextManager` (or confirmed async-aware equivalent) so cross-async-boundary spans stay connected. 4. Register only the instrumentations actually needed (`DocumentLoad`, `FetchInstrumentation`, `XMLHttpRequestInstrumentation`) — registering unused instrumentations increases span volume without benefit. 5. Configure sampling (see `sampling-cardinality-pii-controls.md`) before enabling in production, sized to stated traffic. 6. Name any custom span/attribute per OpenTelemetry Semantic Conventions (see `sampling-cardinality-pii-controls.md`) rather than inventing ad-hoc keys. 7. Run every attribute the instrumentation or custom spans emit through the PII check before enabling in production. ## High-risk assumptions to kill - "Creating the provider means tracing already works" — instrumentations must be separately registered. - "The default context manager is fine for async code" — it is not; use `ZoneContextManager` or a verified equivalent. - "No sampler configured means it's fine for now" — unconfigured sampling at high traffic is a cost/cardinality incident waiting to happen. - "Fetch instrumentation only captures URLs, so it's automatically safe" — captured URLs can include query strings carrying tokens/PII; review `FetchInstrumentation` URL-capture behavior against the PII check. ## When to push back Push back if the user asks for: - shipping `SimpleSpanProcessor` + a network exporter direct to production ("it's simpler"), - no sampler configuration for a stated high-traffic production app, - hardcoding a durable collector auth token into client-shipped exporter `headers`, - registering every available instrumentation "to be thorough" without a stated need. That is not "faster." It is a cost, correctness, and security liability. -
sampling-cardinality-pii-controls.md 7.9 KB
# Sampling, Cardinality, and PII Controls Use this reference when sizing a sampling rate against traffic/cost, naming attributes via OpenTelemetry Semantic Conventions, or auditing an existing RUM/tracing pipeline for PII leakage. ## What people get wrong The naive story is: > "Sampling is just a knob you turn down if the bill gets too high." That treats sampling as a cost afterthought instead of a design input, and it ignores that unsampled, high-cardinality, or PII-bearing telemetry is a *data-governance* problem before it is a *cost* problem — a leaked attribute doesn't get cheaper at a lower sample rate, it gets leaked less often. ## Non-negotiable design rules ### 1. Sample deterministically by trace ID, not randomly per-request `TraceIdRatioBasedSampler` (from `@opentelemetry/sdk-trace-web` / `@opentelemetry/sdk-trace-base`) samples a percentage of traces determined deterministically by the trace ID — any trace sampled at a given ratio is also sampled at any higher ratio. This is the documented building block for consistent, traffic-proportional sampling: ```javascript const { WebTracerProvider, ParentBasedSampler, TraceIdRatioBasedSampler, } = require('@opentelemetry/sdk-trace-web'); const provider = new WebTracerProvider({ sampler: new ParentBasedSampler({ root: new TraceIdRatioBasedSampler(0.1), // sample 10% of root traces }), }); ``` ### 2. Wrap it in `ParentBasedSampler` unless you have a specific reason not to `ParentBasedSampler` respects an incoming trace's sampling decision by default and only delegates the *root*-span sampling decision to the wrapped sampler (e.g., `TraceIdRatioBasedSampler`). This keeps a single distributed trace's sampling decision consistent across service boundaries — a browser span sampled independently of its downstream backend spans produces incomplete, misleading traces. ### 3. Size the ratio against stated traffic volume and the backend's cost/cardinality limits, not a guess Before recommending a ratio: - Get (or explicitly estimate and label as an estimate) the page's traffic volume (sessions/day or requests/day). - Get the backend's pricing/cardinality model (per-span cost, per-attribute cardinality limits, retention window). - Compute expected exported span volume at candidate ratios (e.g., 100k sessions/day × 5 spans/session × 0.1 ratio = 50k spans/day) and state that arithmetic explicitly in the recommendation — never hand back a bare percentage with no traffic math behind it. - Never recommend 100% (unsampled) export for a stated high-traffic production app without an explicit, logged cost review — treat that as the exception requiring justification, not the default. ### 4. Name attributes with OpenTelemetry Semantic Conventions, not ad-hoc keys Use the standard attribute namespaces (e.g., `http.request.method`, `http.response.status_code`, `url.path`, `user_agent.original`) documented in the OpenTelemetry Semantic Conventions rather than inventing project-specific keys like `httpMethod` or `req_status`. Ad-hoc naming breaks cross-service correlation and makes backend queries/dashboards non-portable. When no standard convention covers a genuinely custom business attribute, prefix it clearly (e.g., a documented internal namespace) so it's visually distinguishable from standard conventions, rather than colliding with or shadowing a standard key name. ### 5. Every custom attribute must pass a PII check before it ships Before adding any span attribute, metric label, or `web-vitals` attribution `generateTarget` output, check it against this list. If any apply, the attribute must be scrubbed, hashed, truncated, or dropped before export: - **Raw URLs with query strings** — query strings frequently carry session tokens, search terms, or user-identifying values. Capture path only (`url.path`), or explicitly strip known-sensitive query parameters before capturing `url.full`. - **User identifiers** — no raw email, username, account ID, or customer name as a plain attribute value. If correlation is needed, use an opaque, rotated, non-reversible identifier and document the mapping's access control separately. - **Free-text form input** — never capture form field values (search boxes, comment fields, addresses) as telemetry attributes. - **Precise geolocation** — no raw lat/long or street-level location; use coarse-grained (e.g., country/region) if location is genuinely needed for the stated purpose. - **DOM-derived strings that embed the above** — a CSS selector or `generateTarget` output can accidentally embed a customer name or ID if it's built from live `id`/`class`/`data-*` attributes that the DOM populates with user data; audit before trusting default stringification (see `web-vitals-attribution-wiring.md` rule 1). - **Auth headers/tokens** — never mirror `Authorization`, cookies, or bearer tokens into span attributes, even truncated; treat any such capture as a credential leak, not a debugging convenience. ### 6. Cardinality is a cost multiplier independent of sampling A single high-cardinality attribute (e.g., a raw user ID or full URL used as a span attribute) can multiply backend storage/indexing cost regardless of the trace sampling ratio, because most backends index per-unique-attribute-value. Prefer bounded-cardinality attributes (enum-like values, coarse buckets) for anything intended to be an indexed/filterable dimension; keep genuinely high-cardinality values (if needed at all) as non-indexed payload fields, and confirm the plan against the destination backend's own cardinality documentation. ## Minimal safe implementation flow 1. Get or explicitly estimate traffic volume for the page/app in scope. 2. Get the backend's pricing/cardinality/retention model (or state clearly that it's unknown and flag the gap). 3. Propose a `TraceIdRatioBasedSampler` ratio wrapped in `ParentBasedSampler`, with the traffic-volume arithmetic shown. 4. Enumerate every custom span attribute, metric label, and `generateTarget` output planned; run each through the PII check in rule 5. 5. Map custom attributes to OpenTelemetry Semantic Conventions where a standard key exists; namespace the rest clearly. 6. Flag any attribute that is both high-cardinality and unbounded (raw user input, raw URL) as an indexing/cost risk even if it passes the PII check. 7. State the sampling ratio, PII findings, and cardinality findings together as one review output — do not split cost sizing from privacy review into separate, disconnected passes. ## Adversarial checklist Before approving a RUM/tracing pipeline, answer these: - What is the actual (or explicitly estimated) traffic volume this sampling ratio is sized against? - Does the sampler respect parent sampling decisions (`ParentBasedSampler`), or will browser spans and backend spans disagree on what got sampled? - For every custom attribute: does it match a standard Semantic Convention key, or is it deliberately namespaced as custom? - For every attribute sourced from user-controllable input (URL, form, DOM content, `generateTarget`): has it been explicitly scrubbed, or is it being trusted "because it looked fine in testing"? - Which attributes are indexed/filterable in the destination backend, and are any of those unbounded-cardinality values? If these cannot be answered, the review is incomplete — say so rather than approving. ## When to push back Push back if the user asks for: - 100% unsampled export "to not miss anything," on a stated high-traffic app, with no cost review, - capturing full request URLs or form values "just in case it's useful later," - inventing custom attribute names that duplicate an existing Semantic Convention key under a different spelling, - skipping the PII check because "it's just internal telemetry" — internal-only access does not eliminate the need for a PII/data-governance review; it changes who is affected, not whether it matters. Those are not shortcuts. They are cost incidents and data-governance failures waiting to be found in an audit. -
web-vitals-attribution-wiring.md 6.6 KB
# web-vitals Attribution Build Wiring Use this reference only when writing or reviewing the actual `onLCP`/`onINP`/`onCLS` (or `onFCP`/`onTTFB`) instrumentation call. > Version note: `web-vitals` API surface (attribution option shapes, default thresholds) is version-sensitive. Verify the installed package version against Context7-fetched docs (`/googlechrome/web-vitals`) before prescribing option names. ## What people get wrong The naive story is: > "I'll import `onLCP` from `web-vitals` and send everything to analytics with `reportAllChanges: true` so I never miss a value." Wrong on two counts: 1. The **standard build** (`web-vitals`) only returns the metric name, value, rating, and delta. It does **not** tell you *why* LCP was slow or *which* element shifted for CLS. For root-cause work you need the **attribution build** (`web-vitals/attribution`), which is a separate import path, not a flag on the standard build. 2. `reportAllChanges: true` reports every intermediate value change (e.g., every CLS shift, every LCP candidate), not just the final value — this is a debugging tool that multiplies event volume in production, not a default-on setting. ## Standard build vs. attribution build - **Standard build** (`import {onLCP, onINP, onCLS} from 'web-vitals'`): metric name, `value`, `rating` (`good`/`needs-improvement`/`poor`), `delta`, `id`. Use this for default production RUM export — it is the lower-cardinality, lower-payload choice. - **Attribution build** (`import {onLCP, onINP, onCLS} from 'web-vitals/attribution'`): adds a `metric.attribution` object with phase/target detail (e.g., LCP's `attribution.target`, `attribution.timeToFirstByte`, `attribution.resourceLoadDelay`; INP's `attribution.interactionTarget`, `attribution.inputDelay`, `attribution.processingDuration`, `attribution.presentationDelay`; CLS's `attribution.largestShiftTarget`, `attribution.largestShiftTime`). - Do not send the attribution build's full diagnostic payload to a blanket production analytics endpoint by default — route it to a debug/diagnostic destination or sample it separately, because its per-metric fields substantially increase payload size and can carry more surface area for accidental PII (e.g., a raw element attribute pulled into a target string). ## Non-negotiable design rules ### 1. Choose `generateTarget` deliberately, not by default `AttributionReportOpts.generateTarget` is `(el: Node | null) => string | undefined`. It overrides how a DOM node is stringified in attribution fields (e.g., `attribution.target`, `attribution.largestShiftTarget`, `attribution.interactionTarget`). Do not leave the default CSS-selector-based stringification unexamined if the DOM contains sensitive `id`/`class`/`data-*` values that could leak into telemetry (e.g., a selector embedding a customer name). A safe pattern prioritizes an intentional tracking attribute over the raw selector: ```javascript import {onCLS, onINP, onLCP} from 'web-vitals/attribution'; function generateTarget(el) { // Prefer an explicit, reviewed tracking attribute. if (el?.dataset.trackingId) { return el.dataset.trackingId; } // Fall back to default CSS-selector logic only if that selector // is known not to embed PII (audit before relying on this). return undefined; } onLCP((metric) => sendToAnalytics(metric), {generateTarget}); onCLS((metric) => sendToAnalytics(metric), {generateTarget}); onINP((metric) => sendToAnalytics(metric), {generateTarget}); ``` ### 2. Tune `INPAttributionReportOpts` explicitly, don't accept silent defaults `onINP` in the attribution build accepts `durationThreshold` (default `40`ms — interactions shorter than this are not attributed) and `includeProcessedEventEntries` (default `true` — includes the full array of event-timing entries for the interaction frame). State both defaults explicitly in any review: a lower `durationThreshold` increases event volume by attributing more interactions; leaving `includeProcessedEventEntries: true` on a high-traffic page increases per-event payload size and should be weighed against export cost. ### 3. Keep `reportAllChanges` off in production by default `reportAllChanges` (inherited from `ReportOpts`, default `false`) reports every intermediate metric change instead of just the final value. Only enable it for a scoped debugging session against a non-production or sampled destination — never as the default production wiring. ### 4. Match the reporting transport to the page-lifecycle guarantee you need `web-vitals` callbacks can fire late in the page lifecycle (e.g., on visibility change or page unload for LCP/CLS finalization). Use a transport that survives page unload (e.g., `navigator.sendBeacon`, or `fetch` with `keepalive: true`) rather than a plain `fetch` call that can be cancelled when the page unloads mid-request. ## Minimal safe implementation flow 1. Decide standard vs. attribution build based on whether root-cause detail is needed for this rollout (default: standard build for blanket production RUM). 2. If attribution build: define and review a `generateTarget` function before shipping — do not accept the default selector logic unreviewed. 3. Explicitly set (or explicitly accept and document) `durationThreshold`/`includeProcessedEventEntries` for INP. 4. Leave `reportAllChanges` at its default `false` unless a scoped debug session requires it, and route that debug payload separately. 5. Choose an unload-safe transport (`sendBeacon` / `fetch` with `keepalive`). 6. Run every attribution field emitted through the PII check in `sampling-cardinality-pii-controls.md` before it ships. ## Adversarial checklist Before shipping any `web-vitals` wiring, answer these: - Which build (standard or attribution) is this, and does that match the stated purpose (production dashboard vs. active debugging)? - If attribution build: has `generateTarget` been reviewed against the actual DOM, or is the default selector logic unaudited? - Is `reportAllChanges` on, and if so, is that intentional and scoped to non-production/debug traffic? - What percentile will this data be aggregated to before it's called "the site's LCP/INP/CLS" (per CWV convention, p75)? - Does the reporting transport survive `visibilitychange`/unload, or will late-finalizing metrics be silently dropped? ## When to push back Push back if the user asks for: - shipping the attribution build's full payload to production analytics "just in case," with no PII review of `generateTarget` output, - `reportAllChanges: true` as a permanent production default, - reporting a single lab run's LCP/INP/CLS value as if it were the field p75 the org will be judged on. Those are not shortcuts. They inflate cost, leak data, and misreport what users actually experience.
-
-
metadata.json 1.2 KB
{ "id": "frontend-observability-rum-instrumentation", "name": "Frontend Observability RUM Instrumentation", "type": "skill", "provider": "frontend", "harnesses": [ "claude-code", "cursor", "codex", "gemini", "kiro", "other" ], "summary": "Skill for designing and reviewing Real User Monitoring instrumentation using the web-vitals attribution build and OpenTelemetry Web, with explicit sampling, cardinality, and PII-in-telemetry controls and lab-vs-field evidence labeling.", "source_type": "adapted", "official_docs": [ "https://web.dev/articles/vitals", "https://github.com/GoogleChrome/web-vitals", "https://opentelemetry.io/docs/languages/js/", "https://opentelemetry.io/docs/specs/semconv/" ], "security_notes": "Never recommends capturing full URLs with query strings, user identifiers, or free-text form values as span/metric attributes without explicit scrubbing. Read-only review of existing instrumentation; does not deploy telemetry config changes to production without explicit human sign-off logged outside this skill.", "last_verified": "2026-07-02", "path": "skills/frontend/frontend-observability-rum-instrumentation", "author": "github: VincentChuWaiChow", "version": "0.1.0" } -
SKILL.md 7.4 KB
--- name: frontend-observability-rum-instrumentation description: Design or review browser-side Real User Monitoring instrumentation for Core Web Vitals (LCP, INP, CLS) using the web-vitals attribution build and distributed tracing via OpenTelemetry Web, enforcing lab-vs-field evidence labeling, sampling/cardinality sizing, and PII-in-telemetry controls, with library-specific wiring references loaded only when instrumentation code is actually being written or reviewed. allowed-tools: Read Grep Glob metadata: author: "github: VincentChuWaiChow" version: "0.1.0" updated: "2026-07-02" category: observability --- # Frontend Observability RUM Instrumentation ## Purpose Lab-only performance testing (Lighthouse, synthetic CI runs) systematically misses the device/network diversity of real users, and ad-hoc telemetry instrumentation routinely leaks PII into span attributes or blows up observability cost through unsampled high-cardinality export. This skill wires and reviews field RUM (web-vitals + OpenTelemetry Web) with explicit evidence labeling and privacy/cost guardrails baked in from the start. ## When to use Use this skill when the user asks to: - instrument Core Web Vitals (LCP, INP, CLS) field measurement in a web app, - set up or review OpenTelemetry Web browser tracing (document load, fetch/XHR spans), - review an existing RUM/telemetry pipeline for PII exposure or sampling/cost issues, - interpret or report on Core Web Vitals data and needs a lab-vs-field distinction made explicit, - set or validate a performance budget against field p75 data. ## When NOT to use - Diagnosing *why* a specific LCP/INP/CLS number is high (phase decomposition of an existing regression) — hand off to `core-web-vitals-triage`; this skill instruments the pipeline that produces the data, it does not decompose an already-captured regression. - Bundle size, code-splitting, or caching remediation once a cause is known — hand off to `bundle-budget-code-splitting-review` or `service-worker-cache-strategy-review`. - Server-side or backend-service tracing with no browser-origin span — this skill is scoped to browser-side RUM and OpenTelemetry Web only. ## Context7 Documentation Protocol `web-vitals` and OpenTelemetry Web packages ship breaking API changes across major versions (attribution-build option shapes, sampler class names, exporter constructor options), and memorized snippets go stale. Before writing or reviewing instrumentation code: 1. Call `ToolSearch` with query `"context7"` (or `"select:mcp__Context7__resolve-library-id,mcp__Context7__query-docs"`) to load the Context7 tools if not already loaded this session. 2. Call `mcp__Context7__resolve-library-id` for `web-vitals` and/or `opentelemetry-js` before making any claim about their current API shape. Verified IDs as of this skill's `updated` date: `/googlechrome/web-vitals` and `/open-telemetry/opentelemetry-js`. 3. Call `mcp__Context7__query-docs` for the specific mechanism in scope — e.g. "attribution build generateTarget option", "INPAttributionReportOpts durationThreshold", "WebTracerProvider spanProcessors config", "TraceIdRatioBasedSampler" — before ruling on it. Do this per review; do not reuse a prior session's memory of these APIs. 4. Known facts verified via Context7 as of this skill's `updated` date: the attribution build (`web-vitals/attribution`) accepts `AttributionReportOpts` (`reportAllChanges`, `generateTarget`) on `onCLS`/`onFCP`/`onLCP`/`onTTFB`, and `INPAttributionReportOpts` additionally exposes `durationThreshold` (default `40`) and `includeProcessedEventEntries` (default `true`) on `onINP`. `WebTracerProvider` (from `@opentelemetry/sdk-trace-web`) takes a `spanProcessors` array in its constructor — do not tell a user to call a separate `addSpanProcessor` method as the primary pattern without checking the installed SDK version. `TraceIdRatioBasedSampler` takes a single ratio argument (0–1) and is normally wrapped in `ParentBasedSampler({ root: ... })` so that downstream/parent sampling decisions are respected. 5. If Context7 is unavailable or returns no relevant match, fall back to `official_docs` / `references/*.md` and mark the claim `documentation-based (Context7 unavailable)` rather than presenting it as freshly verified. 6. Never invent a `web-vitals` metric field, attribution property, OpenTelemetry exporter option, or sampler class that no queried source confirms. ## Lean operating rules - Always label every performance number as lab evidence (Lighthouse/CI/synthetic) or field evidence (RUM, real users — state the percentile; p75 is the CWV standard) — never let the two blur into an unlabeled "the site is fast" claim. - Default to the standard `web-vitals` build (`onLCP`/`onINP`/`onCLS`) for production reporting; only use the `/attribution` build's richer diagnostic payload when actively debugging a specific regression, and only send that extra payload to a dev/debug destination, not blanket production export. - Do not enable `reportAllChanges: true` in production by default — it multiplies event volume for a debugging-only benefit. - Before adding any custom span/metric attribute, run it through a PII check: no raw URLs with query strings, no user IDs, no free-text form input, no precise geolocation, unless explicitly scrubbed and justified. - Size the OpenTelemetry sampling rate (e.g. `TraceIdRatioBasedSampler` wrapped in `ParentBasedSampler`) against the stated/estimated traffic volume and the backend's cost/cardinality limits; never recommend 100% unsampled export for a high-traffic app without an explicit cost review. - Use OpenTelemetry Semantic Conventions attribute names instead of inventing ad-hoc span/attribute naming, so telemetry stays queryable/comparable across services. - Treat any recommended performance budget threshold as grounded in web.dev's documented "good" ranges (LCP ≤2.5s, INP ≤200ms, CLS ≤0.1 at p75) unless the org has an explicitly stricter documented SLO. - This skill performs static review and instrumentation-code authoring only; it does not deploy telemetry configuration to production or flip sampling/export settings on a live collector without explicit human sign-off logged outside this skill. ## References Load these only when needed: - [web-vitals attribution build wiring](references/web-vitals-attribution-wiring.md) — use when writing or reviewing the actual `onLCP`/`onINP`/`onCLS` instrumentation call, choosing between standard and attribution builds, or setting `generateTarget`/`durationThreshold`/`reportAllChanges`. - [OpenTelemetry Web tracing wiring](references/opentelemetry-web-tracing-wiring.md) — use only when wiring or reviewing `WebTracerProvider`, `DocumentLoadInstrumentation`, `FetchInstrumentation`, exporter, or sampler configuration. - [Sampling, cardinality, and PII controls](references/sampling-cardinality-pii-controls.md) — use when sizing sampling rates against traffic/cost, naming attributes via Semantic Conventions, or auditing an existing pipeline for PII leakage. ## Response minimum Return, at minimum: - the metric/trace component in scope and whether guidance concerns lab or field measurement, - evidence level for any performance claim (lab-evidence, field-evidence with percentile, or documentation-based), - instrumentation code or review findings with exact API options used, - PII-in-telemetry check result for every attribute touched, - sampling-rate rationale tied to stated traffic volume and cost constraints.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.