e2e
Prove real user journeys across browser, mobile, desktop, API and CLI surfaces using the project's existing test tools. Use for end-to-end verification, runtime UI review or requested visual proof.
Install
npx skills add https://github.com/sgaabdu4/building-flutter-apps/tree/main/.agents/skills/e2e
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install sgaabdu4-building-flutter-apps@llmmart
git clone https://github.com/sgaabdu4/building-flutter-apps.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole sgaabdu4/building-flutter-apps collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
E2E
Routes
Select by actual surface + explicit browser/device requirements. Reuse the project's runner and fixtures; missing runtime/connection → identify the needed capability before setup.
flowchart LR
S{Surface} -->|Web UI| B[Existing browser E2E]
B -->|Recorded proof needed| V[Product Walkthrough: recorded E2E]
S -->|Polished web video requested| D[Product Walkthrough: video delivery]
S -->|Flutter + Riverpod| F[Building Flutter Apps]
S -->|Other Flutter| I[Existing Flutter integration + device tools]
F -->|Recorded proof needed| G[Recorded proof: references/flutter.md]
I -->|Recorded proof needed| G
S -->|Native / React Native / desktop| N[Platform + device runner]
S -->|API / worker / CLI / pure Dart| A[Real request / event / command boundary]
click V "../product-walkthrough-video/SKILL.md"
click D "../product-walkthrough-video/SKILL.md"
click F "../building-flutter-apps/SKILL.md"
click G "references/flutter.md"
OS-owned dialogs → platform control beyond the app tree. Browser exploration → durable regressions in existing tests. API/CLI journeys need UI runtime only when the journey includes it.
Prove the journey
- Establish requested behavior, environment/build, actors, permissions and starting data. Exercise the real user path with stable semantic selectors and observable state waits.
- Choose the smallest cases covering the changed behavior and meaningful failures. Shared state needs independent writer/observer proof; persisted changes need source-of-truth readback. Cover denied/revoked access, retries or relaunch when the requirement depends on them.
- Keep cases isolated and repeatable using existing fixtures and cleanup. Do not substitute a mock, direct API mutation or test-only shortcut for the interaction being proved.
- Assert the visible result and relevant durable effects. A completed click, successful request, clean log or zero exit alone is insufficient. Investigate unexpected app/native/network errors and retries that only pass intermittently.
- On failure, preserve the reproduction and useful evidence; distinguish product, fixture, runner and capture defects. With fix authorization, correct the owner, rerun the original case and affected downstream cases, and retain a meaningful regression in the existing suite. Completion requires the corrected journey + affected checks to pass; otherwise report the exact blocker and remaining proof. Retries or weaker assertions do not resolve a defect.
- A user-reported defect reopens the affected journey's verification. Reproduce it, strengthen the assertion or review that missed it, and repeat the fix/retest loop; previous passing evidence cannot close the new report.
Visual proof and completion
- Capture the smallest useful evidence set. Ordinary regression work does not require video. When screenshots or video are requested or needed, inspect the actual delivered media and confirm its subject, required steps and final state.
- Recorded proof follows the selected route's owner; backend readback + repeatable journey assertions stay here and media checks alone cannot prove acceptance. Captures outside an owned pipeline → existing artifacts + direct inspection; do not invent a compatible report.
- Keep secrets and personal data out of artifacts. Show requested evidence to the user; treat test artifacts as local unless their inclusion as repository assets is authorized.
- Report tested surfaces, outcomes and exact gaps separately. Assertions, persisted state, deployment identity and visual evidence prove different things. An unavailable device, account or unreviewed artifact remains unproven.
Files (building-flutter-apps)
-
agents
-
openai.yaml 278 B
interface: display_name: "E2E" short_description: "Prove user journeys with the existing stack tools" default_prompt: "Use $e2e to verify the requested journey through the real product and report its outcomes and remaining gaps." policy: allow_implicit_invocation: true
-
-
references
-
flutter.md 4.8 KB
# Flutter recorded proof Load when a Flutter journey needs recorded video evidence. Selector, configuration and entrypoint rules (rule 12 `lib/main_dev.dart`, optional runtime interaction, release/native limits) stay in [Flutter runtime E2E](../../building-flutter-apps/references/dart-mcp-e2e-testing.md). Journey assertions, durable-state readback and evidence reporting stay in [E2E](../SKILL.md#prove-the-journey). Marionette taps and reads the running app; it never records. The recorder is a separate OS process capturing the device screen while Marionette drives. ```mermaid flowchart TD R[marionette server registered] --> A[Binding in main] A --> L[flutter run on the booted device] L --> C[connect with the ws VM service URI] C --> S[Start OS recorder in background] S --> D[Drive + assert] D --> X[SIGINT recorder; wait for exit] X --> V[ffprobe + extracted frames; inspect] ``` ## Server registration - Trigger = any Flutter app (a `flutter` SDK dependency in `pubspec.yaml`). Hard Eng setup/update registers the `marionette` server as `dart run marionette_mcp@<pubspec.lock version>` when the lock records `marionette_flutter`, otherwise `dart run marionette_mcp@` (latest). No `dart pub global activate`. Registration alone drives nothing: the app binding below is still required. - Detection runs only when setup installs or an update applies a newer verified revision; existing MCP entries are preserved, so a written pin does not follow a later lock upgrade. After adding or upgrading `marionette_flutter`, edit the `marionette` entry in `.mcp.json` (and `.codex/config.toml`) by hand to the locked version. - Servers load at session start → start a new agent session after registration. Absent tools in the registering session are not a failed install. ## App binding - `marionette_flutter` is a regular dependency, not a dev dependency: `lib/main.dart` imports it. Upstream 0.6.0 requires Flutter ≥ 3.27. Binding and server versions must match; a mismatch surfaces at `connect`. - `MarionetteBinding.ensureInitialized()` = first line of `main()`. A test calling `main()` must not create a second binding, so guard on both `kDebugMode` and `FLUTTER_TEST`: ```dart final isFlutterTest = Platform.environment.containsKey('FLUTTER_TEST'); if (kDebugMode && !isFlutterTest) { MarionetteBinding.ensureInitialized(); } else { WidgetsFlutterBinding.ensureInitialized(); } ``` - It must run before `SentryFlutter.init`, not inside its `appRunner`: Sentry claims the binding first and its zone swallows the resulting error, so the app hangs on the splash screen with no exception or log. - Debug and profile builds only; the VM service does not exist in release. `main()` does not re-run on hot reload → hot restart after adding the binding. ## Connect + drive - `flutter run` on the booted simulator/emulator, take the VM service URI from its output and pass the `ws://.../ws` form to `connect`. `connect` precedes every other tool. - Drive with `get_interactive_elements` → `tap` (prefer `key`, then `identifier`, then `text`/`type`/`coordinates`), `enter_text`, `press_key`, `swipe`, `scroll_to`; read with `take_screenshots` and `get_logs` (needs a configured collector); `hot_reload`/`hot_restart` after code changes. - On iOS/Android the platform keyboard owns field editing → change a field's value with `enter_text`, not `press_key`. ## Record the screen Start the recorder in the background before driving and keep the printed pid; shell state does not survive between commands. iOS simulator: ```sh xcrun simctl io booted recordVideo --codec=h264 --force journey.mp4 & echo $! ``` Android emulator (documented, not verified here): `adb shell screenrecord --time-limit 180 /sdcard/journey.mp4 &`, then `adb pull /sdcard/journey.mp4` after the stop below. 180 seconds is the default cap. Stop with SIGINT and wait for the process to exit before reading the file; it writes the container while shutting down and prints `Recording completed. Writing to disk.` ```sh kill -INT <pid> while kill -0 <pid> 2>/dev/null; do sleep 0.2; done ``` ## Verify the media ```sh ffprobe -v error -select_streams v:0 -count_frames \ -show_entries stream=codec_name,width,height,nb_read_frames:format=duration \ -of default=nw=1 journey.mp4 ``` - `duration` must cover the drive and `nb_read_frames` must exceed 1. - The simulator recorder emits a frame only when the screen changes: a six-second capture of a static screen yields 1 frame and ~0.07s duration. That means nothing moved, not a broken recorder — fix the drive, not the recording command. On a single-frame file both extracted frames are that one frame. - Extract and actually view both ends, plus the video itself; record what was seen as the evidence [E2E](../SKILL.md#visual-proof-and-completion) requires. ```sh ffmpeg -v error -i journey.mp4 -frames:v 1 first.png ffmpeg -v error -sseof -0.5 -i journey.mp4 -frames:v 1 last.png ```
-
-
SKILL.md 3.8 KB
--- name: e2e description: Prove real user journeys across browser, mobile, desktop, API and CLI surfaces using the project's existing test tools. Use for end-to-end verification, runtime UI review or requested visual proof. --- # E2E ## Routes Select by actual surface + explicit browser/device requirements. Reuse the project's runner and fixtures; missing runtime/connection → identify the needed capability before setup. ```mermaid flowchart LR S{Surface} -->|Web UI| B[Existing browser E2E] B -->|Recorded proof needed| V[Product Walkthrough: recorded E2E] S -->|Polished web video requested| D[Product Walkthrough: video delivery] S -->|Flutter + Riverpod| F[Building Flutter Apps] S -->|Other Flutter| I[Existing Flutter integration + device tools] F -->|Recorded proof needed| G[Recorded proof: references/flutter.md] I -->|Recorded proof needed| G S -->|Native / React Native / desktop| N[Platform + device runner] S -->|API / worker / CLI / pure Dart| A[Real request / event / command boundary] click V "../product-walkthrough-video/SKILL.md" click D "../product-walkthrough-video/SKILL.md" click F "../building-flutter-apps/SKILL.md" click G "references/flutter.md" ``` OS-owned dialogs → platform control beyond the app tree. Browser exploration → durable regressions in existing tests. API/CLI journeys need UI runtime only when the journey includes it. ## Prove the journey - Establish requested behavior, environment/build, actors, permissions and starting data. Exercise the real user path with stable semantic selectors and observable state waits. - Choose the smallest cases covering the changed behavior and meaningful failures. Shared state needs independent writer/observer proof; persisted changes need source-of-truth readback. Cover denied/revoked access, retries or relaunch when the requirement depends on them. - Keep cases isolated and repeatable using existing fixtures and cleanup. Do not substitute a mock, direct API mutation or test-only shortcut for the interaction being proved. - Assert the visible result and relevant durable effects. A completed click, successful request, clean log or zero exit alone is insufficient. Investigate unexpected app/native/network errors and retries that only pass intermittently. - On failure, preserve the reproduction and useful evidence; distinguish product, fixture, runner and capture defects. With fix authorization, correct the owner, rerun the original case and affected downstream cases, and retain a meaningful regression in the existing suite. Completion requires the corrected journey + affected checks to pass; otherwise report the exact blocker and remaining proof. Retries or weaker assertions do not resolve a defect. - A user-reported defect reopens the affected journey's verification. Reproduce it, strengthen the assertion or review that missed it, and repeat the fix/retest loop; previous passing evidence cannot close the new report. ## Visual proof and completion - Capture the smallest useful evidence set. Ordinary regression work does not require video. When screenshots or video are requested or needed, inspect the actual delivered media and confirm its subject, required steps and final state. - Recorded proof follows the selected route's owner; backend readback + repeatable journey assertions stay here and media checks alone cannot prove acceptance. Captures outside an owned pipeline → existing artifacts + direct inspection; do not invent a compatible report. - Keep secrets and personal data out of artifacts. Show requested evidence to the user; treat test artifacts as local unless their inclusion as repository assets is authorized. - Report tested surfaces, outcomes and exact gaps separately. Assertions, persisted state, deployment identity and visual evidence prove different things. An unavailable device, account or unreviewed artifact remains unproven.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.