lean-ux
Apply lean thinking to UX: hypothesis-driven design, collaborative sketching, and rapid experiments instead of heavy deliverables. Use when the user mentions "Lean UX", "design hypothesis", "outcome over output", "design studio method", "assumption mapping", "lightweight research
Install
npx skills add https://github.com/wondelai/skills/tree/main/plugins/wondelai-skills/skills/lean-ux
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install wondelai-skills@llmmart
git clone https://github.com/wondelai/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole wondelai/skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Lean UX Framework
A practice-driven approach to UX that replaces heavy deliverables with rapid experimentation, cross-functional collaboration, and continuous learning. Lean UX shifts the question from "What should we design?" to "What do we need to learn?"
Core Principle
Outcomes over outputs. The value of a design is measured not by the fidelity of the deliverable but by the change in user behavior it produces.
The foundation: Traditional UX waterfalls requirements into wireframes, mockups, specs, and code—losing context and hiding untested assumptions at every handoff. Lean UX compresses the distance between idea and evidence: declare assumptions, form hypotheses, run the smallest possible experiment, and let real user behavior settle the argument. Shared understanding replaces documentation; learning velocity replaces pixel perfection.
Scoring
Goal: 10/10. Score a UX process, design plan, or team workflow by the eight-row Quick Diagnostic below: award ~1.25 points per row answered "yes" (8 yeses = 10). Bands:
- 9-10 — assumptions declared, hypotheses with pre-committed success criteria, lowest-fidelity experiments, whole-team design, weekly research, outcome (not output) metrics, dual-track agile, and a recently invalidated hypothesis on the books.
- 5-6 — hypotheses exist but criteria are vague or fidelity is over-invested; design and research still partly siloed.
- <=3 — heavy deliverables, untested assumptions, output-counting, no experiment log.
Always state the current score, the diagnostic rows that failed, and the specific fix for each.
Framework
1. Declaring Assumptions
Core concept: Every design starts with assumptions. Lean UX makes them explicit so they can be prioritized and tested, rather than baked invisibly into specifications.
Why it works: Unspoken assumptions mean teams build on shaky ground and discover problems only after launch; surfacing them early focuses energy on the riskiest ones and reduces the cost of being wrong.
Key insights:
- Business assumptions define what must be true for the business (revenue model, market size, willingness to pay); user assumptions define who users are and how they behave
- Prioritize on two axes: risk (how damaging if wrong) and uncertainty (how little we know)
- Test high-risk, high-uncertainty assumptions first
- Write assumptions collaboratively as a team, not in isolation
Product applications:
| Context | Application | Example |
|---|---|---|
| New feature kick-off | Assumption mapping workshop | "We assume users want to share reports with teammates" |
| Roadmap planning | Rank features by assumption risk | Prioritize features whose success depends on untested beliefs |
| Stakeholder alignment | Expose hidden assumptions across roles | PM assumes pricing works; engineer assumes scale; designer assumes flow |
Ethical boundary: Assumptions must be honest assessments, not post-hoc justifications—if leadership has already committed to a direction, acknowledge the constraint rather than pretending it's open to falsification.
See references/hypothesis-canvas.md when running an assumption workshop or writing a hypothesis — the risk/uncertainty prioritization matrix, business-vs-user assumption split, and fillable hypothesis and sub-hypothesis templates.
2. Hypothesis Statements
Core concept: A hypothesis translates an assumption into a testable prediction, linking a proposed change to a measurable outcome for a specific user segment.
Why it works: Hypotheses force precision—instead of "make onboarding better," the team commits to a prediction that can be proven or disproven, which prevents scope creep and makes the learn step unambiguous.
Key insights:
- Standard format: "We believe [outcome] will happen if [persona] achieves [action] with [feature]"
- Every hypothesis specifies persona, action, outcome, and measurable signal
- Sub-hypotheses break a large bet into independently testable parts
- Agree on what "validated" and "invalidated" look like before running the experiment
Product applications:
| Context | Application | Example |
|---|---|---|
| Feature design | Write hypothesis before wireframing | "We believe trial-to-paid conversion will rise 10% if new users complete a guided setup wizard" |
| A/B tests | Formalize test rationale | "We believe click-through will rise 15% if we move the CTA above the fold" |
| Sprint planning | Attach hypothesis to each story | Story: "filter by date." Hypothesis: "task completion time drops 30%" |
Ethical boundary: Never cherry-pick metrics after the fact to declare a hypothesis validated—pre-commit to success criteria.
See references/outcome-metrics.md when picking the measurable signal for a hypothesis or defining team success — outcomes-vs-outputs, leading-vs-lagging indicator pairs, UX OKRs, and the vanity metrics to avoid.
3. MVPs and Experiments
Core concept: An MVP in Lean UX is the smallest design artifact that can test a hypothesis with real users—a learning tool, not a product launch.
Why it works: A paper prototype tested with five users in a hallway can invalidate a hypothesis that would otherwise consume a full engineering sprint; matching experiment fidelity to assumption risk maximizes learning per unit of effort.
Key insights:
- Experiments range from low fidelity (paper prototypes, concierge tests) to high fidelity (coded A/B tests, Wizard of Oz)
- Choose the lowest-fidelity experiment that can answer the question
- A good experiment has a clear hypothesis, defined audience, measurable signal, and time box
- Proto-personas can stand in for full research when speed matters, but must be validated later
Product applications:
| Context | Application | Example |
|---|---|---|
| Early concept validation | Paper prototype or clickable mockup | Sketch 3 concepts, test with 5 users same day |
| Demand validation | Landing page smoke test | "Sign up for early access" measures real interest |
| Usability validation | Clickable prototype test | Figma prototype tested with 5-8 users |
| Pricing validation | Painted door test | Show pricing page, measure click-through before building billing |
Ethical boundary: Smoke tests and fake doors must not mislead users into believing a product exists—disclose test status and offer an opt-out.
See references/experiment-patterns.md when choosing or designing an experiment — the full catalog of experiment types with when/when-NOT-to-run notes, the experiment selection matrix and fidelity ladder, and a design template.
4. Collaborative Design
Core concept: Design is a team sport. Lean UX replaces the solitary designer-then-handoff model with cross-functional sessions where developers, PMs, and designers sketch solutions together.
Why it works: Developers who helped sketch the solution don't need a 40-page spec to build it—shared understanding replaces documentation, diverse perspectives generate more creative solutions, and handoff waste drops dramatically.
Key insights:
- Design Studio method: diverge (individual sketching), present, critique, converge (refined sketch), iterate
- The goal is informed commitment, not consensus: the team agrees on what to test, not what is "right"
- Cross-functional means engineers, QA, data analysts, and stakeholders sketch too
- Style guides and pattern libraries are living documents; reduce deliverables to the minimum needed for shared understanding (often a whiteboard photo)
Product applications:
| Context | Application | Example |
|---|---|---|
| Sprint kick-off | Design Studio session (90 minutes) | Whole team sketches solutions to the sprint's hypothesis |
| Feature exploration | Collaborative sketching workshop | 6-up sketches: each person draws 6 ideas in 5 minutes |
| Remote teams | Virtual whiteboard sessions | FigJam or Miro board with timed sketch rounds |
Ethical boundary: Collaboration must not become design by committee—a designated designer synthesizes input; the team does not vote on pixels.
See references/collaborative-design.md when facilitating a Design Studio — the step-by-step workshop protocol (timings, materials, remote variants) and how to keep style guides as living documents.
5. Feedback and Research
Core concept: Continuous, lightweight research replaces big-bang usability studies—small research activities embedded in every sprint instead of quarterly reports.
Why it works: Findings only change a decision while it is still cheap to reverse, so research value decays with every sprint between learning and the decision it informs; small weekly studies keep that gap near zero, which a quarterly report never can.
Key insights:
- Research types: usability tests, customer interviews, A/B tests, analytics review, surveys, diary studies
- Five users uncover approximately 85% of usability problems (Nielsen)
- Continuous cadence: recruit weekly, test weekly, synthesize weekly
- The whole team should observe at least some sessions to build empathy
- Proto-personas are refined and eventually replaced by evidence-based personas
Product applications:
| Context | Application | Example |
|---|---|---|
| Weekly usability testing | Test prototype with 3-5 users every Thursday | "Testing Thursday" ritual with rotating facilitators |
| Post-launch learning | Monitor analytics + 3 follow-up interviews | Find drop-off points, interview churned users |
| Persona validation | Compare proto-persona assumptions to interview data | "We assumed power users are marketers; data shows ops managers" |
Ethical boundary: Conduct research with informed consent—participants should understand how their data is used and be free to withdraw.
6. Integration with Agile
Core concept: Lean UX works inside Agile via dual-track development: discovery (learning what to build) and delivery (building it) run in parallel.
Why it works: Design work doesn't fit neatly into a delivery sprint; running discovery one sprint ahead means validated designs are ready when the delivery sprint begins, instead of design forever catching up.
Key insights:
- The discovery track (research + design) feeds the delivery track (engineering + QA), staggered one sprint ahead
- User stories gain a hypothesis and success metric alongside acceptance criteria
- "Definition of Done" for UX includes validated learning, not just shipped pixels
- Backlog items from invalidated hypotheses are removed, not deferred
Product applications:
| Context | Application | Example |
|---|---|---|
| Sprint planning | Include hypothesis validation in sprint goals | "Sprint goal: validate that inline editing cuts task time 20%" |
| Backlog refinement | Attach experiment results to stories | Story moves to delivery only after hypothesis is validated |
| Retrospectives | Review learning velocity alongside delivery velocity | "We validated 4 hypotheses and invalidated 2 this sprint" |
Ethical boundary: Never use Lean UX as an excuse to skip accessibility, security, or compliance—these are non-negotiable quality standards, not assumptions to test.
See references/agile-integration.md when fitting discovery into a delivery cadence — the staggered dual-track sprint mechanics, how stories carry a hypothesis, and a UX Definition of Done.
See references/case-studies.md when you want a worked end-to-end example to model an engagement on — four composite scenarios (enterprise, startup, agency, internal tools) showing assumptions, experiments, and before/after outcome metrics.
Common Mistakes
| Mistake | Why It Fails | Fix |
|---|---|---|
| Treating MVPs as launches | Over-building by conflating MVP with first release | Reframe: MVP = learning tool, not product launch |
| Skipping assumption declaration | Hidden assumptions become expensive surprises | Run a 30-minute assumption mapping session at kick-off |
| Hypothesis without success criteria | Can't tell if the experiment passed | Pre-commit to metric, threshold, and sample size |
| Designer-only design | Handoff waste, misalignment, slow iteration | Run Design Studio sessions with the full team |
| Research as a phase | Feedback arrives too late to matter | Embed lightweight research in every sprint |
| Ignoring invalidated hypotheses | Building features that failed testing | Remove invalidated items from the backlog; pivot or drop |
| Documenting instead of collaborating | 40-page specs nobody reads | Replace specs with shared understanding from co-design |
| Measuring outputs not outcomes | Shipping features that don't change behavior | Define success as behavior change, not delivery |
Quick Diagnostic
Audit any UX process or design plan:
| Question | If No | Action |
|---|---|---|
| Are assumptions explicitly declared? | Hidden assumptions drive decisions | Run an assumption mapping workshop |
| Is there a testable hypothesis? | Building on opinion | Write hypothesis in standard format before designing |
| Is the experiment the lowest fidelity that answers the question? | Over-investing before learning | Downgrade to paper prototype or smoke test |
| Does the whole team participate in design? | Handoff waste and misalignment | Schedule a Design Studio session |
| Is research happening every sprint? | Feedback loop too slow | Establish a weekly testing cadence |
| Are you tracking outcomes, not just outputs? | Shipping without learning | Define behavior-change metrics per feature |
| Does UX work feed into Agile smoothly? | Design bottleneck or sprint-zero trap | Implement dual-track agile with staggered sprints |
| Can you point to a recently invalidated hypothesis? | Not learning; confirmation bias | Review the experiment log and celebrate a pivot |
Further Reading
For the complete methodology, research, and case studies:
- "Lean UX: Designing Great Products with Agile Teams" by Jeff Gothelf & Josh Seiden
- "Sense and Respond" by Jeff Gothelf & Josh Seiden (scaling outcome-focused thinking across organizations)
About the Authors
Jeff Gothelf is an organizational designer, coach, and author who spent over 15 years leading UX teams at companies including TheLadders and Neo Innovation; watching teams waste months on unvalidated deliverables led him to create Lean UX. Josh Seiden is a designer and product strategist with 25+ years of experience who co-founded the interaction design practice at Cooper and was Managing Director at Neo Innovation. Together they co-authored Lean UX and Sense and Respond.
Files (skills)
-
references
-
agile-integration.md 13 KB
# Agile Integration for Lean UX Lean UX was designed to work inside Agile development teams, not alongside them. The central mechanism is dual-track agile, where a discovery track (learning what to build) runs in parallel with a delivery track (building it). This reference covers the practical mechanics of making UX and Agile work together. ## Dual-Track Agile Dual-track agile separates the work of figuring out what to build from the work of building it. Both tracks run continuously and feed each other. ### Track Definitions | Track | Purpose | Activities | Output | |-------|---------|-----------|--------| | **Discovery** | Learn what to build | Hypothesis writing, experiments, user research, collaborative design, prototype testing | Validated hypotheses, tested prototypes, experiment results | | **Delivery** | Build what has been validated | Sprint planning, development, QA, deployment, monitoring | Shippable software, production features | ### How the Tracks Connect ``` DISCOVERY TRACK (Sprint N) DELIVERY TRACK (Sprint N+1) Hypothesis → Experiment Validated design → Development → Prototype → User test → Code → QA → Deploy → Validated design ─────────────────→ (enters delivery backlog) → Invalidated design ──→ (discard or re-hypothesize) ``` **Key rule:** Only validated designs enter the delivery backlog. Invalidated designs are discarded, pivoted, or re-tested. The delivery team never builds something the discovery team has not tested. ### Discovery Track Activities | Week | Activity | Participants | Output | |------|----------|-------------|--------| | Monday | Review last experiment results; write new hypotheses | PM, Designer, Tech Lead | Updated hypothesis log | | Tuesday | Collaborative design session (Design Studio or sketching) | Full team | Sketches, direction to prototype | | Wednesday | Build prototype (paper or clickable) | Designer + Developer pair | Testable artifact | | Thursday | Run user tests (3-5 sessions) | Designer (facilitator), Team (observers) | Raw findings | | Friday | Synthesize results; update backlog; plan next experiment | PM, Designer, Tech Lead | Validated/invalidated hypotheses | ### Delivery Track Activities The delivery track follows standard Agile/Scrum ceremonies but with Lean UX modifications: | Ceremony | Lean UX Modification | |----------|---------------------| | **Sprint Planning** | Each story includes its hypothesis and success metric. Team reviews experiment evidence before committing. | | **Daily Standup** | Include discovery track updates alongside delivery updates. "Yesterday I tested the prototype with 3 users; today I'm synthesizing results." | | **Sprint Review/Demo** | Demo includes experiment results and learnings, not just shipped features. "We shipped feature X AND we learned Y from our experiment." | | **Retrospective** | Review learning velocity alongside delivery velocity. "How many hypotheses did we validate/invalidate? Are we learning fast enough?" | ## Fitting UX Work into Sprints The most common complaint about UX in Agile is that design work does not fit into a sprint. Lean UX solves this with three techniques: staggered sprints, T-shaped participation, and hypothesis-driven stories. ### Staggered Sprints Discovery runs one sprint ahead of delivery. While the delivery team builds features validated in Sprint N, the discovery team validates designs for Sprint N+1. ``` Sprint 1 Discovery: Validate design for Feature A Sprint 1 Delivery: (Build Feature Z from previous discovery) Sprint 2 Discovery: Validate design for Feature B Sprint 2 Delivery: Build Feature A (validated in Sprint 1 Discovery) Sprint 3 Discovery: Validate design for Feature C Sprint 3 Delivery: Build Feature B (validated in Sprint 2 Discovery) ``` **Benefits:** - Design is never a bottleneck; validated designs are always ready when the sprint starts - Discovery and delivery happen simultaneously, not sequentially - If a hypothesis is invalidated, it does not derail the current delivery sprint **Common pitfall:** If discovery gets more than one sprint ahead, validated designs become stale by the time they reach delivery. Keep the gap to exactly one sprint. ### T-Shaped Participation In a Lean UX team, every member has a primary skill (the vertical bar of the T) and broad participation in adjacent activities (the horizontal bar). | Role | Primary Skill | Lean UX Participation | |------|-------------|----------------------| | **Designer** | Visual/interaction design | Facilitates experiments, writes hypotheses, codes simple prototypes | | **Developer** | Code, architecture | Participates in Design Studio, pair-designs with designer, builds experiment infrastructure | | **Product Manager** | Strategy, prioritization | Writes hypotheses, facilitates assumption workshops, defines success metrics | | **QA** | Testing, edge cases | Participates in design critique, identifies testable scenarios, reviews experiment results | ### Hypothesis-Driven User Stories Traditional user stories focus on what to build. Lean UX stories add why (the hypothesis) and how we will know (the metric). **Traditional format:** ``` As a [user], I want [feature], so that [benefit]. Acceptance criteria: [technical specification] ``` **Lean UX format:** ``` As a [user], I want [feature], so that [benefit]. HYPOTHESIS: We believe [outcome] will happen if [persona] achieves [action] with [feature]. SUCCESS METRIC: [metric] will [change] by [amount] within [timeframe]. EXPERIMENT: [How this was validated in discovery] Acceptance criteria: [technical specification] ``` **Example:** ``` As a project manager, I want to filter tasks by due date, so that I can focus on urgent work. HYPOTHESIS: We believe task completion rate will increase by 15% if project managers use a due-date filter to surface overdue tasks. SUCCESS METRIC: Task completion rate measured over 2 weeks post-launch. EXPERIMENT: Validated with clickable prototype (5 users, 4/5 completed filter task in under 10 seconds, all reported it would change their daily workflow). Acceptance criteria: - Filter dropdown with options: Today, This Week, Overdue, Custom Range - Persists across sessions - Works on mobile ``` ## Backlog Management ### Lean UX Backlog Columns | Column | Definition | Entry Criteria | Exit Criteria | |--------|-----------|---------------|---------------| | **Assumptions** | Unvalidated ideas and assumptions | Any team member can add | Prioritized in assumption matrix | | **Hypotheses** | Assumptions converted to testable predictions | Written in standard hypothesis format | Experiment designed and scheduled | | **Testing** | Hypotheses currently being tested | Experiment is actively running | Experiment complete, results analyzed | | **Validated** | Hypotheses confirmed by experiment | Data meets pre-set success threshold | Ready for delivery sprint planning | | **Invalidated** | Hypotheses disproven by experiment | Data falls below pre-set threshold | Archived with learnings; team decides pivot or drop | | **Delivery** | Validated designs in development | Enters delivery sprint | Shipped to production | ### Backlog Grooming for Lean UX During backlog grooming: 1. **Review invalidated hypotheses.** Decide: pivot (new hypothesis for same problem) or drop (problem is not worth solving). 2. **Re-prioritize assumptions.** New information from experiments may change which assumptions are highest risk. 3. **Size experiments, not features.** In discovery, estimate the effort to run an experiment, not the effort to build the final feature. 4. **Remove zombie items.** If a backlog item has not been tested in 3 sprints, it is either not important enough to test or the team lacks conviction. Remove or re-prioritize. ## Working with Engineering Teams ### Building Trust The biggest barrier to Lean UX adoption is often trust between design and engineering. Engineers may resist if they feel: - Design decisions are arbitrary ("the designer just likes it this way") - Requirements change constantly ("they keep changing their mind") - Their input is not valued ("just build what the wireframe says") **Trust-building practices:** | Practice | How It Builds Trust | |----------|-------------------| | Invite engineers to Design Studio | Their ideas are valued; they see the reasoning behind design decisions | | Share experiment results openly | Decisions are evidence-based, not opinion-based | | Pair design sessions | Designer and developer solve problems together; mutual respect grows | | Prototype together | Developer builds a quick prototype while designer directs; fast, collaborative | | Celebrate invalidated hypotheses | Shows the team that being wrong is expected and valuable | ### Handling Disagreements When designers and engineers disagree on a solution: 1. **Frame it as a hypothesis.** "We have two approaches. Let's write a hypothesis for each and test the riskier one." 2. **Use data, not authority.** "The experiment showed users preferred A. Let's go with the evidence." 3. **Time-box the debate.** "We have 10 minutes to decide. If we can't agree, we test both with 5 users and let them decide." ## Definition of Done for UX In traditional Agile, the Definition of Done focuses on engineering quality (code reviewed, tests passing, deployed). Lean UX expands the Definition of Done to include learning. ### Lean UX Definition of Done A feature is "done" when: | Criterion | Description | |-----------|-------------| | **Hypothesis validated** | The experiment met its pre-set success criteria | | **Design tested with users** | At least 5 users have tested the design (prototype or live) | | **Success metric defined** | The team knows exactly what metric to monitor post-launch | | **Instrumented** | Analytics events are in place to measure the success metric | | **Code complete and tested** | Standard engineering DoD (code review, unit tests, QA) | | **Deployed** | Feature is in production (behind a flag or fully rolled out) | | **Post-launch plan** | Team knows when and how they will review post-launch data | ### Post-Launch Learning Loop The Definition of Done extends beyond deployment: | Timeframe | Activity | Owner | |-----------|----------|-------| | **Day 1** | Verify instrumentation is working; check for errors | Engineer + Analyst | | **Week 1** | Review early metric data; compare to hypothesis target | PM + Designer | | **Week 2** | Run 3 follow-up interviews with users of the new feature | Designer | | **Sprint end** | Report results: validated, invalidated, or inconclusive | PM (in sprint review) | ## Sprint Zero Anti-Pattern **The problem:** Many teams use a "Sprint Zero" where designers work ahead for weeks before engineers start building. This creates a waterfall disguised as Agile. **Why it fails:** - Design decisions are made without engineering input - By the time engineers start, context is lost and designs need rework - The feedback loop between design and development is broken - Designers become a bottleneck **The Lean UX alternative:** - Start discovery and delivery simultaneously from day one - The first discovery sprint runs experiments with paper prototypes; there is no delay waiting for "finished designs" - Engineers participate in design sessions from the start - Use staggered sprints to maintain flow without a buffer sprint ## Scaling Lean UX ### Multiple Teams When multiple squads work on the same product: | Challenge | Solution | |-----------|----------| | Hypotheses overlap across teams | Shared hypothesis board visible to all squads | | Inconsistent experiment standards | Shared experiment template and success criteria norms | | Duplicate research | Shared research repository; weekly cross-team research sync | | Diverging design directions | Shared design system; cross-team Design Studio quarterly | ### Lean UX in SAFe / Large-Scale Agile | SAFe Concept | Lean UX Integration | |-------------|---------------------| | **Program Increment (PI) Planning** | Include discovery track objectives alongside delivery objectives | | **Architectural runway** | Discovery track identifies UX patterns needed for future features | | **Enabler stories** | Include experiment infrastructure (analytics, prototype tools, research ops) as enablers | | **Inspect and Adapt** | Review learning velocity and hypothesis validation rate across teams | ## Metrics for Lean UX in Agile Track these metrics to evaluate whether Lean UX is working within your Agile process: | Metric | What It Measures | Target | |--------|-----------------|--------| | **Hypotheses validated per sprint** | Learning velocity | 2-4 per sprint | | **Hypotheses invalidated per sprint** | Willingness to be wrong | At least 1 per sprint (0 means confirmation bias) | | **Time from hypothesis to experiment** | Discovery speed | Less than 1 sprint | | **Backlog items removed due to invalidation** | Waste prevention | At least 1 per quarter | | **Team members observing research** | Shared empathy | All team members observe at least 1 session per sprint | | **Post-launch metrics reviewed** | Closing the learning loop | 100% of shipped features reviewed within 2 weeks | -
case-studies.md 15.6 KB
# Case Studies: Lean UX in Practice These case studies illustrate how Lean UX principles apply across different organizational contexts. Each scenario is a realistic composite based on common patterns, not a specific company. The goal is to show how hypothesis-driven design, collaborative practices, and outcome-focused metrics work in the real world. ## Case Study 1: Enterprise Product Team ### Context A B2B SaaS company with 500 employees builds project management software for mid-market companies. The product team (12 people: 2 designers, 6 engineers, 2 PMs, 1 data analyst, 1 QA) has been shipping features from a roadmap driven by sales requests. Despite shipping 15 major features in the past year, net revenue retention is flat and NPS has declined from 42 to 35. ### The Problem The team is caught in the output trap. Sales submits feature requests, PMs write specs, designers create wireframes, engineers build, and the cycle repeats. No one checks whether shipped features actually improve user outcomes. The roadmap is a conveyor belt of outputs with no feedback loop. ### Lean UX Intervention **Week 1: Assumption Workshop** The PM organized a 60-minute assumption workshop with the full team. They identified 24 assumptions underlying the current roadmap. Prioritization revealed three high-risk, high-uncertainty assumptions: | Assumption | Risk | Uncertainty | |-----------|------|-------------| | "Users want Gantt chart view because sales says they do" | High (3 months of development planned) | High (no user research) | | "Users leave because we lack integrations" | High (churn is the main revenue threat) | High (based on exit survey with 12% response rate) | | "Our onboarding is fine because support tickets are low" | Medium (may be suppressing activation) | High (no onboarding funnel data) | **Week 2-3: Hypotheses and Experiments** The team wrote hypotheses and designed experiments for each assumption: **Hypothesis 1:** "We believe project managers will use Gantt charts at least 3x/week if they have a one-click timeline view in their project dashboard." - **Experiment:** Clickable Figma prototype tested with 8 existing users. - **Result:** 6 of 8 users could not complete the test task. 5 of 8 said they use spreadsheets for timeline views and would not switch. Hypothesis invalidated. - **Decision:** Removed Gantt chart from the roadmap, saving 3 months of development. Pivoted to a lightweight "milestones" view based on user feedback. **Hypothesis 2:** "We believe users churn because they cannot connect to their existing tools. Integration adoption will reduce 90-day churn by 20%." - **Experiment:** Concierge MVP. The team manually set up Zapier integrations for 15 at-risk accounts and measured engagement over 30 days. - **Result:** 9 of 15 accounts used the integration. Churn intent (measured by cancellation page visits) dropped by 35% in the test group. Hypothesis partially validated. - **Decision:** Built native integrations for the top 3 tools identified in the concierge test. **Hypothesis 3:** "We believe onboarding completion is low (baseline unknown) and that users who complete onboarding retain at 2x the rate of those who don't." - **Experiment:** Instrumented the existing onboarding flow to measure completion. - **Result:** Only 28% of new users completed onboarding. Users who completed onboarding had 2.4x higher 30-day retention. Hypothesis validated. - **Decision:** Redesigned onboarding using a Design Studio session. New onboarding tested with 5 users before development. ### Outcomes After 3 Months | Metric | Before Lean UX | After 3 Months | |--------|---------------|----------------| | Features shipped per quarter | 5 | 3 (but all validated) | | Hypotheses tested per quarter | 0 | 11 | | NPS | 35 | 41 | | Onboarding completion | 28% | 52% | | 90-day churn | 18% | 14% | | Roadmap items removed | 0 | 4 (saved ~6 months of development) | ### Key Lesson The team shipped fewer features but produced better outcomes. The Gantt chart cancellation alone saved 3 months of engineering time that was redirected to validated work. The cultural shift was the hardest part: the VP of Sales initially resisted removing Gantt charts from the roadmap but was convinced by the user test videos. ## Case Study 2: Startup ### Context A 6-person startup building a personal finance app for freelancers. The team consists of 1 founder/PM, 1 designer, 3 engineers, and 1 marketer. They have $400K in runway (8 months) and 200 beta users. The product has expense tracking and invoicing but retention is poor: only 15% of users are active after 30 days. ### The Problem The founder has a long list of features to build (tax estimation, bank sync, reporting, team billing) but limited runway. Building the wrong feature next could burn 2-3 months and bring the company closer to failure without improving retention. ### Lean UX Intervention **Day 1: Assumption Mapping** The team spent 90 minutes listing assumptions and plotting them on the prioritization matrix: Top 3 to test: 1. "Freelancers leave because they forget to log expenses" (retention assumption) 2. "Tax estimation is the feature that will make users pay" (monetization assumption) 3. "Users want to connect their bank account for auto-categorization" (value assumption) **Week 1: Rapid Experiments** **Experiment 1 (retention):** Interviewed 8 churned users over Zoom (30 minutes each). Found that 6 of 8 said they "forgot the app existed" after the first week. The problem was not missing features but missing triggers. - **Insight:** Users needed reminders, not more features. - **Quick test:** Sent manual weekly email summaries to 50 active users for 2 weeks. - **Result:** 30-day retention for the email group was 32% vs. 15% for the control group. Validated. **Experiment 2 (monetization):** Created a smoke test landing page for a "tax estimation" feature with a $9/month price tag and "Get Early Access" button. - **Result:** 12% of the 200 beta users clicked. 4% entered their email. Weak signal. Partially invalidated. - **Pivot:** Interviewed the 8 users who clicked. Found they wanted "peace of mind" about taxes, not a calculator. Pivoted hypothesis to a "tax set-aside" feature that automatically suggests how much to save per invoice. **Experiment 3 (bank sync):** Wizard of Oz test. Offered 10 users "automatic bank sync" and manually categorized their transactions for 1 week. - **Result:** 8 of 10 users logged in more frequently during the test week. Qualitative feedback was overwhelmingly positive. Validated. - **Decision:** Prioritized bank sync using a third-party API rather than building from scratch. **Week 2-8: Iterative Build and Test** The team implemented dual-track agile with 1-week sprints: - Discovery: Designer + founder tested 2 hypotheses per week - Delivery: Engineers built validated features ### Outcomes After 2 Months | Metric | Before | After 2 Months | |--------|--------|----------------| | 30-day retention | 15% | 34% | | Weekly active users | 30 | 68 | | Features built | 0 (all planned) | 3 (all validated) | | Runway consumed | N/A | 2 months | | Hypotheses tested | 0 | 14 | | Features removed from backlog | 0 | 5 | ### Key Lesson With limited runway, the startup could not afford to build the wrong feature. Lean UX helped them discover that the retention problem was not about features but about triggers. The weekly email summary (which took 2 hours to build) had more impact than the tax estimation feature (which would have taken 6 weeks). Speed of learning was the competitive advantage. ## Case Study 3: Agency ### Context A digital agency with 40 employees serves mid-market e-commerce clients. The design team (8 designers) typically delivers pixel-perfect mockups, comprehensive wireframe decks, and detailed style guides. Projects take 8-12 weeks from kick-off to handoff. Clients frequently request changes after handoff, causing scope creep and eroding margins. ### The Problem The agency's deliverable-heavy process creates two problems: (1) designers spend weeks on artifacts that change after client review, and (2) the final product often misses user needs because no real user testing happens until after launch. ### Lean UX Intervention **Pilot Project:** A new e-commerce client wanted a redesign of their checkout flow to reduce abandonment (currently 72%). **Week 1: Collaborative Kick-Off** Instead of a traditional creative brief, the agency ran a Lean UX kick-off: 1. **Assumption workshop (2 hours)** with client stakeholders, agency designers, and the client's development team. Identified 16 assumptions about why users abandon checkout. 2. **Hypothesis prioritization:** Top 3 assumptions to test: - "Users abandon because the form is too long (18 fields)" - "Users abandon because shipping costs are hidden until step 4" - "Users abandon because guest checkout is hard to find" 3. **Design Studio (90 minutes):** Agency designers, client's developer, and client's marketing lead sketched solutions together. **Week 2: Prototype and Test** - Built clickable prototype of the top-voted checkout concept in 2 days - Recruited 6 of the client's actual customers - Ran 30-minute moderated usability sessions - Findings: Form length was not the issue (users did not mind the fields). Hidden shipping costs were the primary abandonment trigger (5 of 6 users mentioned it). Guest checkout was fine. **Week 3: Iterate and Retest** - Revised prototype to show shipping cost estimate on the cart page (before checkout) - Retested with 5 new users - All 5 completed checkout. 4 of 5 specifically mentioned that seeing shipping costs early made them more confident. **Week 4: Handoff** - Delivered a tested, validated prototype (not a 60-page wireframe deck) - Client's development team had attended the Design Studio and usability sessions, so they already understood the design - Handoff meeting took 30 minutes instead of the usual 3 hours ### Outcomes | Metric | Before Lean UX | After Lean UX | |--------|---------------|---------------| | Project duration | 10 weeks | 4 weeks | | Designer hours | 320 hours | 140 hours | | Client revision rounds | 4-5 | 1 | | User tests before launch | 0 | 2 rounds (11 users) | | Checkout abandonment | 72% | 58% (post-launch) | | Client satisfaction | "Met expectations" | "Exceeded expectations" | | Profit margin | 18% | 34% | ### Key Lesson The agency discovered that clients did not actually want 60-page wireframe decks; they wanted confidence that the design would work. Testing with real users provided that confidence faster and cheaper than polished deliverables. The agency now offers "Lean UX Sprints" as a service, charging the same fee but delivering in half the time with better results. ## Case Study 4: Internal Tools Team ### Context A 200-person logistics company has a 4-person internal tools team (1 PM, 1 designer, 2 engineers) that builds software for warehouse staff, dispatchers, and customer service agents. The team maintains 6 internal applications. Requests come from department heads, and the team has a 9-month backlog. No user research has ever been conducted because "we know our users; they sit down the hall." ### The Problem Despite a full backlog, the three most recent features shipped in the past 6 months have been underused. The warehouse barcode scanning feature (3 months to build) is used by only 2 of 15 warehouse staff. The dispatcher dashboard (2 months) was abandoned after 1 week because dispatchers reverted to their spreadsheet. ### Lean UX Intervention **Week 1: Assumption Audit** The team reviewed the 9-month backlog and categorized every item by assumption risk: | Backlog Category | Items | Validated | Assumed | Unknown | |-----------------|-------|-----------|---------|---------| | Warehouse tools | 8 | 0 | 5 | 3 | | Dispatcher tools | 6 | 0 | 4 | 2 | | Customer service tools | 5 | 0 | 3 | 2 | Every single backlog item was based on assumptions from department heads, with zero user validation. **Week 2: Go to the Gemba** The designer and PM spent 3 days observing actual users: - **Day 1:** Shadowed 3 warehouse staff for 4 hours each. Discovered that the barcode scanner was too slow (3 seconds per scan vs. the 0.5-second manual process). Staff needed speed, not technology. - **Day 2:** Sat with 2 dispatchers for a full shift. Discovered the dashboard was abandoned because it lacked a critical field (driver phone number) that the spreadsheet had. A 5-minute fix could resurrect the feature. - **Day 3:** Observed 3 customer service agents. Discovered they toggled between 4 different systems to resolve one ticket. The most impactful change would be a unified view, not a new tool. **Week 3-4: Rapid Hypothesis Testing** **Hypothesis 1:** "We believe dispatcher dashboard adoption will reach 80% if we add the driver phone number field and a one-click call button." - **Experiment:** Added the field (30-minute code change). Monitored adoption for 1 week. - **Result:** Adoption went from 0% to 73% in one week. Validated. **Hypothesis 2:** "We believe customer service resolution time will decrease by 25% if agents can see order status, shipping status, and customer history in a single pane." - **Experiment:** Paper prototype of unified view, tested with 4 agents using real (anonymized) ticket scenarios. - **Result:** All 4 agents completed tasks faster and expressed strong preference for the unified view. 3 of 4 identified additional data fields they needed. Validated (with refinements). **Hypothesis 3:** "We believe warehouse scan adoption will increase if scan time is under 1 second." - **Experiment:** Technical spike. Engineer determined that switching to a different barcode library could reduce scan time to 0.4 seconds. Built a prototype in 2 days. - **Result:** Tested with 5 warehouse staff. All 5 preferred the fast scanner. 4 of 5 said they would use it daily. Validated. ### Outcomes After 2 Months | Metric | Before | After | |--------|--------|-------| | Backlog items validated before building | 0% | 100% | | Features adopted by target users | 33% | 90% | | Dispatcher dashboard adoption | 0% | 73% | | Time spent observing users per month | 0 hours | 12 hours | | Backlog items removed (not worth building) | 0 | 7 | | Average time from request to validated solution | 3 months | 2 weeks | ### Key Lesson "We know our users" was the most dangerous assumption the team held. Sitting next to users for a single day revealed that the barcode scanner problem was about speed (not technology), the dashboard problem was about one missing field (not a redesign), and the service tool problem was about context switching (not a new tool). The cheapest Lean UX activity, direct observation, had the highest ROI. ## Cross-Cutting Themes Patterns that appear across all four case studies: | Theme | How It Appeared | Principle | |-------|----------------|-----------| | **Assumptions are invisible until surfaced** | Every team had critical assumptions they had never questioned | Start with an assumption workshop | | **Observation beats opinion** | Watching users revealed insights that surveys and stakeholders missed | Go to the gemba; watch real behavior | | **Small experiments prevent big waste** | A 30-minute code fix, a landing page, or a paper prototype saved months of misdirected effort | Choose the lowest-fidelity experiment that answers the question | | **Invalidation is valuable** | Removing features from the backlog was as impactful as building new ones | Celebrate invalidated hypotheses | | **Shared understanding beats documentation** | Teams that designed together and observed research together needed less handoff | Collaborative design and shared research observation | | **Outcomes reveal the truth** | Output metrics (features shipped) masked failure; outcome metrics (retention, adoption, task time) revealed reality | Measure behavior change, not feature delivery | -
collaborative-design.md 13 KB
# Collaborative Design in Lean UX Lean UX replaces the lone-designer model with cross-functional collaboration. Design is not a phase or a department; it is a team activity. The goal is shared understanding, not comprehensive documentation. When the whole team participates in design, handoff waste disappears and learning velocity increases. ## The Design Studio Method The Design Studio is the signature collaborative design technique in Lean UX. It is a structured, time-boxed workshop where the entire cross-functional team generates, critiques, and refines design solutions together. ### Design Studio Structure | Phase | Duration | Activity | Output | |-------|----------|----------|--------| | **Problem statement** | 5 min | Facilitator presents the hypothesis and constraints | Shared understanding of the problem | | **Individual sketching (diverge)** | 10 min | Each participant sketches 6-8 ideas on paper (6-up template) | Many diverse ideas from diverse perspectives | | **Present and critique** | 3 min per person | Each person presents sketches; team asks questions and gives feedback | Highlighted strong ideas, identified gaps | | **Pair sketching (converge)** | 10 min | Pairs combine the best ideas into refined concepts | Stronger, cross-pollinated concepts | | **Present and critique round 2** | 3 min per pair | Pairs present refined concepts; team evaluates | Top 2-3 concepts to prototype | | **Team converge** | 10 min | Team selects elements from the best concepts for a unified direction | One direction to prototype and test | **Total time:** 60-90 minutes. ### Facilitation Tips - **Enforce time boxes strictly.** Divergent thinking needs pressure. If you give people 30 minutes to sketch, they will agonize over one idea. Give them 5 minutes and they produce six ideas because they cannot overthink. - **No laptops during sketching.** Paper and markers only. Digital tools slow divergent thinking and encourage premature refinement. - **Everyone sketches.** Engineers, product managers, data analysts, stakeholders. Bad drawing is fine. The goal is ideas, not art. - **Critique the idea, not the person.** Frame feedback as "I like how this concept solves X" and "I wonder if Y would be a challenge" rather than "This is wrong." - **Dot voting to prioritize.** Give each participant 3 dots to place on their favorite ideas. This surfaces the team's collective intuition quickly. - **Capture decisions on camera.** Photograph the whiteboard or sketches. This is your documentation. No one needs to write a spec after a Design Studio. ### The 6-Up Template Each participant receives a sheet of paper divided into 6 panels. In 5 minutes, they sketch one idea per panel. This forces breadth over depth. ``` +----------+----------+----------+ | | | | | Idea 1 | Idea 2 | Idea 3 | | | | | +----------+----------+----------+ | | | | | Idea 4 | Idea 5 | Idea 6 | | | | | +----------+----------+----------+ ``` **Rules:** - One idea per box - No erasing; move to the next box - Labels and annotations are encouraged - Stick figures and boxes are valid UI sketches - Star your own favorite at the end ## Collaborative Sketching Sessions Beyond the formal Design Studio, Lean UX teams use shorter collaborative sketching sessions throughout the sprint. ### Quick Sketch Formats | Format | Duration | When to Use | |--------|----------|-------------| | **Crazy 8s** | 8 minutes (1 minute per sketch) | Rapid ideation for a specific screen or interaction | | **Solution sketch** | 15 minutes (one refined idea per person) | When the team has already narrowed the problem space | | **Storyboard** | 20 minutes (6-panel story per person) | When the hypothesis involves a multi-step user journey | | **How Might We brainstorm** | 15 minutes | Reframing the problem before sketching solutions | ### Integrating Sketching into Standups In a Lean UX team, the daily standup can include a 5-minute sketch round when the team encounters a design question. Instead of saying "let's schedule a meeting with the designer," anyone can grab a marker and say "here's what I'm thinking" on a whiteboard. **Key principle:** The cost of a bad sketch is near zero. The cost of a bad specification is a wasted sprint. ## Cross-Functional Design ### Who Participates and Why | Role | What They Bring | Why It Matters | |------|----------------|----------------| | **Designer** | Visual and interaction expertise, user empathy | Synthesizes input into coherent experiences | | **Developer** | Technical feasibility, implementation awareness | Prevents designing things that are impossible or expensive to build | | **Product Manager** | Business context, priorities, constraints | Ensures designs serve business goals and user needs | | **Data Analyst** | Usage patterns, metrics, quantitative evidence | Grounds design decisions in data, not assumptions | | **QA Engineer** | Edge cases, error states, system thinking | Catches problems before they become bugs | | **Stakeholder** | Domain expertise, organizational context | Reduces approval delays; builds buy-in through participation | ### Making Cross-Functional Design Work **Establish ground rules:** - No seniority in the sketching room. The intern's idea gets the same critique as the VP's. - "Yes, and" over "No, but." Build on ideas before criticizing them. - The designer is the synthesizer, not the dictator. They combine the best elements into a coherent design. - Decisions are based on the hypothesis, not personal preference. "Will this test our hypothesis?" is the only valid design criterion during a Design Studio. **Common resistance and how to address it:** | Resistance | Response | |-----------|----------| | "I can't draw." | "We're sketching ideas, not art. Boxes and arrows are perfect." | | "This isn't my job." | "The team that designs together builds faster and with fewer misunderstandings." | | "Just tell me what to build." | "We'll all be faster if we understand why we're building it. Your perspective catches things we miss." | | "We'll design by committee." | "The designer synthesizes. The team contributes perspectives, not pixels." | ## Reducing Waste in UX Deliverables Traditional UX produces documents that no one reads: 60-page wireframe decks, annotated mockups with 200 footnotes, user flow diagrams that are outdated by the time they are printed. Lean UX reduces deliverables to the minimum needed for shared understanding. ### The Deliverable Spectrum | Traditional Deliverable | Lean UX Replacement | Why It's Better | |------------------------|--------------------|-----------------| | 60-page wireframe deck | Whiteboard photo from Design Studio | Team was in the room; they remember the context | | Annotated mockup with spec notes | Figma prototype with developer in the room during design | Developer can ask questions in real time | | Persona document (20 pages) | Proto-persona on a single page | Updated weekly as team learns; not a shelf document | | User journey map (poster-sized) | Storyboard sketch from collaborative session | Created by the team, reflects current understanding | | Usability test report (30 pages) | 5-minute video highlight reel + 3 bullet findings | Team watches video together; shared empathy, not a report | ### The "Just Enough" Documentation Test Before creating a deliverable, ask: 1. **Who needs this information?** If the answer is "the team I sit with," a conversation replaces a document. 2. **Will this be read?** If the document will sit in a folder, replace it with a shared session. 3. **What is the minimum artifact that conveys the decision?** A photo of a whiteboard, a 3-bullet Slack message, or a Loom video often suffices. 4. **Is this for communication or for approval?** Approval artifacts may need more formality, but only as much as the approval process requires. ## Shared Understanding Over Documentation Shared understanding is the state where every team member has the same mental model of what is being built, why it is being built, and how success will be measured. It is the primary output of collaborative design. ### How to Build Shared Understanding | Technique | How It Works | |-----------|-------------| | **Co-located design sessions** | Team sketches together; understanding is built through the act of creating | | **Pair designing** | Designer and developer (or PM) work side by side on the same problem | | **Research observation** | The whole team watches at least 2 usability test sessions per sprint | | **Shared walls** | Physical or virtual walls displaying current hypotheses, experiment results, and design directions | | **Sprint demos with context** | Demo includes not just "what we built" but "what we learned and what we are testing next" | ### Measuring Shared Understanding A simple test: ask each team member independently to answer these three questions: 1. What are we building this sprint? 2. Why are we building it? (What hypothesis are we testing?) 3. How will we know if it worked? If answers diverge, shared understanding is low. Run a collaborative session to realign. ## Style Guides as Living Documents In Lean UX, style guides and design systems are not static reference PDFs. They are living, evolving artifacts maintained collaboratively by designers and developers. ### Living Style Guide Principles | Principle | What It Means | Anti-Pattern | |-----------|-------------|--------------| | **Code is the source of truth** | The style guide is a running code library, not a PDF | Designer updates Figma but not the code; they diverge | | **Joint ownership** | Designers and developers maintain the guide together | Only the design team updates the guide; engineers ignore it | | **Evolve with the product** | New patterns are added as they are built and validated | Style guide is a one-time project that becomes outdated | | **Low ceremony** | Adding a new component should take minutes, not days | New component requires a 3-meeting approval process | ### Style Guide Workflow 1. **Design exploration:** Designer sketches a new pattern during a Design Studio or individually. 2. **Collaborative refinement:** Designer and developer discuss feasibility and implementation. 3. **Build and document:** Developer implements the component; designer reviews. 4. **Add to the guide:** Component is added to the living style guide with usage guidelines. 5. **Iterate:** As the team learns from experiments, components evolve. ### Style Guide Content A minimal living style guide includes: | Section | Contents | |---------|----------| | **Colors** | Primary, secondary, semantic colors (success, error, warning) with hex and variable names | | **Typography** | Font families, sizes, weights, line heights for headings, body, captions | | **Spacing** | Base unit and scale (4px, 8px, 16px, 24px, 32px, 48px) | | **Components** | Buttons, inputs, cards, modals, navigation, tables with states and variants | | **Patterns** | Common layouts, form patterns, empty states, loading states, error states | | **Voice and tone** | Writing style, vocabulary, microcopy patterns | ## Remote Collaborative Design Distributed teams can run all collaborative design activities virtually with some adjustments. ### Virtual Design Studio Setup | Element | Tool Options | Tips | |---------|-------------|------| | **Sketching canvas** | FigJam, Miro, Mural | Pre-create sections for each participant | | **Video** | Zoom, Google Meet | Cameras on; body language matters during critique | | **Timer** | Built-in timer in FigJam/Miro | Visible to all participants; auto-alerts on time | | **Voting** | Dot voting in FigJam/Miro | Anonymous voting reduces bias | | **Capture** | Screenshot + paste into Confluence/Notion | Assign one person to capture decisions and photos | ### Remote Facilitation Adjustments - **Add 50% more time** for each phase compared to in-person (communication overhead) - **Use breakout rooms** for pair sketching instead of "turn to your neighbor" - **Explicit speaking order** during critique rounds to prevent talking over each other - **Async pre-work:** Share the hypothesis and context 24 hours before the session so participants arrive prepared - **Post-session summary:** Send a 3-bullet recap within 1 hour. In remote settings, the shared wall does not exist passively; you must actively push information. ## Anti-Patterns in Collaborative Design | Anti-Pattern | Symptom | Fix | |-------------|---------|-----| | **HiPPO dominance** | Highest-paid person's opinion wins every critique | Anonymous voting; data-driven critique ("does this test our hypothesis?") | | **Design by committee** | Every critique point becomes a mandatory change | Designer synthesizes; critique informs but does not dictate | | **Sketch theater** | People sketch to impress, not to explore | Enforce time pressure; praise quantity over quality | | **No follow-through** | Great ideas from the session are never prototyped | Assign action items at end of session; track in sprint backlog | | **Excluding developers** | Engineers see designs for the first time in sprint planning | Developers attend every Design Studio; pair design weekly | -
experiment-patterns.md 13.6 KB
# Experiment Patterns for Lean UX Experiments are the engine of learning in Lean UX. The right experiment answers the hypothesis with the least effort. Choosing the wrong experiment type wastes time, money, or both. This reference covers the full spectrum of UX experiments, from napkin sketches to coded A/B tests. ## Table of Contents 1. [Types of UX Experiments](#types-of-ux-experiments) 2. [Choosing the Right Experiment](#choosing-the-right-experiment) 3. [Experiment Design Template](#experiment-design-template) 4. [Minimum Viable Tests](#minimum-viable-tests) 5. [Running Experiments in Practice](#running-experiments-in-practice) 6. [Experiment Cheat Sheet](#experiment-cheat-sheet) --- ## Types of UX Experiments ### 1. Paper Prototypes **What it is:** Hand-drawn screens on paper or index cards. A facilitator plays "computer," swapping screens as the user taps or points. **Best for:** Early concept validation, flow testing, information architecture. **Effort:** Very low (30 minutes to create). **Fidelity:** Very low. **Confidence:** Low-medium. Validates flow and concept, not visual design or interaction details. **When to use:** - You have multiple competing concepts and need to narrow down - The hypothesis is about flow or content, not aesthetics - You need to test today, not next week **When NOT to use:** - The hypothesis depends on visual design, animation, or micro-interactions - Users need to interact with real data - Stakeholders will not trust low-fidelity evidence **How to run:** 1. Sketch each screen on a separate sheet or card 2. Write a realistic task scenario for the participant 3. Ask the participant to "tap" or point at what they would interact with 4. Swap screens manually based on their choices 5. Note where they hesitate, get confused, or go off-script ### 2. Clickable Prototypes **What it is:** Interactive mockups built in tools like Figma, Sketch, or InVision. Users click through a realistic-looking interface, but no backend logic exists. **Best for:** Usability testing, flow validation, stakeholder buy-in, developer communication. **Effort:** Medium (1-3 days). **Fidelity:** Medium-high. **Confidence:** Medium-high. Validates flow, layout, and basic usability. **When to use:** - The hypothesis involves user navigation or task completion - You need to test with users who expect a realistic experience - The prototype will also serve as a design reference for developers **When NOT to use:** - A paper prototype would suffice (over-investing) - The hypothesis is about performance, load times, or real data behavior - You need to test with hundreds of users (use coded experiments instead) **How to run:** 1. Build the key screens and link hotspots in your prototyping tool 2. Write 3-5 task scenarios 3. Recruit 5-8 participants matching your persona 4. Run moderated sessions (20-30 minutes each) 5. Track task completion, time on task, errors, and qualitative feedback ### 3. Concierge MVP **What it is:** Deliver the service or experience manually, person-to-person, without building any technology. The user receives the full value, but the backend is entirely human-powered. **Best for:** Validating that the solution genuinely solves the problem before investing in automation. **Effort:** Medium (ongoing manual work per user). **Fidelity:** High (the experience is real). **Confidence:** High. Real behavior with real value delivery. **When to use:** - You are unsure whether the solution concept works at all - The cost of building the automated version is high - You want to deeply understand the user's experience and edge cases **When NOT to use:** - The hypothesis is about scale or technology performance - You need to test with more than 10-20 users simultaneously - The value proposition depends on speed that only automation can provide **Example:** A meal-planning app manually emails personalized weekly meal plans and shopping lists to 10 users based on their dietary preferences, before building the algorithm. ### 4. Wizard of Oz **What it is:** The user interacts with what appears to be a functioning product, but a human behind the scenes is performing the work the technology would eventually do. **Best for:** Testing the user experience of an automated feature before building the automation. **Effort:** Medium (build the frontend; humans operate the backend). **Fidelity:** High from the user's perspective. **Confidence:** High. Users interact with what feels like a real product. **When to use:** - The hypothesis depends on the user experience of an AI, algorithm, or automation feature - Building the actual technology is expensive or risky - You want to learn what the "right" output looks like before training a model **When NOT to use:** - The hypothesis is about system performance or response time - Manual operation cannot replicate the technology's speed - Ethical issues arise from deception (always disclose if legally required) **Example:** A "smart" scheduling assistant that appears to use AI but is actually a team member reading requests and sending calendar invites manually. ### 5. Landing Page / Smoke Test **What it is:** A single web page describing a product or feature that does not yet exist, with a call to action (sign up, pre-order, request access). Measures demand by tracking how many people take the action. **Best for:** Demand validation before building anything. **Effort:** Low (half a day to create). **Fidelity:** Low (no product). **Confidence:** Medium. Measures stated intent, not actual usage. **When to use:** - You need to validate demand before committing development resources - The hypothesis is about whether people want this at all - You want to build an early-access list for future testing **When NOT to use:** - You already know there is demand and need to validate usability - The product concept is hard to explain without a demo - Your audience is internal (use interviews instead) **How to run:** 1. Create a landing page with a clear value proposition, 2-3 key benefits, and a CTA 2. Drive targeted traffic (ads, social media, email, communities) 3. Measure conversion rate (visitors to CTA clicks or sign-ups) 4. Set success threshold before launch (e.g., 5% sign-up rate from 500 visitors) 5. Follow up with sign-ups for qualitative interviews ### 6. A/B Test (Coded Experiment) **What it is:** Two or more versions of a live feature are shown to different user segments. Statistical analysis determines which version performs better on a target metric. **Best for:** Optimizing existing features, validating specific design changes with statistical rigor. **Effort:** High (requires code, traffic, and statistical analysis). **Fidelity:** Production-level. **Confidence:** Very high (if properly powered). **When to use:** - You have sufficient traffic to reach statistical significance - The hypothesis involves a measurable behavior change in an existing product - You need high-confidence evidence to justify a significant investment **When NOT to use:** - Traffic is too low for statistical significance (fewer than 1,000 users per variant) - The concept is entirely new (test with prototypes first) - The change is too small to produce a detectable effect ## Choosing the Right Experiment The decision depends on three factors: what you need to learn, how much confidence you need, and how much you can invest. ### Experiment Selection Matrix | Question to Answer | Best Experiment | Fidelity | Time | Confidence | |-------------------|-----------------|----------|------|------------| | "Does anyone want this?" | Landing page smoke test | Low | 1-2 days | Medium | | "Does the flow make sense?" | Paper prototype | Very low | 1 day | Low-Medium | | "Can users complete this task?" | Clickable prototype | Medium | 3-5 days | Medium-High | | "Does this solution actually work?" | Concierge MVP | High | 1-2 weeks | High | | "Will the automated version feel right?" | Wizard of Oz | High | 1-2 weeks | High | | "Which version performs better?" | A/B test | Production | 2-4 weeks | Very High | ### The Fidelity Ladder Start at the lowest rung that can answer your question. Only climb higher when lower fidelity cannot provide the needed confidence. ``` Level 1: Paper prototype / Sketches ↓ (if concept validated, test usability) Level 2: Clickable prototype (Figma, Sketch) ↓ (if usability validated, test real value) Level 3: Concierge MVP / Wizard of Oz ↓ (if value validated, test at scale) Level 4: Coded experiment / A/B test ↓ (if optimized, ship) Level 5: Production release ``` ## Experiment Design Template Use this template for every experiment, regardless of type: ``` EXPERIMENT DESIGN ================= Date: _______________ Hypothesis ID: _______________ Experimenter: _______________ HYPOTHESIS We believe _______________ will happen if _______________ achieves _______________ with _______________. EXPERIMENT TYPE [ ] Paper prototype [ ] Clickable prototype [ ] Concierge MVP [ ] Wizard of Oz [ ] Landing page test [ ] A/B test [ ] Other: _______________ AUDIENCE Target persona: _______________ Sample size: _______________ Recruitment method: _______________ DESIGN What we will build/prepare: _______________ What the participant will do: _______________ What we will observe/measure: _______________ SUCCESS CRITERIA Primary metric: _______________ Success threshold: _______________ Failure threshold: _______________ TIME BOX Build time: _______________ Run time: _______________ Analysis time: _______________ Total: _______________ RESULTS (fill after experiment) Primary metric result: _______________ Qualitative observations: _______________ Surprises: _______________ Decision: [ ] Validate [ ] Iterate [ ] Pivot [ ] Kill Next step: _______________ ``` ## Minimum Viable Tests A minimum viable test (MVT) is the simplest possible experiment that can answer a specific question. The goal is to learn before you build, not to test what you have already built. ### MVT Examples by Question | Question | Minimum Viable Test | Time | Cost | |----------|-------------------|------|------| | "Do people understand our value proposition?" | Show landing page to 5 people, ask them to explain it back | 2 hours | Free | | "Will users find this navigation intuitive?" | Card sort with 10 users using index cards | 3 hours | Free | | "Is this onboarding flow clear?" | Clickable prototype with 5 users | 2 days | Free | | "Do users prefer layout A or B?" | First-click test on UsabilityHub | 4 hours | $50-100 | | "Will users pay for this feature?" | Add pricing page with "buy" button that leads to waitlist | 1 day | $50 for ads | | "Is this workflow faster than the current one?" | Time-on-task comparison: 5 users on old flow, 5 on prototype | 1 day | Free | ### The 5-User Rule Jakob Nielsen's research shows that 5 users uncover approximately 85% of usability problems. For Lean UX experiments: - **5 users** for qualitative usability tests (prototype tests, task analysis) - **20+ users** for quantitative surveys or preference tests - **1,000+ users per variant** for statistically significant A/B tests Do not over-recruit for qualitative tests. Five users, tested quickly, are better than 50 users tested slowly. ## Running Experiments in Practice ### Weekly Experiment Cadence A mature Lean UX team runs experiments every week. Here is a sample cadence: | Day | Activity | |-----|----------| | **Monday** | Review last week's results. Write new hypotheses. Design this week's experiment. | | **Tuesday** | Build experiment artifact (prototype, landing page, test script). | | **Wednesday** | Recruit participants (or launch ad traffic for smoke tests). | | **Thursday** | Run experiment sessions (usability tests, interviews). | | **Friday** | Synthesize results. Update hypothesis log. Plan next experiment. | ### Remote Experiment Tips - Use screen-sharing tools (Zoom, Lookback) for moderated prototype tests - Unmoderated tools (Maze, UserTesting) scale to more participants but lose qualitative depth - Record sessions (with consent) so the full team can watch asynchronously - Use virtual whiteboards (FigJam, Miro) for collaborative synthesis ### Common Experiment Failures | Failure | Cause | Prevention | |---------|-------|------------| | Leading questions | Facilitator hints at the "right" answer | Use neutral prompts: "What would you do next?" not "Would you click here?" | | Confirmation bias | Team sees only evidence that supports their idea | Assign a devil's advocate; review raw data before discussing | | Too few participants | Results are unreliable | Minimum 5 for qualitative, 1,000+ per variant for A/B | | No success criteria | Any result is interpreted as success | Define thresholds before running the experiment | | Testing too late | Feature is already built; team is reluctant to change | Test early with low-fidelity artifacts; never skip the prototype stage | | Wrong audience | Testing with colleagues instead of real users | Recruit external participants matching the target persona | ## Experiment Cheat Sheet Quick reference for choosing and running experiments: | If you need to learn... | Use this experiment | Minimum time | Participants | |------------------------|-------------------|-------------|-------------| | Does the concept resonate? | Landing page smoke test | 2 days | 200+ visitors | | Does the flow work? | Paper or clickable prototype | 1-2 days | 5 users | | Is the solution valuable? | Concierge MVP | 1-2 weeks | 5-10 users | | Does the "smart" feature feel right? | Wizard of Oz | 1-2 weeks | 5-10 users | | Which design wins? | A/B test | 2-4 weeks | 1,000+ per variant | | What do users really need? | Customer interview | 1 day | 5-8 users | | How do users organize information? | Card sort | 3 hours | 10-15 users | | What do users notice first? | First-click or 5-second test | 4 hours | 20+ users | -
hypothesis-canvas.md 11 KB
# Hypothesis Canvas The hypothesis canvas is the central planning artifact in Lean UX. It transforms vague product ideas into structured, testable predictions. Every design initiative should start here, not in a wireframing tool. ## The Lean UX Hypothesis Format The standard hypothesis statement links four elements into a single testable prediction: ``` We believe [outcome] will happen if [persona] achieves [action] with [feature]. ``` ### Breaking Down the Components | Component | Definition | Example | |-----------|-----------|---------| | **Outcome** | The measurable business or user result you expect | "A 15% increase in trial-to-paid conversion" | | **Persona** | The specific user segment the hypothesis targets | "First-time project managers using our free tier" | | **Action** | The behavior you expect users to perform | "Complete the guided project setup wizard within the first session" | | **Feature** | The design change or product element that enables the action | "A 3-step interactive setup wizard on the dashboard" | ### Complete Hypothesis Examples **E-commerce checkout redesign:** "We believe cart abandonment will drop by 20% if returning shoppers can complete checkout in under 60 seconds with a one-click reorder button on the product page." **SaaS onboarding:** "We believe 7-day retention will increase by 25% if new users who sign up via the marketing site achieve their first successful data import within 10 minutes with an auto-mapping CSV import tool." **Internal tools:** "We believe support ticket resolution time will decrease by 30% if support agents can view the full customer timeline without switching tabs with a unified customer context panel in the ticket view." **Mobile app engagement:** "We believe weekly active usage will increase by 40% if casual fitness users log at least 3 workouts in their first week with a simplified one-tap workout logging feature." ## Assumption Prioritization Matrix Before writing hypotheses, the team must surface and prioritize assumptions. The prioritization matrix plots assumptions on two axes. ### The Two Axes | Axis | Definition | Scale | |------|-----------|-------| | **Risk** | How damaging is it if this assumption is wrong? | Low risk (minor inconvenience) to High risk (project failure) | | **Uncertainty** | How confident are we that this assumption is true? | Low uncertainty (strong evidence) to High uncertainty (pure guess) | ### The Four Quadrants ``` HIGH UNCERTAINTY | +-----------+---------+---------+-----------+ | | | | | Monitor | TEST FIRST | | | (Low R, | (High R, | | | High U) | High U) | | | | | | LOW RISK -------+---------+---------+------- HIGH RISK | | | | | Ignore | Mitigate | | | (Low R, | (High R, | | | Low U) | Low U) | | | | | | +-----------+---------+---------+-----------+ | LOW UNCERTAINTY ``` ### Quadrant Actions | Quadrant | Risk | Uncertainty | Action | |----------|------|-------------|--------| | **Test First** | High | High | Write hypothesis, design experiment, test immediately | | **Monitor** | Low | High | Gather data passively; test if uncertainty persists | | **Mitigate** | High | Low | Build safeguards; standard engineering/design best practices | | **Ignore** | Low | Low | Do not spend time here; proceed with confidence | ### Running the Prioritization Workshop **Time:** 45-60 minutes **Participants:** Product manager, designer, tech lead, and one stakeholder **Steps:** 1. **Generate assumptions (15 min).** Each participant writes assumptions on sticky notes (physical or virtual). One assumption per note. Use prompts: - "Our users are..." - "Users will use this feature because..." - "This will work because..." - "We will make money by..." - "The biggest risk is..." 2. **De-duplicate and cluster (10 min).** Group similar assumptions. Combine duplicates. Give each cluster a label. 3. **Plot on matrix (15 min).** The facilitator reads each assumption aloud. The team discusses and places it on the 2x2 matrix. Disagreement is valuable; it reveals hidden uncertainty. 4. **Select top 3 for testing (10 min).** From the "Test First" quadrant, choose the three assumptions that, if wrong, would most damage the project. These become the first hypotheses. 5. **Write hypotheses (10 min).** Convert each selected assumption into a hypothesis using the standard format. Assign an owner and a target experiment date. ## Business Assumptions vs. User Assumptions Lean UX distinguishes two categories of assumptions. Both must be tested, but they require different experiments. ### Business Assumptions Business assumptions concern the viability and sustainability of the product from the organization's perspective. | Assumption Category | Example | Experiment Type | |---------------------|---------|-----------------| | **Revenue model** | "Users will pay $29/month for this feature" | Pricing page test, pre-order, willingness-to-pay survey | | **Market size** | "There are 50,000 potential customers in our ICP" | Market research, ad campaign response rates | | **Cost structure** | "We can deliver this for less than $5/user/month" | Concierge MVP cost tracking | | **Channel** | "Users will discover us through organic search" | SEO experiment, content test | | **Competitive advantage** | "Our solution is 3x faster than alternatives" | Comparative usability test | ### User Assumptions User assumptions concern the people who will use the product, their behaviors, needs, and context. | Assumption Category | Example | Experiment Type | |---------------------|---------|-----------------| | **Who they are** | "Our primary user is a mid-level marketing manager" | Customer interviews, analytics demographics | | **What they need** | "Users need to generate reports weekly" | Usage analytics, interview, diary study | | **Current behavior** | "Users currently use spreadsheets for this task" | Contextual inquiry, survey | | **Motivation** | "Users will switch because our tool saves 2 hours/week" | Time-on-task comparison, prototype test | | **Barriers** | "Users will not adopt if setup takes more than 10 minutes" | Onboarding funnel analysis, usability test | ### Connecting the Two A complete Lean UX canvas pairs business and user assumptions: ``` BUSINESS ASSUMPTION: Users will pay $29/month └─ USER ASSUMPTION: Users value the time saved enough to justify $29 └─ HYPOTHESIS: We believe paid conversion will reach 5% if marketing managers who complete 3 reports with our auto-generated template feature save at least 2 hours per week. ``` ## Sub-Hypotheses Large hypotheses often need decomposition. A sub-hypothesis isolates one variable from the parent hypothesis so it can be tested independently. ### When to Use Sub-Hypotheses - The parent hypothesis involves multiple unknowns - Testing the parent hypothesis requires building too much - The team disagrees on which component is the riskiest ### Decomposition Example **Parent hypothesis:** "We believe monthly active users will increase by 30% if new users complete a personalized onboarding flow with an AI-powered recommendation engine." **Sub-hypotheses:** | # | Sub-Hypothesis | Tests | |---|---------------|-------| | 1 | "We believe new users will engage with a personalized onboarding flow (measured by 70% completion rate)" | Clickable prototype test with 8 users | | 2 | "We believe AI recommendations during onboarding will feel relevant (measured by >4/5 relevance rating)" | Wizard of Oz test: manual recommendations presented as AI | | 3 | "We believe users who complete personalized onboarding will return within 7 days at 2x the rate of standard onboarding" | A/B test with coded prototype | ### Sub-Hypothesis Decision Tree After testing sub-hypotheses: - **All pass:** Proceed to build the parent feature. - **Some pass, some fail:** Redesign the failing component; retest. - **All fail:** Pivot. The parent hypothesis is likely wrong. ## Hypothesis Tracking Log Maintain a living document (spreadsheet or wiki) to track all hypotheses across sprints. | ID | Hypothesis | Status | Experiment | Metric | Target | Actual | Decision | |----|-----------|--------|------------|--------|--------|--------|----------| | H-001 | Trial-to-paid +10% with setup wizard | Testing | Prototype test | Completion rate | 70% | -- | -- | | H-002 | Cart abandonment -20% with one-click reorder | Validated | A/B test | Abandonment rate | 60% | 58% | Ship | | H-003 | Support time -30% with context panel | Invalidated | Usability test | Task time | 4 min | 6 min | Pivot | ### Review Cadence - **Weekly:** Update status of active experiments. - **Sprint boundary:** Review validated/invalidated count. Celebrate invalidations as learning. - **Quarterly:** Review patterns. Which assumption categories are most often wrong? Adjust future prioritization. ## Hypothesis Canvas Template Use this canvas at the start of every initiative: ``` LEAN UX HYPOTHESIS CANVAS ========================== Project: _______________ Date: _______________ Team: _______________ ASSUMPTIONS (top 3 from prioritization) 1. _______________ 2. _______________ 3. _______________ HYPOTHESIS #1 We believe _______________ will happen if _______________ achieves _______________ with _______________. Success metric: _______________ Target: _______________ Experiment type: _______________ Time box: _______________ HYPOTHESIS #2 We believe _______________ will happen if _______________ achieves _______________ with _______________. Success metric: _______________ Target: _______________ Experiment type: _______________ Time box: _______________ SUB-HYPOTHESES (if needed) 1a. _______________ 1b. _______________ OUTCOME (fill after experiment) Result: _______________ Learning: _______________ Decision: [ ] Validate & ship [ ] Iterate [ ] Pivot [ ] Kill Next hypothesis: _______________ ``` ## Anti-Patterns | Anti-Pattern | Why It Fails | Fix | |-------------|-------------|-----| | Writing hypotheses after building | Retroactive justification, not real testing | Hypotheses must exist before any design work | | Vague outcomes ("improve UX") | Cannot be measured or falsified | Use specific metrics with numeric targets | | Testing the safe assumption first | Wastes time; risky assumptions remain hidden | Use the prioritization matrix; test high-risk, high-uncertainty first | | One giant hypothesis per quarter | Too many variables; impossible to learn from failure | Decompose into sub-hypotheses testable in 1-2 weeks | | No pre-set success criteria | Team rationalizes any result as success | Define pass/fail thresholds before the experiment begins | | Hypothesis written by one person | Lacks diverse perspectives; blind spots persist | Run collaborative assumption workshop with full team | -
outcome-metrics.md 12.9 KB
# Outcome Metrics for Lean UX Lean UX shifts the definition of success from "did we ship it?" to "did it change behavior?" This reference covers how to choose, define, and track metrics that measure real outcomes rather than outputs. ## Outcomes vs. Outputs The distinction between outcomes and outputs is the philosophical core of Lean UX. ### Definitions | Term | Definition | Example | |------|-----------|---------| | **Output** | Something the team produces (a feature, a design, a release) | "We shipped the new onboarding flow" | | **Outcome** | A measurable change in user or business behavior resulting from the output | "New user 7-day retention increased from 25% to 38%" | ### Why the Distinction Matters Teams that measure outputs can ship features that fail silently. The feature is "done," the story is closed, and no one checks whether it actually helped. Teams that measure outcomes discover quickly when a shipped feature does not produce the expected behavior change, and they iterate or roll back before wasting more effort. ### The Output Trap The output trap occurs when teams optimize for velocity (stories per sprint, features per quarter) instead of impact. Symptoms include: - A long list of shipped features but flat or declining engagement metrics - Stakeholders asking "what did we ship?" instead of "what did we learn?" - Roadmaps measured by feature count, not behavior change - Retrospectives that review delivery speed but never experiment results ### Shifting the Conversation | Output-Focused Question | Outcome-Focused Question | |------------------------|-------------------------| | "When will this feature be done?" | "When will we know if this feature works?" | | "How many story points did we complete?" | "How many hypotheses did we validate?" | | "What's on the roadmap for Q3?" | "What user behavior are we trying to change in Q3?" | | "Can we ship by the end of the sprint?" | "Can we measure the result by the end of the sprint?" | ## Leading vs. Lagging Indicators Not all metrics are created equal. Leading indicators predict future results; lagging indicators confirm past results. Lean UX teams focus on leading indicators because they are actionable in the short term. ### Definitions | Type | Definition | Characteristics | Example | |------|-----------|----------------|---------| | **Leading** | Predicts future outcomes; changes quickly in response to actions | Actionable, fast feedback, sometimes noisy | "Onboarding completion rate" (predicts retention) | | **Lagging** | Confirms past outcomes; changes slowly | Reliable, slow feedback, hard to influence directly | "Monthly revenue" (confirms product-market fit) | ### Leading-Lagging Pairs Every important lagging indicator has one or more leading indicators that predict it. Lean UX teams identify these pairs and focus experiments on the leading indicators. | Lagging Indicator | Leading Indicator(s) | Why the Pair Works | |-------------------|---------------------|-------------------| | **Monthly revenue** | Trial-to-paid conversion rate, activation rate | Users who activate and convert drive future revenue | | **Annual churn rate** | Weekly engagement score, feature adoption rate | Users who stop engaging will eventually churn | | **NPS score** | Task completion rate, support ticket volume | Users who complete tasks without help are more likely to recommend | | **Customer lifetime value** | Feature breadth usage, upgrade path completion | Users who use more features stay longer and spend more | | **Market share** | Organic referral rate, branded search volume | Growing word-of-mouth predicts market share gains | ### Choosing Leading Indicators for Experiments For each hypothesis, identify the leading indicator that will show change first: | Hypothesis About | Lagging Indicator | Leading Indicator to Track | |-----------------|-------------------|---------------------------| | "New onboarding will improve retention" | 30-day retention | Onboarding completion rate (measurable in 1 day) | | "Simplified pricing will increase revenue" | Quarterly revenue | Pricing page click-through rate (measurable in 1 week) | | "In-app help will reduce support load" | Monthly support tickets | Help article engagement rate (measurable in 1 day) | | "Social features will increase engagement" | MAU | Invite sent rate, social action completion (measurable in 1 week) | ## Choosing Metrics That Matter ### The HEART Framework Google's HEART framework provides a structured way to choose metrics for UX outcomes: | Category | Definition | Example Metrics | |----------|-----------|-----------------| | **Happiness** | User satisfaction, perceived ease of use | NPS, CSAT, SUS score, satisfaction survey | | **Engagement** | Depth and frequency of interaction | DAU/MAU ratio, session length, actions per session | | **Adoption** | New users or new feature uptake | New accounts, feature activation rate, first-use completion | | **Retention** | Users who continue to use the product over time | Day-7 retention, Day-30 retention, churn rate | | **Task Success** | Ability to complete core tasks efficiently | Task completion rate, time on task, error rate | ### Applying HEART to a Hypothesis For each hypothesis, select 1-2 HEART categories most relevant to the expected outcome: | Hypothesis | Primary HEART Category | Metric | |-----------|----------------------|--------| | "Setup wizard improves first-day experience" | Adoption | Setup completion rate | | "Keyboard shortcuts speed up power users" | Task Success | Time on task for key workflows | | "Redesigned dashboard increases daily usage" | Engagement | DAU/MAU ratio | | "Simplified cancellation flow reduces complaints" | Happiness | CSAT score for cancellation flow | | "Weekly email digest brings users back" | Retention | Day-30 retention for email recipients vs. non-recipients | ### Metric Selection Checklist Before committing to a metric for a hypothesis, verify: - [ ] **Measurable:** Can we actually collect this data with our current instrumentation? - [ ] **Attributable:** Can we attribute changes to our experiment (not external factors)? - [ ] **Timely:** Will we see results within the experiment's time box? - [ ] **Actionable:** Will the result change what we do next? - [ ] **Leading:** Does this metric predict future outcomes, or only confirm past ones? ## OKRs for UX Teams OKRs (Objectives and Key Results) are a goal-setting framework that aligns well with Lean UX because Key Results are outcome-based, not output-based. ### Writing Outcome-Based OKRs | Component | Output-Focused (Bad) | Outcome-Focused (Good) | |-----------|---------------------|----------------------| | **Objective** | "Ship the new dashboard" | "Help users make faster data-driven decisions" | | **Key Result 1** | "Complete dashboard redesign by March 15" | "Dashboard users find their top metric in under 5 seconds (currently 18 sec)" | | **Key Result 2** | "Add 3 new chart types" | "Daily dashboard visits increase from 40% to 65% of active users" | | **Key Result 3** | "Write dashboard documentation" | "Support tickets about 'finding data' decrease by 50%" | ### OKR Templates for UX **Template 1: Feature-level OKR** ``` Objective: [User behavior change we want to see] KR1: [Leading metric] moves from [baseline] to [target] KR2: [Happiness or task success metric] reaches [threshold] KR3: [Adoption or engagement metric] reaches [threshold] ``` **Template 2: Team-level OKR (Learning velocity)** ``` Objective: Increase our team's learning velocity KR1: Validate or invalidate 8+ hypotheses per quarter KR2: Run user tests every sprint (0 sprints missed) KR3: 100% of shipped features have post-launch metrics reviewed within 2 weeks ``` **Template 3: Product-level OKR** ``` Objective: Improve new user activation KR1: Day-1 activation rate increases from 30% to 50% KR2: Time to first value decreases from 12 minutes to under 5 minutes KR3: Activation funnel drop-off at step 3 decreases by 40% ``` ## Measuring Behavior Change Lean UX ultimately measures whether a design changes behavior. Behavior change is the bridge between an output (what we shipped) and an outcome (what happened because we shipped it). ### Behavior Change Levels | Level | Definition | Measurement Method | Example | |-------|-----------|-------------------|---------| | **Awareness** | User knows the feature exists | Feature visibility rate, tooltip hover rate | "70% of users saw the new filter option" | | **Trial** | User tries the feature at least once | First-use rate, feature activation rate | "35% of users who saw the filter used it at least once" | | **Adoption** | User incorporates the feature into regular workflow | Weekly active use rate, feature frequency | "20% of users use the filter at least 3x per week" | | **Habit** | User uses the feature automatically without prompting | Unprompted use rate, time-to-first-use | "15% of users start their session by applying the filter" | ### Cohort Analysis for Behavior Change Track behavior change by cohort (users grouped by when they first encountered the change): | Cohort | Week 1 Trial | Week 2 Adoption | Week 4 Habit | Drop-off | |--------|-------------|-----------------|-------------|----------| | Jan 6-12 | 42% | 28% | 15% | 64% | | Jan 13-19 | 45% | 31% | 18% | 60% | | Jan 20-26 | 48% | 35% | 22% | 54% | If the trial-to-habit drop-off decreases over time, your iterations are working. ## Vanity Metrics to Avoid Vanity metrics make teams feel productive without revealing whether the product is improving. They are dangerous because they can mask failure. ### The Vanity Metric Test A metric is vanity if: 1. It only goes up (total signups, total page views) 2. It does not help you make a decision 3. It cannot be tied to a specific action or experiment 4. It looks impressive in a slide deck but does not change your behavior ### Common Vanity Metrics and Their Alternatives | Vanity Metric | Why It Misleads | Actionable Alternative | |---------------|----------------|----------------------| | **Total registered users** | Includes inactive, churned, and fake accounts | Monthly active users (MAU) | | **Page views** | Bots, accidental clicks, confusion loops inflate it | Engaged sessions (sessions with meaningful actions) | | **Time on site** | Could mean confusion, not engagement | Task completion rate + time on task for key flows | | **App downloads** | Download does not equal usage | Day-7 retention, activation rate | | **Social media followers** | Follower count does not equal engagement | Engagement rate (actions / impressions) | | **Feature count** | More features does not mean better product | Feature adoption rate (% of users using each feature) | | **Story points completed** | Measures delivery speed, not impact | Hypotheses validated per sprint | | **NPS alone** | Single number hides the distribution | NPS by cohort, segment, and feature + qualitative follow-up | ## Metric Dashboards for Lean UX ### Team Dashboard A Lean UX team dashboard should be visible to the entire team (physical wall or always-open screen) and updated at least weekly. **Essential dashboard elements:** | Section | Contents | Update Frequency | |---------|----------|-----------------| | **Current hypotheses** | Active experiments with status (testing, validated, invalidated) | Real-time | | **Key outcomes** | 3-5 outcome metrics with trend lines and targets | Weekly | | **Leading indicators** | 2-3 leading metrics that predict the key outcomes | Daily | | **Experiment log** | Last 5 experiments with results and decisions | Per experiment | | **Learning backlog** | Questions the team wants to answer next | Sprint boundary | ### Stakeholder Dashboard Stakeholders need a different view: less operational detail, more strategic outcome tracking. | Section | Contents | Purpose | |---------|----------|---------| | **OKR progress** | Key results with current vs. target | Shows strategic progress | | **Outcome trends** | 3-month trend lines for primary outcomes | Shows direction of change | | **Validated hypotheses** | Count + top learnings this quarter | Shows learning is happening | | **Pivots and kills** | Features removed or pivoted based on evidence | Shows discipline and waste prevention | ## Metric Anti-Patterns | Anti-Pattern | Why It Fails | Fix | |-------------|-------------|-----| | **Measuring too many things** | Team cannot focus; every metric is someone's priority | Choose 1 primary metric per hypothesis; max 3 OKR key results | | **No baseline** | Cannot tell if a metric improved or declined | Establish baseline before running any experiment | | **Post-hoc metric selection** | Cherry-picking the metric that showed improvement | Pre-commit to the metric in the hypothesis statement | | **Ignoring qualitative data** | Numbers say "what" but not "why" | Pair every quantitative metric with 3-5 user interviews | | **Dashboard but no action** | Data is collected but never reviewed or acted upon | Schedule weekly metric review; assign action items | | **Comparing averages** | Averages hide segment differences | Use cohort analysis and segment breakdowns |
-
-
SKILL.md 15.7 KB
--- name: lean-ux description: 'Apply lean thinking to UX: hypothesis-driven design, collaborative sketching, and rapid experiments instead of heavy deliverables. Use when the user mentions "Lean UX", "design hypothesis", "outcome over output", "design studio method", "assumption mapping", "lightweight research", "too much design documentation", or "get the team designing together". Also trigger when reducing design-documentation overhead, getting cross-functional teams to co-design, or running fast usability experiments. Covers hypothesis statements, MVPs for UX, and cross-functional collaboration. For Build-Measure-Learn, see lean-startup. For usability audits, see ux-heuristics.' license: MIT metadata: author: wondelai version: "1.4.0" --- # Lean UX Framework A practice-driven approach to UX that replaces heavy deliverables with rapid experimentation, cross-functional collaboration, and continuous learning. Lean UX shifts the question from "What should we design?" to "What do we need to learn?" ## Core Principle **Outcomes over outputs.** The value of a design is measured not by the fidelity of the deliverable but by the change in user behavior it produces. **The foundation:** Traditional UX waterfalls requirements into wireframes, mockups, specs, and code—losing context and hiding untested assumptions at every handoff. Lean UX compresses the distance between idea and evidence: declare assumptions, form hypotheses, run the smallest possible experiment, and let real user behavior settle the argument. Shared understanding replaces documentation; learning velocity replaces pixel perfection. ## Scoring **Goal: 10/10.** Score a UX process, design plan, or team workflow by the eight-row Quick Diagnostic below: award ~1.25 points per row answered "yes" (8 yeses = 10). Bands: - **9-10** — assumptions declared, hypotheses with pre-committed success criteria, lowest-fidelity experiments, whole-team design, weekly research, outcome (not output) metrics, dual-track agile, and a recently invalidated hypothesis on the books. - **5-6** — hypotheses exist but criteria are vague or fidelity is over-invested; design and research still partly siloed. - **<=3** — heavy deliverables, untested assumptions, output-counting, no experiment log. Always state the current score, the diagnostic rows that failed, and the specific fix for each. ## Framework ### 1. Declaring Assumptions **Core concept:** Every design starts with assumptions. Lean UX makes them explicit so they can be prioritized and tested, rather than baked invisibly into specifications. **Why it works:** Unspoken assumptions mean teams build on shaky ground and discover problems only after launch; surfacing them early focuses energy on the riskiest ones and reduces the cost of being wrong. **Key insights:** - Business assumptions define what must be true for the business (revenue model, market size, willingness to pay); user assumptions define who users are and how they behave - Prioritize on two axes: risk (how damaging if wrong) and uncertainty (how little we know) - Test high-risk, high-uncertainty assumptions first - Write assumptions collaboratively as a team, not in isolation **Product applications:** | Context | Application | Example | |---------|-------------|---------| | **New feature kick-off** | Assumption mapping workshop | "We assume users want to share reports with teammates" | | **Roadmap planning** | Rank features by assumption risk | Prioritize features whose success depends on untested beliefs | | **Stakeholder alignment** | Expose hidden assumptions across roles | PM assumes pricing works; engineer assumes scale; designer assumes flow | **Ethical boundary:** Assumptions must be honest assessments, not post-hoc justifications—if leadership has already committed to a direction, acknowledge the constraint rather than pretending it's open to falsification. See [references/hypothesis-canvas.md](references/hypothesis-canvas.md) when running an assumption workshop or writing a hypothesis — the risk/uncertainty prioritization matrix, business-vs-user assumption split, and fillable hypothesis and sub-hypothesis templates. ### 2. Hypothesis Statements **Core concept:** A hypothesis translates an assumption into a testable prediction, linking a proposed change to a measurable outcome for a specific user segment. **Why it works:** Hypotheses force precision—instead of "make onboarding better," the team commits to a prediction that can be proven or disproven, which prevents scope creep and makes the learn step unambiguous. **Key insights:** - Standard format: "We believe [outcome] will happen if [persona] achieves [action] with [feature]" - Every hypothesis specifies persona, action, outcome, and measurable signal - Sub-hypotheses break a large bet into independently testable parts - Agree on what "validated" and "invalidated" look like before running the experiment **Product applications:** | Context | Application | Example | |---------|-------------|---------| | **Feature design** | Write hypothesis before wireframing | "We believe trial-to-paid conversion will rise 10% if new users complete a guided setup wizard" | | **A/B tests** | Formalize test rationale | "We believe click-through will rise 15% if we move the CTA above the fold" | | **Sprint planning** | Attach hypothesis to each story | Story: "filter by date." Hypothesis: "task completion time drops 30%" | **Ethical boundary:** Never cherry-pick metrics after the fact to declare a hypothesis validated—pre-commit to success criteria. See [references/outcome-metrics.md](references/outcome-metrics.md) when picking the measurable signal for a hypothesis or defining team success — outcomes-vs-outputs, leading-vs-lagging indicator pairs, UX OKRs, and the vanity metrics to avoid. ### 3. MVPs and Experiments **Core concept:** An MVP in Lean UX is the smallest design artifact that can test a hypothesis with real users—a learning tool, not a product launch. **Why it works:** A paper prototype tested with five users in a hallway can invalidate a hypothesis that would otherwise consume a full engineering sprint; matching experiment fidelity to assumption risk maximizes learning per unit of effort. **Key insights:** - Experiments range from low fidelity (paper prototypes, concierge tests) to high fidelity (coded A/B tests, Wizard of Oz) - Choose the lowest-fidelity experiment that can answer the question - A good experiment has a clear hypothesis, defined audience, measurable signal, and time box - Proto-personas can stand in for full research when speed matters, but must be validated later **Product applications:** | Context | Application | Example | |---------|-------------|---------| | **Early concept validation** | Paper prototype or clickable mockup | Sketch 3 concepts, test with 5 users same day | | **Demand validation** | Landing page smoke test | "Sign up for early access" measures real interest | | **Usability validation** | Clickable prototype test | Figma prototype tested with 5-8 users | | **Pricing validation** | Painted door test | Show pricing page, measure click-through before building billing | **Ethical boundary:** Smoke tests and fake doors must not mislead users into believing a product exists—disclose test status and offer an opt-out. See [references/experiment-patterns.md](references/experiment-patterns.md) when choosing or designing an experiment — the full catalog of experiment types with when/when-NOT-to-run notes, the experiment selection matrix and fidelity ladder, and a design template. ### 4. Collaborative Design **Core concept:** Design is a team sport. Lean UX replaces the solitary designer-then-handoff model with cross-functional sessions where developers, PMs, and designers sketch solutions together. **Why it works:** Developers who helped sketch the solution don't need a 40-page spec to build it—shared understanding replaces documentation, diverse perspectives generate more creative solutions, and handoff waste drops dramatically. **Key insights:** - Design Studio method: diverge (individual sketching), present, critique, converge (refined sketch), iterate - The goal is informed commitment, not consensus: the team agrees on what to test, not what is "right" - Cross-functional means engineers, QA, data analysts, and stakeholders sketch too - Style guides and pattern libraries are living documents; reduce deliverables to the minimum needed for shared understanding (often a whiteboard photo) **Product applications:** | Context | Application | Example | |---------|-------------|---------| | **Sprint kick-off** | Design Studio session (90 minutes) | Whole team sketches solutions to the sprint's hypothesis | | **Feature exploration** | Collaborative sketching workshop | 6-up sketches: each person draws 6 ideas in 5 minutes | | **Remote teams** | Virtual whiteboard sessions | FigJam or Miro board with timed sketch rounds | **Ethical boundary:** Collaboration must not become design by committee—a designated designer synthesizes input; the team does not vote on pixels. See [references/collaborative-design.md](references/collaborative-design.md) when facilitating a Design Studio — the step-by-step workshop protocol (timings, materials, remote variants) and how to keep style guides as living documents. ### 5. Feedback and Research **Core concept:** Continuous, lightweight research replaces big-bang usability studies—small research activities embedded in every sprint instead of quarterly reports. **Why it works:** Findings only change a decision while it is still cheap to reverse, so research value decays with every sprint between learning and the decision it informs; small weekly studies keep that gap near zero, which a quarterly report never can. **Key insights:** - Research types: usability tests, customer interviews, A/B tests, analytics review, surveys, diary studies - Five users uncover approximately 85% of usability problems (Nielsen) - Continuous cadence: recruit weekly, test weekly, synthesize weekly - The whole team should observe at least some sessions to build empathy - Proto-personas are refined and eventually replaced by evidence-based personas **Product applications:** | Context | Application | Example | |---------|-------------|---------| | **Weekly usability testing** | Test prototype with 3-5 users every Thursday | "Testing Thursday" ritual with rotating facilitators | | **Post-launch learning** | Monitor analytics + 3 follow-up interviews | Find drop-off points, interview churned users | | **Persona validation** | Compare proto-persona assumptions to interview data | "We assumed power users are marketers; data shows ops managers" | **Ethical boundary:** Conduct research with informed consent—participants should understand how their data is used and be free to withdraw. ### 6. Integration with Agile **Core concept:** Lean UX works inside Agile via dual-track development: discovery (learning what to build) and delivery (building it) run in parallel. **Why it works:** Design work doesn't fit neatly into a delivery sprint; running discovery one sprint ahead means validated designs are ready when the delivery sprint begins, instead of design forever catching up. **Key insights:** - The discovery track (research + design) feeds the delivery track (engineering + QA), staggered one sprint ahead - User stories gain a hypothesis and success metric alongside acceptance criteria - "Definition of Done" for UX includes validated learning, not just shipped pixels - Backlog items from invalidated hypotheses are removed, not deferred **Product applications:** | Context | Application | Example | |---------|-------------|---------| | **Sprint planning** | Include hypothesis validation in sprint goals | "Sprint goal: validate that inline editing cuts task time 20%" | | **Backlog refinement** | Attach experiment results to stories | Story moves to delivery only after hypothesis is validated | | **Retrospectives** | Review learning velocity alongside delivery velocity | "We validated 4 hypotheses and invalidated 2 this sprint" | **Ethical boundary:** Never use Lean UX as an excuse to skip accessibility, security, or compliance—these are non-negotiable quality standards, not assumptions to test. See [references/agile-integration.md](references/agile-integration.md) when fitting discovery into a delivery cadence — the staggered dual-track sprint mechanics, how stories carry a hypothesis, and a UX Definition of Done. See [references/case-studies.md](references/case-studies.md) when you want a worked end-to-end example to model an engagement on — four composite scenarios (enterprise, startup, agency, internal tools) showing assumptions, experiments, and before/after outcome metrics. ## Common Mistakes | Mistake | Why It Fails | Fix | |---------|-------------|------| | **Treating MVPs as launches** | Over-building by conflating MVP with first release | Reframe: MVP = learning tool, not product launch | | **Skipping assumption declaration** | Hidden assumptions become expensive surprises | Run a 30-minute assumption mapping session at kick-off | | **Hypothesis without success criteria** | Can't tell if the experiment passed | Pre-commit to metric, threshold, and sample size | | **Designer-only design** | Handoff waste, misalignment, slow iteration | Run Design Studio sessions with the full team | | **Research as a phase** | Feedback arrives too late to matter | Embed lightweight research in every sprint | | **Ignoring invalidated hypotheses** | Building features that failed testing | Remove invalidated items from the backlog; pivot or drop | | **Documenting instead of collaborating** | 40-page specs nobody reads | Replace specs with shared understanding from co-design | | **Measuring outputs not outcomes** | Shipping features that don't change behavior | Define success as behavior change, not delivery | ## Quick Diagnostic Audit any UX process or design plan: | Question | If No | Action | |----------|-------|--------| | Are assumptions explicitly declared? | Hidden assumptions drive decisions | Run an assumption mapping workshop | | Is there a testable hypothesis? | Building on opinion | Write hypothesis in standard format before designing | | Is the experiment the lowest fidelity that answers the question? | Over-investing before learning | Downgrade to paper prototype or smoke test | | Does the whole team participate in design? | Handoff waste and misalignment | Schedule a Design Studio session | | Is research happening every sprint? | Feedback loop too slow | Establish a weekly testing cadence | | Are you tracking outcomes, not just outputs? | Shipping without learning | Define behavior-change metrics per feature | | Does UX work feed into Agile smoothly? | Design bottleneck or sprint-zero trap | Implement dual-track agile with staggered sprints | | Can you point to a recently invalidated hypothesis? | Not learning; confirmation bias | Review the experiment log and celebrate a pivot | ## Further Reading For the complete methodology, research, and case studies: - [*"Lean UX: Designing Great Products with Agile Teams"*](https://www.amazon.com/Lean-UX-Designing-Great-Products/dp/1098116305?tag=wondelai00-20) by Jeff Gothelf & Josh Seiden - [*"Sense and Respond"*](https://www.amazon.com/Sense-Respond-Successful-Organizations-Continuously/dp/1633691888?tag=wondelai00-20) by Jeff Gothelf & Josh Seiden (scaling outcome-focused thinking across organizations) ## About the Authors **Jeff Gothelf** is an organizational designer, coach, and author who spent over 15 years leading UX teams at companies including TheLadders and Neo Innovation; watching teams waste months on unvalidated deliverables led him to create Lean UX. **Josh Seiden** is a designer and product strategist with 25+ years of experience who co-founded the interaction design practice at Cooper and was Managing Director at Neo Innovation. Together they co-authored *Lean UX* and *Sense and Respond*.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.