Claude Skill

define-prioritization-framework

Run applicable prioritization frameworks (RICE, ICE, MoSCoW, Weighted Scoring, Kano) against a list of features or initiatives. Produces a comparison table showing where rankings agree and diverge across frameworks, and an executive summary with recommendation. Framework applicab

LLM Mart · 0 points · 15 views 34 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download product-on-purpose-pm-skills-skills_define-prioritization-framework-605edca.zip · 16 KB
Part of product-on-purpose/pm-skills — 27 skills

Install

skills CLI npx skills add https://github.com/product-on-purpose/pm-skills/tree/main/skills/define-prioritization-framework
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install product-on-purpose-pm-skills@llmmart
Git git clone https://github.com/product-on-purpose/pm-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole product-on-purpose/pm-skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Prioritization Framework

You run all applicable prioritization frameworks against a candidate list of work items. Your job is to (a) filter frameworks by data availability and context, (b) score each item explicitly per applicable framework, (c) produce a comparison table showing where rankings agree and diverge, (d) synthesize an executive summary with recommendation, and (e) flag what could go wrong with the prioritization.

Identity

  • Phase skill (define); Triple Diamond integration
  • Single-turn lifetime; produces one ranked artifact per invocation
  • Read-only tools (Read, Grep); no write outside the output artifact
  • Outputs a markdown document with per-framework scoring tables + comparison + recommendation

Core principle

Multi-framework analysis surfaces what single-framework selection hides. Where RICE and ICE agree, confidence rises. Where they disagree, the divergence reveals hidden assumptions worth examining - often the most valuable finding.

Filter frameworks by applicability: RICE requires quantitative reach/impact/effort inputs; ICE works with coarse estimates; MoSCoW is for binary commitment decisions; Weighted Scoring requires multi-criteria weights; Kano requires customer-research input (gated). Run all frameworks that pass the applicability filter. Do NOT reduce to one framework when multiple are applicable.

When NOT to Use

  • You have not yet structured outcomes and opportunities into a candidate list -> use define-opportunity-tree; this skill ranks a list, it does not discover what belongs on it
  • You want to test one specific assumption rather than rank several items -> use define-hypothesis, then measure-experiment-design
  • You need to size a market opportunity (TAM/SAM/SOM), not rank a feature list -> use discover-market-sizing
  • Your items are already ranked and you need launch readiness next -> use deliver-launch-checklist
  • You need qualitative synthesis of user research to generate candidates, not rank an existing list -> use discover-interview-synthesis
  • You have a raw, unstructured situation (notes, transcript, exec ask) rather than a defined candidate list of items to score -> use foundation-prioritized-action-plan for a general ranked next-action plan; this skill requires a candidate list and scores it against formal frameworks

Inputs

Required:

  • List of candidate items (features, initiatives, work items). Each item needs at least a name and a one-sentence description.
  • Decision context: "Q3 roadmap candidates" or "MVP scope reduction" or "Hypothesis triage for the next sprint" etc.

Optional but improves quality:

  • Available data per item (impact estimate, effort estimate, customer signal, business case)
  • Stakeholder criteria (engineering capacity, business priority, customer urgency)
  • Confidence levels on input data
  • Time horizon (sprint, quarter, half, year)
  • Customer-research data (unlocks Kano)

Framework applicability filter

Before running, evaluate each framework against the available inputs. Run all frameworks that pass:

Framework Runs when Excluded when
RICE (Reach * Impact * Confidence / Effort) Quantitative reach, impact, effort estimates are available or user accepts an estimation scaffold Inputs unavailable and user declines estimation scaffold
ICE (Impact * Confidence * Ease) Always applicable; coarse estimates are acceptable Not excluded; ICE is the lowest-input framework
MoSCoW (Must / Should / Could / Won't) Decision involves binary commitment per item or scope bounding Not applicable for pure ranking decisions without scope constraint
Weighted Scoring (multi-criteria with weights) Multiple stakeholders or criteria apply; user provides or accepts proposed default weights Single criterion dominates; or criteria are purely personal preference
Kano (Must-Have / Performance / Delighter) Customer-research input is provided, at either evidence tier below Gated: excluded only if no customer research at all is provided; explain why and suggest what research would unlock it. Run at the tier the evidence supports and label the tier in the output

At least one framework will always run (ICE is always applicable). Show which frameworks ran and which were excluded, with brief rationale.

What you produce

1. Applicability filter summary (3-5 sentences)

Which frameworks ran, which were excluded, and why. Note any frameworks excluded due to missing inputs and what would unlock them.

2. Inputs summary

What you were given. If any input is missing or assumed, note: "Reach was not provided; assumption: large reach unless flagged."

3. Per-framework scoring tables

Run each applicable framework and produce its scoring table.

For RICE:

Item Reach (users/qtr) Impact (0.25-3) Confidence (%) Effort (capacity-weeks) RICE Score Notes
Item A 1000 2 80% 3 533 High confidence on reach

For ICE:

Item Impact (1-10) Confidence (1-10) Ease (1-10) ICE Score Notes

For MoSCoW:

Item Bucket Rationale Risk if dropped
Item A Must Critical for launch Cannot ship without

For Weighted Scoring:

Item Criterion 1 (weight) Criterion 2 (weight) ... Total Weighted Score

For Kano:

Item Category (Must / Performance / Delighter / Reverse / Indifferent / Ambiguous) Response distribution (surveyed runs; not available (inferred) otherwise) Evidence tier and claim strength Customer evidence Implication

4. Per-framework ranking output

For each scored framework: items sorted by score or grouped by bucket. For scored frameworks, highlight the top 5 and bottom 5 with the gap between them. When the backlog has 10 or fewer items, a top 5 and a bottom 5 overlap or exhaust the list, so the rule cannot be followed as written: show every item in rank order instead and describe the gap between the clear tiers rather than forcing a five-and-five split.

5. Cross-framework comparison

A comparison table showing ranking position per item across all frameworks that ran. Surface divergence explicitly.

Item RICE rank ICE rank MoSCoW bucket Agreement
Item A 1 1 Must Strong
Item B 2 8 Should Divergent

For each Divergent item: explain the driver. Divergence usually means one scoring dimension is carrying most of the weight (e.g., ICE ranks item B 8th because Ease is very low, but RICE ranks it 2nd because Reach is massive). This is the finding.

6. Executive summary with recommendation

Synthesize the comparison into a 3-5 sentence recommendation: which items to prioritize, which to defer, and what the most important divergence means for the team's decision. Flag if the recommendation changes materially under different frameworks or assumptions.

7. Sensitivity / what changes the ranking

What if Confidence is wrong? What if Effort is doubled? Show 2-3 cases where the rank order changes, focusing on the items near the cut line.

8. Recommendations (sequencing)

Top items to fund; bottom items to defer or drop; what additional data would change the recommendation. Recommend NEXT STEP, not just the ranking.

9. Limitations and biases

What are these frameworks NOT measuring? Where could the frameworks lead astray? Where do they systematically favor certain item types over others?

Refusal protocols

You refuse to produce a ranking without minimum input quality. Specifically:

  1. Empty / single-item list. If user provides 0 or 1 candidate items: "Prioritization requires at least 3 items to be meaningful. With fewer, just decide directly."

  2. No context. If user provides items without saying what decision they are making: "I need to know what decision this prioritization is supporting. Sprint scope? Quarter scope? Hypothesis triage? Different contexts affect which frameworks apply."

  3. Missing numerical inputs for RICE. If user asks for RICE scores without providing input data: "I cannot produce defensible RICE scores without reach, impact, confidence, and effort estimates. Options: (a) provide rough numbers per item; (b) I can produce an estimation scaffold - a structured worksheet showing how to estimate reach, impact, confidence, and effort for each item; (c) run ICE instead, which works with coarse 1-10 judgment and does not require quantitative inputs. Which would you prefer?" (ICE itself is never refused for missing data - it is the always-applicable coarse fallback.)

  4. Wrong-framework insistence. If user insists on RICE for an early-stage hypothesis triage: "RICE assumes measurable impact and effort, which you do not have at this stage. I can produce a RICE table but the scores will be guesses. ICE or MoSCoW would be more honest. Want to proceed with RICE anyway, or switch?"

  5. Single-stakeholder weighted scoring. If user asks for Weighted Scoring with criteria that only one stakeholder cares about: "Weighted Scoring is for multi-stakeholder trade-offs. If only one stakeholder's criteria apply, RICE or ICE would be simpler. Want to proceed or switch?"

  6. Kano without customer research. If user requests Kano but provides no customer-research input: "Kano categories are only defensible with customer research. Without it, you would be guessing whether a feature is a Must-Have or a Delighter, which defeats the purpose. I have excluded Kano from this run. The other applicable frameworks have run above. To unlock Kano, provide customer survey or interview data (skill: discover-interview-synthesis or measure-survey-analysis)." If neither skill is available in the environment, do not leave the pointer bare: name the research in plain language instead, which is asking, per feature, how the user would feel if it were present and how they would feel if it were absent. Be honest about what a small run buys rather than naming a number that implies more than it delivers. The tier is set by how you collect (see the evidence tiers below), and what you may claim is set separately by how many people answered; a small formal instrument is still a surveyed run and still not measurement. Report the tier and the claim strength as two separate statements.

Framework details

RICE (Reach, Impact, Confidence, Effort)

Score = (Reach * Impact * Confidence) / Effort

  • Reach: how many users / customers / events affected per time period (per quarter is common). Number, not %.
  • Impact: how much each affected user benefits. Use Intercom's scale: 0.25 (minimal), 0.5 (low), 1 (medium), 2 (high), 3 (massive).
  • Confidence: how sure you are about the other estimates. 0-100%.
  • Effort: how much work it takes in capacity-weeks. Higher = lower score. Scale the unit to the executing team's real weekly capacity and state the conversion in the output. A notional 40-hour engineering week is the right unit for a staffed team and the wrong one for a solo maintainer with five hours a week: at that capacity every small item rounds to "under one week" and the Effort dimension stops discriminating between items entirely. Size in whatever a week actually buys that team, name the unit once, and keep it consistent across all items.

ICE (Impact, Confidence, Ease)

Score = Impact * Confidence * Ease

All three on 1-10 scale. Coarse but fast. Use when you need to triage 30+ ideas quickly. Do not use for committing significant capital.

MoSCoW (Must / Should / Could / Won't)

  • Must have: required for launch / release / commitment
  • Should have: important but not critical
  • Could have: nice to include if time/budget permits
  • Won't have (this time): explicitly out of scope

Strong commitment communication; weak relative ranking within buckets.

Weighted Scoring

Multi-criteria with explicit weights per criterion.

Score = Sum over criteria (Weight_i * Score_i)

Use when stakeholders disagree on what matters. Make the disagreement explicit via the weights.

Default criteria if not user-provided: business value, customer value, effort, risk, strategic fit - all at equal weight (20% each). Equal weights is itself a choice. Flag this explicitly: "These starting weights are equal; adjust them to reflect what your org actually values." Never silently apply weights.

Kano

Categorize features by how their presence / absence affects customer satisfaction:

  • Must-Have: absence causes dissatisfaction; presence is taken for granted
  • Performance: more is better in a linear way
  • Delighter: presence delights; absence does not dissatisfy
  • Reverse: presence dissatisfies (rare)
  • Indifferent: customers do not care either way

Requires customer-research input to populate categories defensibly. Gated - excluded from the run if no research input is provided (see refusal #6).

Evidence tiers, because "customer research" spans a wide range and the classification's trustworthiness varies with it. State which tier the run used, in the output, next to the categories:

Tier What it means How to run and label
Surveyed Formal Kano question pairs, functional and dysfunctional, asked per feature Classify normally. Label Kano (surveyed). Whether the categories are measured is a separate question from how they were collected: see the adequacy note below, and do not write "directly measured" until it is satisfied
Inferred Summarized research: interview themes, survey means, support-ticket patterns, or analytics that speak to satisfaction but were not collected as Kano pairs Classify, and label Kano (inferred). Say which signal drove each non-obvious category, and treat a Delighter or Must-Have call as a hypothesis to confirm rather than a finding

Do not refuse a run because the evidence is inferred rather than surveyed. Refuse only when there is no customer research at all. Downgrading to the inferred tier and saying so is more useful than excluding the framework, and it keeps the confidence claim honest.

Tier is about method; adequacy is about sample, and they are independent. A formal Kano instrument run on a handful of people is still surveyed, because that is how it was collected, and it is still not measurement. measure-survey-analysis sets n under 100 as direction-only and n under 30 as too small for segment claims, and a Kano category is a per-feature claim. So label the tier by method, then state separately what the sample supports: below those thresholds, report the categories as directional and do not use "directly measured", "validated", or a percentage breakdown of respondents by category. A small formal survey and a large one earn the same tier and different claims.

Clearing the count is necessary and not sufficient. Reaching n does not license "validated"; it only removes one reason to refuse the word. Adequacy is also about valid paired responses per feature (a respondent who skipped the dysfunctional half of a pair does not count toward that feature), who was recruited and whether they represent the population the decision applies to, and instrument quality. measure-survey-analysis is explicit that recruitment bias prevents generalization regardless of size. So a run with a large but self-selected sample is still directional, and any measured claim is limited to the people who actually answered. If you cannot say why the respondents represent the users the roadmap serves, do not write "validated" at any n.

An inferred run has no distribution, and must not invent one. Interview themes, support patterns and analytics do not produce respondent counts per Kano category. Write not available (inferred) in that column and carry the qualitative signal and its source base instead. A fabricated spread would be worse than the missing one, because it would make an inferred call look surveyed.

This skill does not define "clearly leads" as a number, and that is deliberate. A threshold that would be right for a five-item consumer backlog is wrong for a two-item enterprise one, and inventing a house cutoff here would be the same move as inventing a sample size: a number with nothing behind it, carrying more authority than the judgment it replaced. Report the distribution and let the reader see the margin. If you cannot look at the distribution and say which category leads, that is what Ambiguous is for.

And signal strength, which is the condition a clean sample can still fail. measure-survey-analysis sets its confidence label on sample, methodology and signal strength, and the third is the one an adequacy checklist tends to drop. A representative 400-response run split near-evenly across Must-Have, Performance and Delighter for a feature has cleared every condition above and still has no answer. Report the per-feature distribution, not just the winning category, and if no category clearly leads, classify that feature Ambiguous and say what would resolve it. An unstable plurality is not a roadmap input, and "validated" is the wrong word for one at any sample size.

Cross-skill composition

  • Output of this skill feeds into: a future roadmap-sequencing skill (unshipped; would rank, then sequence), deliver-launch-checklist (Must-Have items become launch criteria), sprint-planning workflows
  • Inputs to this skill often come from: develop-solution-brief, define-opportunity-tree, define-hypothesis, discover-interview-synthesis
  • Adversarial review via: utility-pm-critic (challenges assumed inputs, framework applicability, and divergence explanations)

Output Format

Use the template in references/TEMPLATE.md to structure the output. See references/EXAMPLE.md for a complete worked multi-framework run.

Quality Checklist

Before finalizing, verify:

  • At least 3 candidate items and a stated decision context
  • Applicability filter summary names which frameworks ran and which were excluded, with rationale
  • All applicable frameworks ran (not reduced to one when several apply)
  • Every score traces to a provided input or a flagged assumption (no silent fabrication)
  • Cross-framework comparison explains each divergent item by naming the driving dimension
  • Weighted Scoring (if run) loudly flags that the weights are a choice
  • Kano is excluded with an explanation when no customer research is provided
  • If Kano ran surveyed: every item carries its evidence tier and claim strength, and the per-feature response distribution is reported rather than only the winning category
  • If Kano ran inferred: the distribution column reads not available (inferred) rather than a fabricated spread, and each item names the qualitative signal that drove it and how many sources carried that signal
  • If Kano ran: any feature whose distribution shows no clear leader is recorded as Ambiguous with what would resolve it, rather than assigned a weak plurality
  • Executive summary gives a recommendation and a next step, not just a ranking

Cross-references

  • Template: references/TEMPLATE.md
  • Examples: references/EXAMPLE.md + library samples in library/skill-output-samples/define-prioritization-framework/
Files (pm-skills)
  • evals
    • trigger-fixtures.json 4 KB
      {
        "schema": 1,
        "skill": "define-prioritization-framework",
        "runs_per_query": 3,
        "trigger_threshold": 0.5,
        "queries": [
          {
            "q": "Run RICE and ICE against these 15 roadmap candidates and show me where they disagree",
            "expect": "trigger",
            "split": "train"
          },
          {
            "q": "We have a backlog of 20 feature requests; score them and tell me what to build first",
            "expect": "trigger",
            "split": "train",
            "notes": "Intent-only phrasing, no framework keyword"
          },
          {
            "q": "Rank these Q3 initiatives with MoSCoW so we can scope the release",
            "expect": "trigger",
            "split": "train"
          },
          {
            "q": "Compare our candidate list across RICE, ICE, and weighted scoring; I want to see the consensus and the outliers",
            "expect": "trigger",
            "split": "train"
          },
          {
            "q": "Leadership wants a defensible ranking of these eight initiatives before the planning offsite",
            "expect": "trigger",
            "split": "train",
            "notes": "Intent-only phrasing"
          },
          {
            "q": "We surveyed customers on the top pain points; run Kano against these feature candidates now that we have that data",
            "expect": "trigger",
            "split": "train"
          },
          {
            "q": "Score this list of technical debt items against business priority, engineering cost, and customer urgency",
            "expect": "trigger",
            "split": "validation"
          },
          {
            "q": "Two PMs disagree about what ships next quarter; run the numbers so the debate is about assumptions, not opinions",
            "expect": "trigger",
            "split": "validation",
            "notes": "Intent-only phrasing"
          },
          {
            "q": "Apply a weighted scoring model across these twelve candidates using our four criteria",
            "expect": "trigger",
            "split": "validation"
          },
          {
            "q": "Give me an executive summary recommendation after running every applicable framework on this feature list",
            "expect": "trigger",
            "split": "validation"
          },
          {
            "q": "We have forty raw ideas and no structure yet; connect them back to customer opportunities before we rank anything",
            "expect": "no-trigger",
            "split": "train",
            "near_miss_of": "define-opportunity-tree",
            "notes": "No candidate list to rank yet; the tree structures ideas first, per When NOT to Use"
          },
          {
            "q": "Build the outcome-to-opportunity-to-solution map for our retention goal before we even have a list to score",
            "expect": "no-trigger",
            "split": "validation",
            "near_miss_of": "define-opportunity-tree"
          },
          {
            "q": "Here is a messy Slack thread about a production incident; tell me what to do next and how to execute it",
            "expect": "no-trigger",
            "split": "train",
            "near_miss_of": "foundation-prioritized-action-plan",
            "notes": "A raw situation needing a full action plan, not a defined list to rank"
          },
          {
            "q": "I have an executive's rambling notes from a hallway conversation; turn this into the single most important next effort and how to run it",
            "expect": "no-trigger",
            "split": "validation",
            "near_miss_of": "foundation-prioritized-action-plan"
          },
          {
            "q": "How big is the market opportunity for this new product line? Give me TAM, SAM, and SOM",
            "expect": "no-trigger",
            "split": "train",
            "near_miss_of": "discover-market-sizing",
            "notes": "Sizing a market, not ranking a feature list, per When NOT to Use"
          },
          {
            "q": "Why does my Kubernetes pod keep getting OOMKilled under normal load?",
            "expect": "no-trigger",
            "split": "train",
            "notes": "Unrelated engineering ask"
          },
          {
            "q": "Write a regex to validate international phone numbers",
            "expect": "no-trigger",
            "split": "validation",
            "notes": "Unrelated technical ask"
          },
          {
            "q": "Recommend a good running route near Golden Gate Park",
            "expect": "no-trigger",
            "split": "train",
            "notes": "Unrelated"
          }
        ]
      }
      
  • references
    • EXAMPLE.md 6.2 KB
      ---
      artifact: prioritization-framework
      version: "1.0"
      created: 2026-05-21
      status: complete
      context: Project-management SaaS - prioritizing 6 candidate features for the Q3 roadmap with RICE + ICE + MoSCoW
      ---
      
      # Prioritization: Q3 Roadmap Candidates (Project-Management SaaS)
      
      > All reach, impact, effort, and confidence values below are illustrative `[fictional]` PM inputs for this scenario; replace them with your own estimates.
      
      ## Applicability Filter Summary
      
      We have reach, impact, effort, and confidence estimates per feature, so **RICE** and **ICE** both run. The decision also bounds Q3 scope, so **MoSCoW** runs as a commitment view. **Weighted Scoring** is excluded (no competing multi-stakeholder criteria were provided). **Kano** is excluded: no customer-research data was supplied. To unlock Kano, run a Kano survey on these six features.
      
      ## Inputs Summary
      
      Six Q3 candidate features with PM-supplied estimates. Reach is measured in affected users per quarter; effort in engineering-weeks. Confidence reflects how solid the estimates are.
      
      ## Per-Framework Scoring
      
      ### RICE
      <!-- Score = (Reach * Impact * Confidence) / Effort -->
      
      | Item | Reach (users/qtr) | Impact (0.25-3) | Confidence (%) | Effort (eng-wk) | RICE Score | Notes |
      |---|---|---|---|---|---|---|
      | Guest sharing links | 5,000 | 1 | 80% | 1 | 4,000 | Cheap, broad |
      | Bulk task editing | 8,000 | 1 | 90% | 2 | 3,600 | High-confidence quick win |
      | Mobile offline mode | 12,000 | 2 | 60% | 8 | 1,800 | Big reach, big effort |
      | SSO / SAML | 2,000 | 3 | 90% | 4 | 1,350 | Narrow reach, high per-user impact |
      | Custom dashboards | 4,000 | 2 | 70% | 5 | 1,120 | Mid on everything |
      | AI task suggestions | 15,000 | 1 | 40% | 10 | 600 | Huge reach, low confidence, high effort |
      
      ### ICE
      <!-- Score = Impact * Confidence * Ease, each 1-10 -->
      
      | Item | Impact (1-10) | Confidence (1-10) | Ease (1-10) | ICE Score | Notes |
      |---|---|---|---|---|---|
      | SSO / SAML | 9 | 9 | 6 | 486 | High-value, well-understood |
      | Guest sharing links | 6 | 8 | 9 | 432 | Easy and solid |
      | Bulk task editing | 6 | 9 | 8 | 432 | Easy and solid |
      | Custom dashboards | 7 | 7 | 5 | 245 | Middling |
      | Mobile offline mode | 8 | 6 | 3 | 144 | Valuable but hard |
      | AI task suggestions | 7 | 4 | 2 | 56 | Speculative and hard |
      
      ### MoSCoW (Q3 scope bound)
      
      | Item | Bucket | Rationale | Risk if dropped |
      |---|---|---|---|
      | SSO / SAML | Must | Three enterprise deals are blocked on it | Lose committed enterprise revenue |
      | Guest sharing links | Must | Competitive parity; churn risk without it | Continued competitive losses |
      | Bulk task editing | Should | High-value quick win | Slower power-user workflows |
      | Custom dashboards | Should | Requested but not blocking | Mild dissatisfaction |
      | Mobile offline mode | Could | Valuable but 8 eng-weeks | Mobile users wait another quarter |
      | AI task suggestions | Won't (this time) | Low confidence, 10 eng-weeks | Defer until validated |
      
      ## Per-Framework Ranking Output
      
      Each scoring table above is sorted high to low, so the per-framework ranking is the row order shown (top item first, lowest last). The side-by-side rank positions, and the items where the frameworks disagree, are consolidated in the Cross-Framework Comparison below.
      
      ## Cross-Framework Comparison
      
      | Item | RICE rank | ICE rank | MoSCoW bucket | Agreement |
      |---|---|---|---|---|
      | Guest sharing links | 1 | 2 | Must | Strong |
      | Bulk task editing | 2 | 3 | Should | Strong |
      | Mobile offline mode | 3 | 5 | Could | Divergent |
      | SSO / SAML | 4 | 1 | Must | Divergent |
      | Custom dashboards | 5 | 4 | Should | Close |
      | AI task suggestions | 6 | 6 | Won't | Strong (agree: defer) |
      
      **Divergent - SSO / SAML (RICE 4th, ICE 1st, MoSCoW Must):** RICE's Reach term punishes SSO because it only touches 2,000 users. ICE has no reach term, so SSO's high per-user impact and high confidence push it to the top. MoSCoW agrees with ICE because the 2,000 users are concentrated, high-value enterprise accounts with blocked deals. **The divergence reveals that RICE under-weights revenue-concentrated features.** This is the most important finding in the analysis.
      
      **Divergent - Mobile offline (RICE 3rd, ICE 5th):** RICE rewards the large reach (12,000); ICE penalizes the low Ease (3, an 8-week build). The driver is effort vs. reach.
      
      ## Executive Summary with Recommendation
      
      Fund **Guest sharing** and **Bulk task editing** first: they top both scored frameworks and are cheap, so they are unambiguous wins. Fund **SSO / SAML** despite its 4th-place RICE score - the RICE Reach term misleads here because the 2,000 affected users are enterprise accounts with revenue already blocked, which ICE and MoSCoW both surface. Defer **AI task suggestions** (all three frameworks agree it is not ready). The recommendation is robust except for SSO, whose ranking depends entirely on whether you weight raw reach (RICE) or strategic revenue concentration (ICE + MoSCoW); given the blocked deals, weight the latter.
      
      ## Sensitivity / What Changes the Ranking
      
      - If SSO's enterprise deals were not actually blocked, its Impact drops and it falls in ICE too, validating RICE's lower placement - so confirm the blocked-deal claim before committing.
      - If Mobile offline's effort came in at 4 eng-weeks instead of 8, its RICE score doubles to 3,600 and it jumps to a clear Should.
      - AI task suggestions stays last unless confidence rises above ~70%; a cheap spike to de-risk it would change its standing more than any other input.
      
      ## Recommendations (Sequencing)
      
      - **Fund now:** Guest sharing, Bulk task editing, SSO / SAML
      - **Fund if capacity allows:** Custom dashboards
      - **Defer:** Mobile offline (revisit if effort drops), AI task suggestions (revisit after a confidence-building spike)
      - **Data that would change this:** Confirm the SSO blocked-deal value; re-estimate Mobile offline effort; run a spike on AI suggestions
      
      ## Limitations and Biases
      
      - RICE systematically under-ranks high-value, low-reach features (the SSO problem); do not let it auto-decide enterprise/strategic items.
      - None of these frameworks measure sequencing dependencies (e.g., if SSO must ship before an enterprise launch). Pair this ranking with a roadmap view.
      - All scores rest on PM estimates; the cross-framework agreement is only as good as those inputs.
      
    • TEMPLATE.md 2.6 KB
      ---
      artifact: prioritization-framework
      version: "1.0"
      created: <YYYY-MM-DD>
      status: draft
      ---
      
      # Prioritization: [Decision Context]
      
      ## Applicability Filter Summary
      <!-- Which frameworks ran, which were excluded, and why. Note what would unlock an excluded framework. -->
      
      - **Ran:** [e.g., RICE, ICE, MoSCoW]
      - **Excluded:** [e.g., Kano - no customer research; Weighted Scoring - single criterion]
      
      ## Inputs Summary
      <!-- What you were given. Flag any missing or assumed input. -->
      
      [Item list and available data per item; note assumptions]
      
      ## Per-Framework Scoring
      
      ### RICE
      <!-- Score = (Reach * Impact * Confidence) / Effort -->
      
      | Item | Reach (users/qtr) | Impact (0.25-3) | Confidence (%) | Effort (eng-wk) | RICE Score | Notes |
      |---|---|---|---|---|---|---|
      | [Item A] | [N] | [0.25-3] | [%] | [N] | [score] | [note] |
      
      ### ICE
      <!-- Score = Impact * Confidence * Ease, each 1-10 -->
      
      | Item | Impact (1-10) | Confidence (1-10) | Ease (1-10) | ICE Score | Notes |
      |---|---|---|---|---|---|
      | [Item A] | [N] | [N] | [N] | [score] | [note] |
      
      ### MoSCoW
      
      | Item | Bucket (Must/Should/Could/Won't) | Rationale | Risk if dropped |
      |---|---|---|---|
      | [Item A] | [Bucket] | [Why] | [Risk] |
      
      <!-- Add Weighted Scoring and/or Kano tables here only if those frameworks passed the applicability filter. -->
      
      ## Per-Framework Ranking Output
      <!-- For each scored framework, list items sorted high-to-low by score. With 5+ items, call out the top 5 and bottom 5 and name the score gap between them. (If the scoring tables above are already sorted, summarize the ranking here rather than repeating.) -->
      
      - **RICE ranking:** [items sorted by score, high to low]
      - **ICE ranking:** [items sorted by score, high to low]
      
      ## Cross-Framework Comparison
      <!-- Rank position per item across frameworks. Explain every Divergent row by naming the driving dimension. -->
      
      | Item | RICE rank | ICE rank | MoSCoW bucket | Agreement |
      |---|---|---|---|---|
      | [Item A] | [1] | [1] | [Must] | [Strong] |
      | [Item B] | [2] | [8] | [Should] | [Divergent - why] |
      
      ## Executive Summary with Recommendation
      <!-- 3-5 sentences: what to prioritize, what to defer, what the key divergence means -->
      
      [Recommendation]
      
      ## Sensitivity / What Changes the Ranking
      <!-- 2-3 cases where the order flips, focused on items near the cut line -->
      
      - [If Confidence on Item X is wrong, then ...]
      - [If Effort on Item Y doubles, then ...]
      
      ## Recommendations (Sequencing)
      
      - **Fund now:** [Items]
      - **Defer / drop:** [Items]
      - **Data that would change this:** [What to gather]
      
      ## Limitations and Biases
      <!-- What the frameworks do NOT measure; where they could mislead -->
      
      - [Limitation 1]
      - [Limitation 2]
      
  • HISTORY.md 7.2 KB
    # define-prioritization-framework - Version History
    
    | Version | Date | Release | Effort | Type | Summary |
    |---------|------|---------|--------|------|---------|
    | 1.3.0 | 2026-08-16 | v2.33.0 | C-14 | minor | Field-reported calibration (#252): top/bottom highlight rule scales below 10 items, RICE Effort unit scales to real team capacity, Kano gains surveyed and inferred evidence tiers. |
    | 1.2.0 | 2026-07-05 | v2.31.0 | WS-Z5 | minor | Reciprocal When NOT to Use pointer to `foundation-prioritized-action-plan`; collision pair declared with new trigger fixtures. |
    | 1.1.0 | 2026-07-04 | v2.30.0 | M-35 | minor | Added a "When NOT to Use" section with five reciprocal boundary pointers, including the bidirectional edge back to `define-opportunity-tree` (which already pointed here). Closes a one-way gap in the cross-skill reciprocity mesh flagged by the 2026-07-04 deep audit. Also normalized the "Output format" and "Quality checklist" headings to their canon spelling and resolved a phantom `deliver-roadmap` pointer in Cross-skill composition (both WS-T8b/f, no re-bump). |
    | 1.0.0 | 2026-05-21 | v2.18.0 | - | baseline | Prior published version: runs the applicable prioritization frameworks (RICE, ICE, MoSCoW, Weighted Scoring, Kano) against a candidate list, filtered by data availability, surfacing where rankings agree and diverge plus an executive recommendation. |
    
    ## 1.3.0 (2026-08-16)
    
    Field-reported calibration ([#252](https://github.com/product-on-purpose/pm-skills/issues/252)), from an end-to-end run on a real 8-item backlog in a solo-maintainer context. The report was positive about the skill overall and singled out the convergence and divergence analysis as its most valuable output, so these are calibration fixes rather than defect repairs.
    
    **Top and bottom highlight rule now scales below 10 items.** The rule read "highlight the top 5 and bottom 5"; with 8 items, 5 and 5 overlap or exhaust the list and the rule cannot be followed as written. At 10 or fewer items the skill now shows every item in rank order and describes the gap between clear tiers instead of forcing a five-and-five split.
    
    **RICE Effort unit now scales to real team capacity.** The unit was fixed as eng-weeks or person-weeks. For a solo operator with roughly five hours a week, a notional 40-hour week makes every small item round to "under one week" and the Effort dimension stops discriminating entirely. The unit is now capacity-weeks, sized to what a week actually buys that team, with the conversion stated in the output.
    
    **Kano gains explicit evidence tiers.** The applicability gate asked for "customer-research input" without saying whether summarized signals qualified or whether formal Kano question pairs were required. It now distinguishes a surveyed tier (functional and dysfunctional pairs per feature, classify normally) from an inferred tier (interview themes, survey means, support patterns; classify, label as inferred, and treat a Delighter or Must-Have call as a hypothesis to confirm). Refusal is reserved for having no customer research at all, since downgrading the tier and saying so is more useful than excluding the framework.
    
    **Kano unlock suggestion no longer leaves a bare pointer** ([#253](https://github.com/product-on-purpose/pm-skills/issues/253), folded into this same unreleased version). When Kano is excluded for having no research at all, the refusal suggests `discover-interview-synthesis` or `measure-survey-analysis` to unlock it. Under a partial install neither may exist, so the refusal now names the research in plain language as a fallback: asking, per feature, how the user would feel if it were present and if it were absent. It deliberately does not name a sample size. An earlier draft said roughly 20 to 30 responses, which contradicted this repo's own contract in `measure-survey-analysis` (n < 100 is direction-only, n < 30 is too small for segment claims) for a Kano category that is itself a per-feature claim. The fallback now points at the surveyed and inferred evidence tiers instead, so the sample earns its tier rather than a number implying measurement.
    
    Minor rather than patch: the Kano tiering lets the skill run a case it previously refused or fudged, and adds an optional tier label to the output, which is additive behavior under the versioning tie-breaker.
    
    ## 1.2.0 (2026-07-05)
    
    Released in [v2.31.0](../../site/src/content/docs/releases/Release_v2.31.0.md). Effort: WS-Z5 (eval backfill wave 1, R-16).
    
    The WS-Z5 fixture backfill declared `foundation-prioritized-action-plan` as a new collision pair for this skill in `scripts/trigger-eval-roster.yaml`, but the reciprocal "When NOT to Use" pointer was never added. The enforcing `check-reciprocal-boundary-pointers` gate caught the gap. Adds one bullet distinguishing a raw, unstructured situation from a defined candidate list ready for formal framework scoring. No other content change.
    
    ## 1.1.0 (2026-07-04)
    
    Released in [v2.30.0](../../site/src/content/docs/releases/Release_v2.30.0.md). Effort: M-35 (trust repair sweep).
    
    The 2026-07-04 deep audit found this skill had no "When NOT to Use" section at all, and that `define-opportunity-tree`'s existing pointer to it was one-directional (opportunity-tree pointed here; nothing pointed back). This release adds the section and closes that edge.
    
    ### Changes
    - Added a "When NOT to Use" section with pointers to `define-opportunity-tree`, `define-hypothesis`, `measure-experiment-design`, `discover-market-sizing`, `deliver-launch-checklist`, and `discover-interview-synthesis`.
    - The `define-opportunity-tree` <-> `define-prioritization-framework` edge is now bidirectional (opportunity-tree required no edit; it already pointed here).
    - Heading-normalization sweep (WS-T8b, folded into this same v2.30.0 row rather than a separate bump): "Output format" to "Output Format" and "Quality checklist" to "Quality Checklist", two of the catalog's drifted heading-spelling instances the 2026-07-04 deep audit flagged.
    - Dedup fix (WS-T8f, folded into this same v2.30.0 row rather than a separate bump): resolved the phantom `deliver-roadmap` pointer in Cross-skill composition. `deliver-roadmap` is not a shipped skill; the backtick-wrapped reference read as a resolvable link when it was an intentional forward-reference (it remains allowlisted in `scripts/check-skill-cross-references.sh` for exactly this reason). Reworded to name a future roadmap-sequencing capability in plain prose, with no link.
    
    No change to the framework-scoring flow, refusal protocols, or output contract.
    
    ## 1.0.0 (2026-05-21)
    
    Released in [v2.18.0](../../site/src/content/docs/releases/Release_v2.18.0.md).
    
    Initial release: runs all applicable prioritization frameworks (RICE, ICE, MoSCoW, Weighted Scoring, Kano) against a candidate list, filtered by data availability and context, then produces a cross-framework comparison and an executive recommendation. Kano is gated on customer research; missing inputs produce an estimation scaffold rather than fabricated scores.
    
    ### Contract established
    - Filters frameworks by applicability rather than reducing to one
    - Refuses to fabricate scores; produces an estimation scaffold when input data is missing
    - Output: per-framework scoring tables, cross-framework comparison, executive summary, sensitivity analysis
    
  • SKILL.md 19.7 KB
    ---
    name: define-prioritization-framework
    description: Run applicable prioritization frameworks (RICE, ICE, MoSCoW, Weighted Scoring, Kano) against a list of features or initiatives. Produces a comparison table showing where rankings agree and diverge across frameworks, and an executive summary with recommendation. Framework applicability is filtered by data availability; Kano requires customer research. Refuses to fabricate scores; produces an estimation scaffold when input data is missing.
    license: Apache-2.0
    metadata:
      phase: define
      version: "1.3.0"
      updated: 2026-08-16
      category: planning
      frameworks: [triple-diamond, prioritization]
      author: product-on-purpose
    ---
    <!-- PM-Skills | https://github.com/product-on-purpose/pm-skills | Apache 2.0 -->
    # Prioritization Framework
    
    You run all applicable prioritization frameworks against a candidate list of work items. Your job is to (a) filter frameworks by data availability and context, (b) score each item explicitly per applicable framework, (c) produce a comparison table showing where rankings agree and diverge, (d) synthesize an executive summary with recommendation, and (e) flag what could go wrong with the prioritization.
    
    ## Identity
    
    - Phase skill (define); Triple Diamond integration
    - Single-turn lifetime; produces one ranked artifact per invocation
    - Read-only tools (Read, Grep); no write outside the output artifact
    - Outputs a markdown document with per-framework scoring tables + comparison + recommendation
    
    ## Core principle
    
    **Multi-framework analysis surfaces what single-framework selection hides.** Where RICE and ICE agree, confidence rises. Where they disagree, the divergence reveals hidden assumptions worth examining - often the most valuable finding.
    
    Filter frameworks by applicability: RICE requires quantitative reach/impact/effort inputs; ICE works with coarse estimates; MoSCoW is for binary commitment decisions; Weighted Scoring requires multi-criteria weights; Kano requires customer-research input (gated). Run all frameworks that pass the applicability filter. Do NOT reduce to one framework when multiple are applicable.
    
    ## When NOT to Use
    
    - You have not yet structured outcomes and opportunities into a candidate list -> use `define-opportunity-tree`; this skill ranks a list, it does not discover what belongs on it
    - You want to test one specific assumption rather than rank several items -> use `define-hypothesis`, then `measure-experiment-design`
    - You need to size a market opportunity (TAM/SAM/SOM), not rank a feature list -> use `discover-market-sizing`
    - Your items are already ranked and you need launch readiness next -> use `deliver-launch-checklist`
    - You need qualitative synthesis of user research to generate candidates, not rank an existing list -> use `discover-interview-synthesis`
    - You have a raw, unstructured situation (notes, transcript, exec ask) rather than a defined candidate list of items to score -> use `foundation-prioritized-action-plan` for a general ranked next-action plan; this skill requires a candidate list and scores it against formal frameworks
    
    ## Inputs
    
    Required:
    
    - List of candidate items (features, initiatives, work items). Each item needs at least a name and a one-sentence description.
    - Decision context: "Q3 roadmap candidates" or "MVP scope reduction" or "Hypothesis triage for the next sprint" etc.
    
    Optional but improves quality:
    
    - Available data per item (impact estimate, effort estimate, customer signal, business case)
    - Stakeholder criteria (engineering capacity, business priority, customer urgency)
    - Confidence levels on input data
    - Time horizon (sprint, quarter, half, year)
    - Customer-research data (unlocks Kano)
    
    ## Framework applicability filter
    
    Before running, evaluate each framework against the available inputs. Run all frameworks that pass:
    
    | Framework | Runs when | Excluded when |
    |---|---|---|
    | **RICE** (Reach * Impact * Confidence / Effort) | Quantitative reach, impact, effort estimates are available or user accepts an estimation scaffold | Inputs unavailable and user declines estimation scaffold |
    | **ICE** (Impact * Confidence * Ease) | Always applicable; coarse estimates are acceptable | Not excluded; ICE is the lowest-input framework |
    | **MoSCoW** (Must / Should / Could / Won't) | Decision involves binary commitment per item or scope bounding | Not applicable for pure ranking decisions without scope constraint |
    | **Weighted Scoring** (multi-criteria with weights) | Multiple stakeholders or criteria apply; user provides or accepts proposed default weights | Single criterion dominates; or criteria are purely personal preference |
    | **Kano** (Must-Have / Performance / Delighter) | Customer-research input is provided, at either evidence tier below | **Gated:** excluded only if no customer research at all is provided; explain why and suggest what research would unlock it. Run at the tier the evidence supports and label the tier in the output |
    
    At least one framework will always run (ICE is always applicable). Show which frameworks ran and which were excluded, with brief rationale.
    
    ## What you produce
    
    ### 1. Applicability filter summary (3-5 sentences)
    
    Which frameworks ran, which were excluded, and why. Note any frameworks excluded due to missing inputs and what would unlock them.
    
    ### 2. Inputs summary
    
    What you were given. If any input is missing or assumed, note: "Reach was not provided; assumption: large reach unless flagged."
    
    ### 3. Per-framework scoring tables
    
    Run each applicable framework and produce its scoring table.
    
    **For RICE:**
    
    | Item | Reach (users/qtr) | Impact (0.25-3) | Confidence (%) | Effort (capacity-weeks) | RICE Score | Notes |
    |---|---|---|---|---|---|---|
    | Item A | 1000 | 2 | 80% | 3 | 533 | High confidence on reach |
    
    **For ICE:**
    
    | Item | Impact (1-10) | Confidence (1-10) | Ease (1-10) | ICE Score | Notes |
    |---|---|---|---|---|---|
    
    **For MoSCoW:**
    
    | Item | Bucket | Rationale | Risk if dropped |
    |---|---|---|---|
    | Item A | Must | Critical for launch | Cannot ship without |
    
    **For Weighted Scoring:**
    
    | Item | Criterion 1 (weight) | Criterion 2 (weight) | ... | Total Weighted Score |
    |---|---|---|---|---|
    
    **For Kano:**
    
    | Item | Category (Must / Performance / Delighter / Reverse / Indifferent / **Ambiguous**) | Response distribution (surveyed runs; `not available (inferred)` otherwise) | Evidence tier and claim strength | Customer evidence | Implication |
    |---|---|---|---|---|---|
    
    ### 4. Per-framework ranking output
    
    For each scored framework: items sorted by score or grouped by bucket. For scored frameworks, highlight the top 5 and bottom 5 with the gap between them. **When the backlog has 10 or fewer items**, a top 5 and a bottom 5 overlap or exhaust the list, so the rule cannot be followed as written: show every item in rank order instead and describe the gap between the clear tiers rather than forcing a five-and-five split.
    
    ### 5. Cross-framework comparison
    
    A comparison table showing ranking position per item across all frameworks that ran. Surface divergence explicitly.
    
    | Item | RICE rank | ICE rank | MoSCoW bucket | Agreement |
    |---|---|---|---|---|
    | Item A | 1 | 1 | Must | Strong |
    | Item B | 2 | 8 | Should | Divergent |
    
    For each Divergent item: explain the driver. Divergence usually means one scoring dimension is carrying most of the weight (e.g., ICE ranks item B 8th because Ease is very low, but RICE ranks it 2nd because Reach is massive). This is the finding.
    
    ### 6. Executive summary with recommendation
    
    Synthesize the comparison into a 3-5 sentence recommendation: which items to prioritize, which to defer, and what the most important divergence means for the team's decision. Flag if the recommendation changes materially under different frameworks or assumptions.
    
    ### 7. Sensitivity / what changes the ranking
    
    What if Confidence is wrong? What if Effort is doubled? Show 2-3 cases where the rank order changes, focusing on the items near the cut line.
    
    ### 8. Recommendations (sequencing)
    
    Top items to fund; bottom items to defer or drop; what additional data would change the recommendation. Recommend NEXT STEP, not just the ranking.
    
    ### 9. Limitations and biases
    
    What are these frameworks NOT measuring? Where could the frameworks lead astray? Where do they systematically favor certain item types over others?
    
    ## Refusal protocols
    
    You refuse to produce a ranking without minimum input quality. Specifically:
    
    1. **Empty / single-item list.** If user provides 0 or 1 candidate items: "Prioritization requires at least 3 items to be meaningful. With fewer, just decide directly."
    
    2. **No context.** If user provides items without saying what decision they are making: "I need to know what decision this prioritization is supporting. Sprint scope? Quarter scope? Hypothesis triage? Different contexts affect which frameworks apply."
    
    3. **Missing numerical inputs for RICE.** If user asks for RICE scores without providing input data: "I cannot produce defensible RICE scores without reach, impact, confidence, and effort estimates. Options: (a) provide rough numbers per item; (b) I can produce an estimation scaffold - a structured worksheet showing how to estimate reach, impact, confidence, and effort for each item; (c) run ICE instead, which works with coarse 1-10 judgment and does not require quantitative inputs. Which would you prefer?" (ICE itself is never refused for missing data - it is the always-applicable coarse fallback.)
    
    4. **Wrong-framework insistence.** If user insists on RICE for an early-stage hypothesis triage: "RICE assumes measurable impact and effort, which you do not have at this stage. I can produce a RICE table but the scores will be guesses. ICE or MoSCoW would be more honest. Want to proceed with RICE anyway, or switch?"
    
    5. **Single-stakeholder weighted scoring.** If user asks for Weighted Scoring with criteria that only one stakeholder cares about: "Weighted Scoring is for multi-stakeholder trade-offs. If only one stakeholder's criteria apply, RICE or ICE would be simpler. Want to proceed or switch?"
    
    6. **Kano without customer research.** If user requests Kano but provides no customer-research input: "Kano categories are only defensible with customer research. Without it, you would be guessing whether a feature is a Must-Have or a Delighter, which defeats the purpose. I have excluded Kano from this run. The other applicable frameworks have run above. To unlock Kano, provide customer survey or interview data (skill: `discover-interview-synthesis` or `measure-survey-analysis`)." If neither skill is available in the environment, do not leave the pointer bare: name the research in plain language instead, which is asking, per feature, how the user would feel if it were present and how they would feel if it were absent. Be honest about what a small run buys rather than naming a number that implies more than it delivers. The tier is set by how you collect (see the evidence tiers below), and what you may claim is set separately by how many people answered; a small formal instrument is still a surveyed run and still not measurement. Report the tier and the claim strength as two separate statements.
    
    ## Framework details
    
    ### RICE (Reach, Impact, Confidence, Effort)
    
    `Score = (Reach * Impact * Confidence) / Effort`
    
    - Reach: how many users / customers / events affected per time period (per quarter is common). Number, not %.
    - Impact: how much each affected user benefits. Use Intercom's scale: 0.25 (minimal), 0.5 (low), 1 (medium), 2 (high), 3 (massive).
    - Confidence: how sure you are about the other estimates. 0-100%.
    - Effort: how much work it takes in capacity-weeks. Higher = lower score. **Scale the unit to the executing team's real weekly capacity and state the conversion in the output.** A notional 40-hour engineering week is the right unit for a staffed team and the wrong one for a solo maintainer with five hours a week: at that capacity every small item rounds to "under one week" and the Effort dimension stops discriminating between items entirely. Size in whatever a week actually buys that team, name the unit once, and keep it consistent across all items.
    
    ### ICE (Impact, Confidence, Ease)
    
    `Score = Impact * Confidence * Ease`
    
    All three on 1-10 scale. Coarse but fast. Use when you need to triage 30+ ideas quickly. Do not use for committing significant capital.
    
    ### MoSCoW (Must / Should / Could / Won't)
    
    - Must have: required for launch / release / commitment
    - Should have: important but not critical
    - Could have: nice to include if time/budget permits
    - Won't have (this time): explicitly out of scope
    
    Strong commitment communication; weak relative ranking within buckets.
    
    ### Weighted Scoring
    
    Multi-criteria with explicit weights per criterion.
    
    `Score = Sum over criteria (Weight_i * Score_i)`
    
    Use when stakeholders disagree on what matters. Make the disagreement explicit via the weights.
    
    **Default criteria if not user-provided:** business value, customer value, effort, risk, strategic fit - all at equal weight (20% each). **Equal weights is itself a choice.** Flag this explicitly: "These starting weights are equal; adjust them to reflect what your org actually values." Never silently apply weights.
    
    ### Kano
    
    Categorize features by how their presence / absence affects customer satisfaction:
    
    - Must-Have: absence causes dissatisfaction; presence is taken for granted
    - Performance: more is better in a linear way
    - Delighter: presence delights; absence does not dissatisfy
    - Reverse: presence dissatisfies (rare)
    - Indifferent: customers do not care either way
    
    Requires customer-research input to populate categories defensibly. **Gated** - excluded from the run if no research input is provided (see refusal #6).
    
    **Evidence tiers, because "customer research" spans a wide range and the classification's trustworthiness varies with it.** State which tier the run used, in the output, next to the categories:
    
    | Tier | What it means | How to run and label |
    |---|---|---|
    | **Surveyed** | Formal Kano question pairs, functional and dysfunctional, asked per feature | Classify normally. Label **Kano (surveyed)**. Whether the categories are *measured* is a separate question from how they were collected: see the adequacy note below, and do not write "directly measured" until it is satisfied |
    | **Inferred** | Summarized research: interview themes, survey means, support-ticket patterns, or analytics that speak to satisfaction but were not collected as Kano pairs | Classify, and label **Kano (inferred)**. Say which signal drove each non-obvious category, and treat a Delighter or Must-Have call as a hypothesis to confirm rather than a finding |
    
    Do not refuse a run because the evidence is inferred rather than surveyed. Refuse only when there is no customer research at all. Downgrading to the inferred tier and saying so is more useful than excluding the framework, and it keeps the confidence claim honest.
    
    **Tier is about method; adequacy is about sample, and they are independent.** A formal Kano
    instrument run on a handful of people is still **surveyed**, because that is how it was collected,
    and it is still not measurement. `measure-survey-analysis` sets n under 100 as direction-only and n
    under 30 as too small for segment claims, and a Kano category is a per-feature claim. So label the
    tier by method, then state separately what the sample supports: below those thresholds, report the
    categories as **directional** and do not use "directly measured", "validated", or a percentage
    breakdown of respondents by category. A small formal survey and a large one earn the same tier and
    different claims.
    
    **Clearing the count is necessary and not sufficient.** Reaching n does not license "validated"; it
    only removes one reason to refuse the word. Adequacy is also about **valid paired responses per
    feature** (a respondent who skipped the dysfunctional half of a pair does not count toward that
    feature), **who was recruited** and whether they represent the population the decision applies to,
    and **instrument quality**. `measure-survey-analysis` is explicit that recruitment bias prevents
    generalization regardless of size. So a run with a large but self-selected sample is still
    directional, and any measured claim is limited to the people who actually answered. If you cannot
    say why the respondents represent the users the roadmap serves, do not write "validated" at any n.
    
    **An inferred run has no distribution, and must not invent one.** Interview themes, support
    patterns and analytics do not produce respondent counts per Kano category. Write
    `not available (inferred)` in that column and carry the qualitative signal and its source base
    instead. A fabricated spread would be worse than the missing one, because it would make an inferred
    call look surveyed.
    
    **This skill does not define "clearly leads" as a number, and that is deliberate.** A threshold
    that would be right for a five-item consumer backlog is wrong for a two-item enterprise one, and
    inventing a house cutoff here would be the same move as inventing a sample size: a number with
    nothing behind it, carrying more authority than the judgment it replaced. Report the distribution
    and let the reader see the margin. If you cannot look at the distribution and say which category
    leads, that is what **Ambiguous** is for.
    
    **And signal strength, which is the condition a clean sample can still fail.**
    `measure-survey-analysis` sets its confidence label on sample, methodology **and signal strength**,
    and the third is the one an adequacy checklist tends to drop. A representative 400-response run
    split near-evenly across Must-Have, Performance and Delighter for a feature has cleared every
    condition above and still has no answer. Report the per-feature distribution, not just the winning
    category, and if no category clearly leads, classify that feature **Ambiguous** and say what would
    resolve it. An unstable plurality is not a roadmap input, and "validated" is the wrong word for
    one at any sample size.
    
    ## Cross-skill composition
    
    - Output of this skill feeds into: a future roadmap-sequencing skill (unshipped; would rank, then sequence), `deliver-launch-checklist` (Must-Have items become launch criteria), sprint-planning workflows
    - Inputs to this skill often come from: `develop-solution-brief`, `define-opportunity-tree`, `define-hypothesis`, `discover-interview-synthesis`
    - Adversarial review via: `utility-pm-critic` (challenges assumed inputs, framework applicability, and divergence explanations)
    
    ## Output Format
    
    Use the template in `references/TEMPLATE.md` to structure the output. See `references/EXAMPLE.md` for a complete worked multi-framework run.
    
    ## Quality Checklist
    
    Before finalizing, verify:
    
    - [ ] At least 3 candidate items and a stated decision context
    - [ ] Applicability filter summary names which frameworks ran and which were excluded, with rationale
    - [ ] All applicable frameworks ran (not reduced to one when several apply)
    - [ ] Every score traces to a provided input or a flagged assumption (no silent fabrication)
    - [ ] Cross-framework comparison explains each divergent item by naming the driving dimension
    - [ ] Weighted Scoring (if run) loudly flags that the weights are a choice
    - [ ] Kano is excluded with an explanation when no customer research is provided
    - [ ] If Kano ran **surveyed**: every item carries its evidence tier and claim strength, and the per-feature response distribution is reported rather than only the winning category
    - [ ] If Kano ran **inferred**: the distribution column reads `not available (inferred)` rather than a fabricated spread, and each item names the qualitative signal that drove it and how many sources carried that signal
    - [ ] If Kano ran: any feature whose distribution shows no clear leader is recorded as **Ambiguous** with what would resolve it, rather than assigned a weak plurality
    - [ ] Executive summary gives a recommendation and a next step, not just a ranking
    
    ## Cross-references
    
    - Template: `references/TEMPLATE.md`
    - Examples: `references/EXAMPLE.md` + library samples in `library/skill-output-samples/define-prioritization-framework/`
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related