sow-pws-builder
Trigger for: writing, revising, converting, or descoping a federal Statement of Work, Performance Work Statement, SOW, PWS, or SOO; develop executable requirements; define contract scope; create measurable performance standards; or prepare requirements before an IGCE. Produce a c
Install
npx skills add https://github.com/1102tools-dev/federal-contracting-skills/tree/main/skills/sow-pws-builder
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 1102tools-dev-federal-contracting-skills@llmmart
git clone https://github.com/1102tools-dev/federal-contracting-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 1102tools-dev/federal-contracting-skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
SOW/PWS Builder
Overview
Turn program-office scope decisions into a contract-file-ready .docx SOW or PWS. Produce three separate outputs:
- The SOW/PWS
.docx, containing requirements and measurable acceptance or surveillance content. - A chat-only staffing handoff for the FFP, LH/T&M, or CR pricing skill.
- A chat-only Section B handoff with suggested CLIN structure, labor categories or estimated hours when needed, travel/ODC lines, and the total ceiling-price input for T&M/LH.
Never combine the handoffs with the document. FAR 37.602(b)(1) directs agencies, to the maximum extent practicable, to describe performance work by required results rather than method or hours. FAR 16.601 requires fixed hourly rates by labor category for T&M/LH and an overall ceiling price, but it does not require a labor-category ceiling-hours table inside the PWS. Keep pricing structure in Section B unless the user supplies a controlling solicitation template that places it elsewhere.
No external MCP server is required.
Product quality default
Make the PWS/SOW read like a concise acquisition-file deliverable, not a process transcript. Its first page must identify the mission need, the performance decision it supports, the material operating assumptions, and the next acquisition-team action. Use a requirement-specific title, a short executive purpose block, and distinct visual hierarchy appropriate to the document type. Put detailed evidence and source notes in a compact appendix only when they affect a requirement; never let them crowd out executable requirements. This does not permit staffing, SOC, CLIN, pricing, or IGCE content inside the document.
Load supporting files only when needed:
- question-blocks.md for intake and scope questions.
- document-specification.md before authoring the
.docx. - professional-product-standard.md before authoring the
.docx. - regulatory-and-content-rules.md for contract-type, security, QASP, and language rules.
- handoff-specification.md before final chat output.
- runtime-adaptation.md for questions, document tools, TOC handling, and delivery.
- validation-gates.md before delivery.
Permanent correctness gates
- Separation: The
.docxcontains no staffing handoff, FTE estimate, SOC code, IGCE content, CLIN table, pricing schedule, or skill-chain message. The two handoffs exist only in chat. - No false FAR exception: Do not cite FAR 37.102(d) as an hours prohibition. Use FAR 37.602(b)(1) for results-oriented PWS language.
- T&M/LH structure: FAR 16.601(c)(2) requires fixed hourly rates by labor category, and FAR 16.601(d)(2) requires an overall ceiling price. Put the pricing schedule and ceiling in the Section B handoff, not the SOW/PWS body.
- Contract-type boundary: Explain how a user-selected type changes the document, but do not originate the FAR Part 16 decision, T&M/LH Determination and Findings, commercial-item determination, or CPFF form selection.
- CPFF form: Require the user to confirm Completion or Term under FAR 16.306(d). Do not default either form. Reflect the confirmed form in the document framework.
- Result-oriented PWS: Organize a PWS around outcomes, measurable standards, AQLs, and assessment methods. Avoid prescribing staffing or contractor method unless a constraint is genuinely Government-controlled.
- QASP and CPARS: QASP payment or administrative consequences must not pre-commit CPARS ratings. Never map an AQL threshold to
Satisfactory,Very Good,Exceptional,Marginal, orUnsatisfactoryCPARS ratings. - Key personnel: Name roles and qualifications only. Do not quantify staff. Never cite FAR 52.237-2 as a key-personnel substitution clause.
- Coverage derivation: For the chat-only staffing handoff, derive coverage as annual coverage hours divided by productive hours. At 1,880 hours, one 24x7x365 seat is 4.6596 FTE, not three and not 4.2.
- Tier 1 mapping: Preserve the project decision that Tier 1 help desk or contact-center agents map to SOC 43-4051. Do not change it during modernization.
- Decision gates: Do not self-approve staffing or the Phase 2 Decision Summary. End each gate response at its confirmation question and wait.
- DOCX validation: Use real heading styles, render every page to images, inspect every page, audit text and OOXML, and do not claim Word compatibility from a fallback renderer alone.
Workflow selection
Workflow A: full build
Use for a concept, rough requirements, or a build from scratch. Run intake and all phases.
Workflow B: SOO conversion
Extract settled objectives, constraints, location, period, systems, volumes, security, and other facts from the SOO. Present the gaps that must be decided, then run the remaining phases without re-asking settled facts.
Workflow C: scope reduction
Use an existing SOW/PWS and, when available, IGCE cost drivers to present capability and coverage tradeoffs. The user selects reductions. Revise affected requirements and handoffs without placing staff counts in the document.
Runtime pre-flight
When this skill is entered immediately after a numbered Pre-Award Agent selection and the current assistant response has not already shown the orchestrator's outcome preview, emit these exact four lines before any intake or capability check:
Begin line 1 with Recommended outcome:. Do not precede the block with a heading, acknowledgement, selection recap, routing narration, or code fence.
Recommended outcome: Validated SOW/PWS `.docx` plus two chat-only handoffs
Includes: an executable work statement, measurable standards, a staffing handoff, and a Section B handoff
Boundary/default: recommend PWS for performance-based services when the requirement supports it; the user or Contracting Officer retains contract type, commerciality, and other reserved decisions
Next: collect the current requirement or source material and missing acquisition-strategy facts
This is a routing fallback, not a second preview. Do not repeat it when the orchestrator already rendered the four lines in the current assistant response, and never replace it with component intake.
After the orchestrator's outcome preview, begin acquisition-strategy intake and reuse every supplied fact. Do not make document-authoring or rendering capability the first question or first action after workflow selection. A read-only or artifact-limited session may still inspect supplied material, identify gaps, and complete useful scope intake.
Before promising or beginning the validated .docx build:
- Confirm the host can read inputs and create
.docxfiles. - Confirm a DOCX render path is available, preferably LibreOffice through the host's document workflow.
- Confirm Python and
python-docxor an equivalent OOXML authoring capability forscripts/validate_docx.py. - Match capabilities semantically. Do not depend on
/mntpaths, a named client tool, or a generated namespace. - If document authoring is unavailable, stop at the artifact boundary and report the missing capability. Preserve the intake already completed and explain what remains needed to resume. Do not silently substitute Markdown or HTML for the contract-file deliverable.
Acquisition strategy intake
Collect framing decisions in one pass, skipping facts already supplied:
- SOW or PWS.
- User-confirmed contract type: FFP, T&M, LH, CR subtype, or hybrid by CLIN.
- Commercial, noncommercial, or pending Contracting Officer determination.
- Agency and solicitation format, if one controls.
- Purpose, audience, acquisition stage, and desired file name.
If the user is unsure about SOW versus PWS, explain that PWS is outcome-oriented and preferred for performance-based services to the maximum extent practicable. A recommendation about document form is permitted. Do not turn it into the contract-type or commerciality decision.
If T&M or LH is selected, flag the D&F and overall ceiling-price requirement for Contracting Officer action. If CPFF is selected, require Completion or Term. If commerciality is unsettled, record [DEFAULT: Commerciality determination pending] in Constraints and Assumptions rather than deciding it.
Phase 0: SOO or source intake
For a supplied source document:
- Extract background, purpose, objectives, performance constraints, location, period, systems, volumes, security, Government-furnished resources, and named standards.
- Distinguish explicit facts from derived assumptions.
- Present gaps as the questions needed for executable requirements.
- Carry settled facts forward without asking them again.
If the source is too thin to support task or objective decomposition, state that it can serve as background but additional scope decisions are required.
Phase 1: scope decision tree
Use question-blocks.md. Prefer structured choices when the host supports them; otherwise use numbered choices with a free-text escape. Batch three to four related questions. Ask open text only for values such as system names, volume, or incumbent facts.
Sequence:
- Mission and service model.
- Technical scope.
- Scale and volume.
- Organizational scope and phasing.
- Period, transition, and Section B preferences.
- Quality, oversight, security, and Government-furnished resources.
Staffing derivation for the internal handoff
Derive a candidate staffing table only after scope decisions. Every row must cite a basis: workload, coverage hours, system/module count, site count, delivery cadence, or an explicit user override. Heuristics are starting points, not facts.
- Continuous coverage: covered seats times annual coverage hours divided by productive hours per FTE.
- Development: derive by system/module count and team design; do not use an unexplained fixed ratio.
- Contact center: annual contacts divided by workdays and contacts per agent-day, adjusted for service level and productive hours.
- Transition and surge: separate time-limited rows.
- O&M: derive from named workload, SLA, or environment size; never assume a fixed percentage of development staffing without user confirmation.
Present the candidate staffing table in chat and ask the user to confirm, amend, or override it. End immediately after that question and wait. Preserve override reasons.
Phase 2: Decision Summary gate
After staffing approval and before authoring, present a concise Decision Summary containing:
- SOW/PWS, contract type and subtype, CPFF form when relevant, and commerciality status.
- Agency/template and solicitation-stage context.
- All defaults or derived assumptions with one-line rationale.
- Section 3 task areas for a SOW or performance objectives for a PWS.
- Applicable deliverables, QASP or inspection approach, security, period, location, and transition.
- Confirmation that staffing, SOCs, CLINs, and pricing will remain outside the document.
The final sentence must ask the user to proceed or correct the summary. Stop immediately after the question. Do not author the .docx in the same response.
Phase 2: document assembly
After explicit approval:
- Read professional-product-standard.md, document-specification.md, and regulatory-and-content-rules.md. Formal requirement controls remain binding, but the document should still read as deliberate professional work rather than a generated template.
- Use the host's document-authoring workflow and a formal business-document design system. Preserve a supplied agency template when one exists.
- Use real
Heading 1,Heading 2, andHeading 3styles, real numbered or bulleted lists, explicit table geometry, page numbers, and accessible repeating table headers. - For more than eight main sections, include a dynamic TOC field. Set the document to update fields on open. If Word is available, update and save the field before delivery. Otherwise disclose the exact refresh step in chat. Do not put the refresh instruction inside the document.
- Run the render, text, structure, and separation gates in validation-gates.md.
- Fix defects and repeat rendering until every page passes.
Do not generate a second DOCX for either handoff.
Phase 3: final validation and handoff
Before delivery:
- Confirm every requirement maps to a deliverable or measurable standard.
- Confirm every deliverable has format, trigger or due date, and acceptance criteria.
- Confirm document type, contract framework, security, location, period, and transition are consistent.
- Confirm QASP consequences contain no CPARS rating commitments.
- Confirm the document contains no staffing, SOC, IGCE, CLIN, pricing, or skill-chain leakage.
- Run
scripts/validate_docx.py <document> --document-type <sow|pws>. - Render and inspect every page.
Deliver the .docx, then present both chat-only handoffs using handoff-specification.md. The pricing skill consumes the approved staffing table without repeating decomposition. The final message must state that the SOW/PWS is the contract-file artifact and the handoffs are internal workpapers outside it.
When a dynamic TOC has not been updated in Word, end with: Open the document in Word, press Ctrl+A (Cmd+A on Mac), then F9 to populate or refresh the Table of Contents, and save.
Out of scope
- IGCE rates, wraps, fee, or cost calculations.
- Acquisition plan, market research report, source-selection plan, J&A, or clause matrix.
- Final contract-type, commerciality, D&F, or inherently-governmental determination.
- Full standalone QASP beyond the embedded summary.
- Section I or L clause selection.
MIT © James Jenrette / 1102tools. Source: github.com/1102tools-dev/federal-contracting-skills
Files (federal-contracting-skills)
-
agents
-
openai.yaml 295 B
interface: display_name: "SOW/PWS Builder" short_description: "Build contract-ready federal work statements" default_prompt: "Use $sow-pws-builder to turn my scope decisions or SOO into a contract-ready SOW or PWS and separate pricing handoffs." policy: allow_implicit_invocation: true
-
-
references
-
document-specification.md 6.1 KB
# SOW/PWS Document Specification Use a supplied agency template as the design and section authority. Without one, use US Letter portrait, 1-inch margins, a restrained formal business style, page numbers, and real Word styles. ## First-page acquisition brief Before the detailed body, make page one useful to a program manager and Contracting Officer. Use a compact, document-native brief that states the acquisition outcome, the performance results that control acceptance, the most consequential Government inputs or unresolved assumptions, the transition posture, and the next review action. This is not a generic executive summary and must not contain staffing, FTE, SOC, CLIN, price, or IGCE content. A reader should be able to understand what the requirement buys and what still blocks release without reconstructing the document from later sections. ## Front matter - Plain title paragraph, document type, requirement title, agency or office, solicitation/contract placeholder, version/date, and distribution marking when supplied. - Dynamic Table of Contents after the title block for documents with more than eight main sections. - Do not include a staffing, pricing, IGCE, or Section B summary in front matter. ## Section order Preserve this order. Renumber when an agency template requires it, but do not merge unrelated sections. ### 1. Introduction - Purpose - Background - Scope summary - Applicable documents and standards - Confirmed contract framework only when it changes performance obligations; for CPFF, state the confirmed Completion or Term form ### 2. Definitions and Acronyms Define terms that materially affect performance. Omit common terms that do not need contractual definition. ### 3. Requirements For an SOW, organize by task area: - Task number and title - Contractor action - Subtasks or sequence only when Government-directed method is necessary - Deliverables - Government-furnished resources For a PWS, organize by performance objective: - Objective number and title - Required outcome - Performance standard - AQL - Assessment method - Contract-administration consequence or incentive when applicable Do not state FTE, staffing count, SOC, annual labor hours, proposed rate, CLIN, or price. ### 4. Deliverables Use a table only because deliverables are repeated comparable records: `ID | Title | Description | Format | Frequency | Due Date or Trigger | Acceptance Criteria | Approving Role` Include transition, status, system documentation, training, or final artifacts only when the scope requires them. Do not add generic deliverables that have no mapped requirement. ### 5. Period of Performance State base, options, phase boundaries, mobilization, and extension placeholders supplied by the user. Do not include Section B CLIN numbering. ### 6. Place of Performance State sites, remote-work boundaries, travel destinations, access hours, and Government facility constraints. Do not infer duty locations. ### 7. Government-Furnished Property and Information Identify each item, availability, condition, access date, maintenance responsibility, and return/disposition requirement. State `None identified` only when the user confirms it. ### 8. Security and Privacy Tailor to the confirmed information and facility context. Address safeguarding, CUI/privacy, suitability, clearances, facility requirements, incident reporting, access termination, and agency-specific authorities. Do not invent classification or clearance. ### 9. Key Personnel Name roles, minimum qualifications, certifications, substitution notice, and approval process. Do not include number of people, FTEs, SOCs, hours, or the full labor-category list. ### 10. Reporting and Oversight State report content, cadence, meetings, data sources, Government roles, escalation, and record retention. Avoid duplicating the Deliverables table. ### 11. QASP Summary or Inspection and Acceptance - PWS: title `QASP Summary`; use metric, standard, AQL, method, frequency, and payment/administrative consequence. - SOW: title `Inspection and Acceptance`; use deliverable, criterion, method, review period, approving role, and remedy. Do not map thresholds to CPARS ratings. ### 12. Transition - Transition-in: incumbent cooperation, knowledge transfer, access, inventory, parallel operation, readiness, and acceptance. - Transition-out: data and property return, documentation, successor support, access termination, and schedule. ### 13. Constraints and Assumptions Use: `ID | Assumption or Constraint | Basis | CO or Program Office Action` Prefix derived defaults with `[DEFAULT]`. Record pending contract type, commerciality, CPFF form, security, Government data, volume, or other unresolved decisions. Do not hide assumptions in narrative. ## Appendices Use only when applicable: - Current environment - Historical workload and volume - System interfaces - Acronyms Never create an appendix for staffing, FTE allocation, IGCE handoff, CLIN structure, pricing, or estimated labor hours. ## Language - SOW: `The Contractor shall <specific action>.` - PWS: `The Contractor shall achieve, maintain, or ensure <measurable result>.` - Avoid standalone `support`, undefined `as needed`, unnamed `best practices`, and coordination with no required outcome. - Every requirement must be testable through inspection, demonstration, analysis, test, or an identified performance record. - Use `Contracting Officer`, `Contracting Officer's Representative`, and `Program Office` consistently. ## DOCX construction - Real Heading 1, 2, and 3 styles. - Real numbered and bulleted lists; no typed bullet characters. - Explicit DXA table widths, cell margins, repeating header rows, and no fixed row heights. - Keep headings with following content and repeat table headers across pages. - Use page-number fields in the footer. - Set update-fields-on-open for the TOC. - Do not place internal instructions, test notes, prompt text, skill names, or file-system paths in the document. - Reject a technically valid document when the rendered first page is metadata-heavy, a table looks like an unfilled template, substantive sections merely repeat the brief, or a customer still has to invent acceptance logic, owners, or next actions. -
handoff-specification.md 2.9 KB
# Chat-Only Handoff Specification Emit both handoffs after the `.docx` is saved and validated. They are internal Government workpapers in the conversation, never files and never document sections or appendices. ## 1. Staffing handoff Use this heading and notice: ```text === STAFFING HANDOFF TABLE: FOR IGCE BUILDER === Internal Government workpaper. Not part of the SOW/PWS contract deliverable. Do not paste this table into the contract file. ``` Fixed columns: `Labor Category | SOC Code | FTE | Phase | Hours/Yr | Notes` Rules: - Include every user-approved labor category. - Preserve user overrides and derivation basis in Notes. - Derive continuous coverage from annual coverage hours divided by productive hours. - Keep at least four decimals in coverage calculations; round presentation only with disclosure. - Preserve Tier 1 help desk or contact-center SOC 43-4051. - Identify hybrid CLIN or contract-type routing in Notes. - Do not include rates, burden, fee, price, or fair-and-reasonable language. End with: `This approved staffing handoff is ready for the FFP, LH/T&M, or CR pricing skill selected by contract type. The pricing skill should preserve it and ask only for missing pricing inputs.` ## 2. Section B handoff Use this heading and notice: ```text === SECTION B HANDOFF TABLE === Contract-administration workpaper. Not part of the SOW/PWS contract deliverable. Use it to draft Section B of the solicitation shell, not the SOW/PWS body. ``` Base columns: `Line | Description | Contract Type | Pricing Unit or Basis | Period | Notes` Build rows from the user's preferred structure: by period, function, deliverable, or hybrid. Add separate travel and ODC lines when in scope. For T&M/LH, add a second table when the user wants estimated hours for rate evaluation: `Labor Category | Contractor or Source | Base Estimated Hours | Option Hours | Total Estimated Hours | Fixed Hourly Rate Input` Also show: `Overall T&M/LH Ceiling Price Input: <user-confirmed amount or PENDING CO DECISION>` Do not label estimated hours as a FAR-mandated labor-category ceiling. FAR 16.601 requires fixed hourly rates by labor category and an overall ceiling price; the Contracting Officer owns the final Section B structure. For CPFF, carry the confirmed Completion or Term form in Notes. For hybrid requirements, make each line's type explicit. End with: `This is a suggested Section B starting point. The Contracting Officer owns the final line-item structure, estimated hours, rates, and ceiling.` ## Separation audit The final `.docx` must not contain: - Either handoff heading or notice - `IGCE Builder`, `build the IGCE`, or pricing-skill names - `SOC Code`, FTE values, staffing derivation, or internal workpaper language - A CLIN or Section B table - Proposed hourly rates, burden, wraps, fee, or prices The conversation may contain all of those because the handoffs are explicitly outside the document. -
professional-product-standard.md 5.5 KB
# Professional product standard This file is the canonical source for the copy packaged with each 1102tools skill. The packaged copies must remain identical to this file. ## Product judgment Produce a finished professional work product, not a record of the process used to create it. Write for the person who must understand, use, approve, or act on the result. Exercise editorial judgment. Include material that improves the reader's understanding or decision. Omit material merely because it was collected, available, or easy to generate. Every page, section, table, and visual must earn its place. Match the structure, length, voice, and visual treatment to the assignment. Do not reuse a universal report outline. A short decision card, formal contract-file document, analytical workbook, landscape, timeline, or longer consulting report may all be correct products for different requests. Lead with the useful output. Research mechanics, process narration, methodology, limitations, and compliance controls are secondary unless the reader's purpose makes one of them the product. ## Controlled freedom Route rules define the substantive outcome and genuine formal boundaries. They do not prescribe identical headings, page counts, layouts, or section order unless a law, supplied template, calculation model, or downstream interface requires it. Choose the clearest form for each idea: - prose for explanation and judgment; - tables for real comparison or repeated fields; - cards, profiles, matrices, timelines, charts, and callouts when they improve comprehension; - appendices only for material the intended reader may reasonably need. Do not put paragraph-length narrative in narrow table cells. Do not repeat the same information as a callout, table, and prose section unless each form serves a different reader need. Use a restrained, coherent design with clear hierarchy, readable typography, comfortable spacing, and accessible contrast. Treat examples as quality references, not templates to copy. ## Paid-value standard The primary artifact should contain the analysis, comparison, requirements, model, or operating guidance the customer is paying to receive. Audit material must remain subordinate. Before delivery, remove: - generic background the intended reader already knows; - duplicated findings or actions; - query logs, tool operations, sanitized parameters, and internal record mechanics; - generic owners, gates, scenarios, or warnings invented to fill a template; - boilerplate disclaimers repeated in multiple sections; - empty sections and tables that merely announce missing content. When evidence is insufficient for the promised product, say so plainly and provide the useful narrower result or acquisition plan for the missing evidence. Do not pad an evidence gap into a document that resembles a completed analysis. ## Reader-facing source citations Internal evidence identifiers such as `E001` may remain in a private research record for backward compatibility, but they never appear in a customer-facing artifact. Assign each distinct reader-visible source an identifier in order of first appearance: `S1`, `S2`, `S3`, with no leading zeros. Reuse the same identifier wherever that source supports another claim. Cite sources beside the supported claim using forms such as `[S1]`, `[S1, S4]`, or `[S1-S3]`. End a sourced report with a concise `Source Register`. Each entry uses the corresponding identifier and provides enough information to verify the source: publisher or organization, title or record identity, relevant date, and a clickable public URL or supplied-document locator. Deduplicate sources. Do not display internal source-class tokens or a query-by-query research log. Native legal and acquisition citations such as `FAR 10.001`, a docket number, PIID, UEI, or document section remain in their ordinary form. Add an `S` citation when the artifact also needs a link to the supporting source; do not replace the native citation with an opaque source number. For workbooks, source notes and benchmark rows may use `S` identifiers that resolve to a Sources or Raw Data register. Formula cells and internal validation IDs are not reader-facing source citations. ## Proportionate boundaries Accuracy, authority boundaries, unresolved decisions, and limitations remain mandatory when material. Present them once, in the least intrusive form that keeps the product honest. A concise note or callout is preferable to a recurring legalistic section when the reader needs the answer more than a compliance lecture. Do not weaken a formal SOW/PWS, OT project description, acquisition-policy status analysis, or auditable cost model for stylistic reasons. Formal and mathematical requirements remain hard constraints. Apply taste to hierarchy, selection, explanation, and delivery view, not to the removal of necessary substance. ## Final editorial review After technical validation and rendering, review the complete artifact as a demanding customer: - Is the useful result apparent immediately? - Did the author select and synthesize rather than dump everything collected? - Does each section materially advance the reader's work? - Are sources integrated credibly without dominating the product? - Are limitations accurate and proportionate? - Does the artifact feel composed for this assignment rather than populated from a universal template? - Is this work product an experienced professional could confidently sell? Revise until the answer to every applicable question is yes. Structural validation is a technical floor, not the release decision. -
question-blocks.md 3.5 KB
# Intake and Scope Question Blocks Skip answers already present. Use a structured question tool when available; otherwise use numbered choices and accept number, label, or free text. Every choice set needs a free-text escape. ## Framing 1. Document form: SOW, PWS, or help deciding. 2. User-confirmed contract framework: FFP, T&M, LH, CPFF, CPAF, CPIF, hybrid by CLIN, or pending CO decision. 3. For CPFF: Completion or Term form, or pending CO decision. 4. Commerciality: commercial, noncommercial, or pending CO determination. 5. Agency, solicitation template, acquisition stage, audience, and file name. Explain consequences but do not make FAR Part 16, commerciality, or D&F decisions. ## Block 1: mission and service model - Core service and mission outcome. - Government staff with contractor augmentation, fully contracted service, or hybrid. - Coverage window, days per week, holidays, surge, and on-call expectations. - Single site, multiple sites, remote, or hybrid. - Current-state problem and success condition. ## Block 2: technical scope - New build, COTS/SaaS configuration, O&M, transition, or hybrid. - Systems, applications, facilities, or assets in scope. - Interfaces and integration count. - Data migration sources, volume, retention, and quality. - Automation or AI functions and human-review boundaries. - Government standards, architecture, tools, or processes that are genuine constraints. ## Block 3: scale and volume - Annual and peak transaction, ticket, contact, case, or workload volume. - User population and concurrency. - Number of sites, systems, environments, or organizational units. - Growth, seasonality, surge, backlog, and response-time expectations. - Service levels and failure impacts. ## Block 4: organization and phasing - Offices and stakeholders served. - Pilot, phased rollout, or full deployment. - Dependencies, Government decisions, and external approvals. - Incumbent transition and knowledge-transfer facts. - Government-furnished property, information, facilities, and access. ## Block 5: period and Section B preferences - Base and option periods. - Transition-in and transition-out durations. - Desired Section B structure: by period, function, deliverable, or hybrid. - Travel and ODC lines. - For T&M/LH: labor categories, estimated hours for evaluation if desired, separate contractor/subcontractor rate treatment when applicable, and overall ceiling-price input. Keep all pricing schedule material out of the SOW/PWS. ## Block 6: quality, oversight, and security - Deliverables, cadence, format, due date, and acceptance authority. - Measurable performance standards and AQLs. - Inspection, demonstration, analysis, test, service-level reports, or sampling method. - Contract-administration consequence when a threshold is missed. - Key-personnel roles and qualifications, without headcount. - Information type, CUI, privacy, system authorization, suitability, clearance, facility, and incident reporting. - Reporting and meeting cadence. ## Staffing derivation questions Ask only what is needed to support the internal workpaper: - Productive hours per FTE. - Coverage seats and exact annual coverage hours. - Workload productivity basis, including contacts per agent-day or tasks per analyst-month. - System/module count and team design. - Time-limited transition or surge effort. - User override and reason. Preserve Tier 1 help desk/contact-center SOC 43-4051. Use 15-1232 only when the work is explicitly computer user support rather than Tier 1 intake, triage, routing, or customer-service contact handling. -
regulatory-and-content-rules.md 3.5 KB
# Regulatory and Content Rules Use current official text and the user's agency supplement. These anchors do not replace Contracting Officer review. ## Performance-based services - FAR 37.102: performance-based acquisition is preferred for services to the maximum extent practicable. - FAR 37.602(b)(1): describe a PWS by required results rather than method or number of hours to the maximum extent practicable. - FAR 37.602(b)(2): enable assessment against measurable performance standards. - FAR 46.401 and Subpart 46.4: Government contract quality-assurance responsibilities. Do not cite FAR 37.102(d) as an hours or staffing restriction. That paragraph addresses nonpersonal service contracts. ## T&M and Labor-Hour - FAR 16.601(c)(2): separate fixed hourly rates by labor category, including wages, overhead, G&A, and profit. - FAR 16.601(d)(1): Contracting Officer D&F that no other type is suitable. - FAR 16.601(d)(2): overall ceiling price that the contractor exceeds at its own risk. - FAR 52.232-7: payment mechanics for labor and materials. These rules affect Section B. They do not create a requirement for a labor-category ceiling-hours table inside the PWS. Keep rates, estimated hours for evaluation, and the ceiling-price schedule in the chat-only Section B handoff. ## CPFF Require user confirmation of Completion or Term form under FAR 16.306(d): - Completion describes a definite goal or target and specifies an end product. - Term obligates a specified level of effort for a stated period. - Completion is preferred when the work can be defined well enough to estimate completion. Do not default Term merely because the requirement is R&D. Record a pending decision when the user cannot confirm. ## Commerciality and contract type Do not determine commerciality from a service label. Do not select FFP, T&M, LH, or CR. Explain document consequences and mark the decision pending for Contracting Officer action. ## Security For unclassified work, tailor safeguarding, CUI, privacy, suitability, and agency requirements to actual information and system context. Do not insert NIST, CMMC, FISMA, privacy, or suitability language without a scope basis. For classified work, surface the need for the applicable DD Form 254, Security Classification Guide, facility clearance, personnel access, derivative-classification and OPSEC requirements, incident reporting, and agency-specific authorities. Use placeholders when the controlling document has not been supplied. Do not invent classification guidance. ## Key personnel Substitution is agency- and contract-specific. Never cite FAR 52.237-2; it concerns protection of Government buildings, equipment, and vegetation. Use supplied agency clauses or contract-specific language. FAR 52.237-3 may support continuity duties but is not a generic key-personnel substitution clause. ## QASP and CPARS QASP metrics support surveillance and contract administration. CPARS is a separate past-performance evaluation under FAR Subpart 42.15. Do not pre-commit CPARS rating labels from metric results. Use a consequence actually supported by the contract framework, such as withholding acceptance, re-performance, a cure notice, an agreed deduction, or an incentive specifically established in the solicitation. Do not invent a payment deduction or termination trigger. ## Scope reduction Describe prior and revised scope by capability, coverage, volume, frequency, or performance level. Keep FTE reductions in the chat-only workpaper. When savings come from an IGCE, label them as estimate deltas and preserve the assumptions that produced them. -
runtime-adaptation.md 1.7 KB
# Runtime Adaptation ## Questions Use the host's structured question tool when it exists. Otherwise present numbered options in chat and accept numbers, labels, or free text. Never expose a host-specific tool name to the user. Batch independent questions. Stop at staffing approval and Decision Summary gates. Do not use tool availability as a reason to self-approve. ## Document authoring Use the host's document capability, Python with python-docx, or direct OOXML. Resolve this skill's references and scripts relative to the skill directory. Do not assume fixed mount paths, file-presentation helpers, or a client-specific document skill exists. If no `.docx` authoring path exists, stop. Markdown and HTML are not equivalent deliverables. ## Rendering Use a real office rendering engine, preferably LibreOffice headless, to convert the latest `.docx` to page images. Inspect every page at 100% zoom. A PDF renderer is evidence about layout, not proof of Microsoft Word behavior. When Word is available, open the final document there for TOC, field, and target-application verification. Record the exact surface tested. ## Table of Contents Use real heading styles and a dynamic TOC field for documents with more than eight main sections. Set `w:updateFields` in document settings. If the TOC cannot be updated in Word before delivery, keep the field and give the refresh instruction in chat. Do not place the instruction inside the contract document. ## Delivery Use the host's artifact-presentation capability when available. Otherwise save to the requested or current directory and report the absolute path. Deliver only the SOW/PWS `.docx`; the handoffs remain chat-only. -
validation-gates.md 3 KB
# SOW/PWS DOCX Validation Gates Run every gate against the final file, then rerun after any change. ## Structural audit Run: ```text python scripts/validate_docx.py output.docx --document-type pws ``` The validator checks: - The file is a valid DOCX ZIP. - Core sections exist and appear in order. - Heading paragraphs use real Word heading styles. - A document with more than eight main sections contains a dynamic TOC field and update-fields-on-open setting. - No staffing, SOC, IGCE, CLIN, pricing, or skill-chain leakage appears in document text. - PWS documents contain measurable performance language and a QASP Summary. - SOW documents contain Inspection and Acceptance. - Forbidden FAR 37.102(d) hours claims and FAR 52.237-2 key-personnel citations are absent. - CPARS rating labels are not tied to QASP content. ## Text and semantic audit Extract all document text and verify: - Every requirement maps to a deliverable or performance standard. - Every deliverable has a trigger or due date and acceptance criteria. - Place, period, security, Government-furnished resources, and transition are consistent. - `[DEFAULT]` assumptions identify required owner action. - No prompt text, local path, internal citation token, test instruction, or TOC-refresh instruction leaked into the file. - No vague standalone `support`, undefined `as needed`, or untestable `best practices` requirement remains. - Page one states the acquisition outcome, controlling performance results, consequential unresolved Government inputs, transition posture, and next review action without leaking staffing or pricing. - Each material unknown names a Government owner and a closeout action; the body contains executable requirements and acceptance evidence rather than repeating the first-page brief. ## Render audit Render the latest DOCX to page PNGs with a real office engine and inspect every page at 100% zoom. Check: - No clipping, overlap, broken glyphs, or orphaned headings. - Tables fit the page, repeat headers, wrap correctly, and use readable widths. - Header, footer, and page numbers are aligned. - TOC placement is correct. If the dynamic field is not populated, retain it and disclose the Word refresh step in chat. - No large blank gaps or accidental blank pages. - The first page reads as an acquisition decision product, not a cover sheet or metadata dump. ## Target application LibreOffice rendering is not proof of Microsoft Word behavior. When Word is available, verify the final file in Word, refresh fields, and save. Record the exact surface in `test.md`. ## Separation fault injection At minimum, confirm the validator rejects test copies containing: 1. `STAFFING HANDOFF TABLE`. 2. `SOC Code` or an FTE staffing table. 3. `CLIN HANDOFF TABLE` or a Section B pricing table. 4. `FAR 37.102(d)` as the basis for results-versus-hours language. 5. `FAR 52.237-2` as a key-personnel clause. 6. A PWS missing QASP Summary or measurable standards. Document the exact automated and manual tests in `test.md`.
-
-
scripts
-
validate_docx.py 6.9 KB
#!/usr/bin/env python3 """Validate SOW/PWS DOCX structure, headings, TOC, and artifact separation.""" from __future__ import annotations import argparse import json import re import sys import zipfile from pathlib import Path from docx import Document CORE_SECTIONS = [ "introduction", "definitions and acronyms", "requirements", "deliverables", "period of performance", "place of performance", "government-furnished property and information", "security and privacy", "key personnel", "reporting and oversight", "quality", "transition", "constraints and assumptions", ] FORBIDDEN = { "staffing handoff": re.compile(r"staffing\s+handoff\s+table", re.I), "IGCE skill plumbing": re.compile(r"\bIGCE\s+Builder\b|\bbuild\s+the\s+IGCE\b", re.I), "Section B handoff": re.compile(r"(?:CLIN|Section\s+B)\s+handoff\s+table", re.I), "SOC code": re.compile(r"\bSOC\s+Code\b", re.I), "FTE staffing": re.compile(r"\bFTEs?\b", re.I), "CLIN content": re.compile(r"\bCLINs?\b", re.I), "labor-category ceiling-hours table": re.compile(r"Labor\s+Category\s+Ceiling\s+Hours", re.I), "wrong FAR 37 citation": re.compile(r"FAR\s+37\.102\s*\(d\)", re.I), "wrong key-personnel clause": re.compile(r"FAR\s+52\.237-2", re.I), "TOC refresh instruction": re.compile(r"(?:Ctrl|Cmd)\+A.{0,40}F9", re.I | re.S), "local runtime path": re.compile(r"/(?:mnt|tmp|Users)/|[A-Za-z]:\\", re.I), } def normalize_heading(text: str) -> str: text = re.sub(r"^\s*(?:section\s+)?\d+(?:\.\d+)*[.):-]?\s*", "", text, flags=re.I) return re.sub(r"\s+", " ", text).strip().lower() def document_text(document: Document) -> str: parts = [paragraph.text for paragraph in document.paragraphs] for table in document.tables: for row in table.rows: for cell in row.cells: parts.append(cell.text) return "\n".join(parts) def heading_level(paragraph: object) -> int | None: style = getattr(paragraph, "style", None) name = getattr(style, "name", "") or "" match = re.fullmatch(r"Heading\s+([1-9])", name, re.I) return int(match.group(1)) if match else None def has_toc_and_update(path: Path) -> tuple[bool, bool]: with zipfile.ZipFile(path) as archive: document_xml = archive.read("word/document.xml").decode("utf-8", errors="replace") settings_xml = archive.read("word/settings.xml").decode("utf-8", errors="replace") has_toc = bool(re.search(r"<w:instrText[^>]*>[^<]*\bTOC\b", document_xml, re.I)) update = bool(re.search(r"<w:updateFields\b[^>]*(?:w:val=[\"'](?:true|1)[\"'])?", settings_xml, re.I)) return has_toc, update def validate(path: Path, document_type: str) -> dict[str, object]: failures: list[str] = [] try: if not zipfile.is_zipfile(path): return {"status": "fail", "failures": ["file is not a valid DOCX ZIP"]} with zipfile.ZipFile(path) as archive: names = set(archive.namelist()) for required in ("[Content_Types].xml", "word/document.xml", "word/styles.xml"): if required not in names: failures.append(f"missing DOCX part: {required}") document = Document(path) except (OSError, ValueError, zipfile.BadZipFile) as exc: return {"status": "fail", "failures": [f"cannot read DOCX: {exc}"]} text = document_text(document) for label, pattern in FORBIDDEN.items(): if pattern.search(text): failures.append(f"document contains forbidden {label}") h1 = [ normalize_heading(paragraph.text) for paragraph in document.paragraphs if heading_level(paragraph) == 1 and paragraph.text.strip() ] if len(h1) < len(CORE_SECTIONS): failures.append(f"only {len(h1)} Heading 1 sections found; expected at least {len(CORE_SECTIONS)}") positions: list[int] = [] for expected in CORE_SECTIONS: if expected == "quality": candidates = [ index for index, value in enumerate(h1) if value in {"qasp summary", "inspection and acceptance"} ] else: candidates = [index for index, value in enumerate(h1) if value == expected] if not candidates: failures.append(f"missing Heading 1 section: {expected}") else: positions.append(candidates[0]) if positions and positions != sorted(positions): failures.append("core Heading 1 sections are out of order") if document_type == "pws": if "qasp summary" not in h1: failures.append("PWS is missing a Heading 1 QASP Summary") if not re.search(r"\bAQL\b|Acceptable Quality Level", text, re.I): failures.append("PWS does not contain an AQL") if not re.search(r"performance standard|assessment method|method of assessment", text, re.I): failures.append("PWS does not contain measurable performance or assessment language") else: if "inspection and acceptance" not in h1: failures.append("SOW is missing a Heading 1 Inspection and Acceptance section") if re.search(r"\bCPARS\b", text, re.I) and re.search( r"\b(?:Exceptional|Very Good|Satisfactory|Marginal|Unsatisfactory)\b", text, re.I ): failures.append("document ties CPARS language to rating labels") if len(h1) > 8: try: has_toc, update = has_toc_and_update(path) except (KeyError, OSError, zipfile.BadZipFile) as exc: failures.append(f"cannot audit TOC settings: {exc}") else: if not has_toc: failures.append("document with more than eight sections has no dynamic TOC field") if not update: failures.append("document does not set updateFields-on-open") return { "status": "pass" if not failures else "fail", "document_type": document_type, "heading_1_count": len(h1), "table_count": len(document.tables), "failures": failures, } def main() -> int: parser = argparse.ArgumentParser(description="Validate a SOW or PWS DOCX.") parser.add_argument("document", type=Path) parser.add_argument("--document-type", required=True, choices=("sow", "pws")) parser.add_argument("--json", action="store_true") args = parser.parse_args() if not args.document.is_file(): print(f"ERROR: document not found: {args.document}", file=sys.stderr) return 2 result = validate(args.document, args.document_type) if args.json: print(json.dumps(result, indent=2, sort_keys=True)) elif result["status"] == "pass": print("DOCX structure, content separation, and TOC settings passed.") else: print("VALIDATION FAILED") for failure in result["failures"]: print(f"- {failure}") return 0 if result["status"] == "pass" else 1 if __name__ == "__main__": raise SystemExit(main())
-
-
SKILL.md 14.5 KB
--- name: sow-pws-builder description: > Trigger for: writing, revising, converting, or descoping a federal Statement of Work, Performance Work Statement, SOW, PWS, or SOO; develop executable requirements; define contract scope; create measurable performance standards; or prepare requirements before an IGCE. Produce a contract-file-ready .docx plus separate chat-only staffing and Section B handoffs. Never place FTEs, SOC codes, staffing estimates, IGCE content, CLINs, or pricing schedules inside the SOW/PWS body. Do NOT use for the IGCE, market research report, acquisition plan, QASP, or clause-selection matrix. --- # SOW/PWS Builder ## Overview Turn program-office scope decisions into a contract-file-ready `.docx` SOW or PWS. Produce three separate outputs: 1. The SOW/PWS `.docx`, containing requirements and measurable acceptance or surveillance content. 2. A chat-only staffing handoff for the FFP, LH/T&M, or CR pricing skill. 3. A chat-only Section B handoff with suggested CLIN structure, labor categories or estimated hours when needed, travel/ODC lines, and the total ceiling-price input for T&M/LH. Never combine the handoffs with the document. FAR 37.602(b)(1) directs agencies, to the maximum extent practicable, to describe performance work by required results rather than method or hours. FAR 16.601 requires fixed hourly rates by labor category for T&M/LH and an overall ceiling price, but it does not require a labor-category ceiling-hours table inside the PWS. Keep pricing structure in Section B unless the user supplies a controlling solicitation template that places it elsewhere. No external MCP server is required. ## Product quality default Make the PWS/SOW read like a concise acquisition-file deliverable, not a process transcript. Its first page must identify the mission need, the performance decision it supports, the material operating assumptions, and the next acquisition-team action. Use a requirement-specific title, a short executive purpose block, and distinct visual hierarchy appropriate to the document type. Put detailed evidence and source notes in a compact appendix only when they affect a requirement; never let them crowd out executable requirements. This does not permit staffing, SOC, CLIN, pricing, or IGCE content inside the document. Load supporting files only when needed: - [question-blocks.md](references/question-blocks.md) for intake and scope questions. - [document-specification.md](references/document-specification.md) before authoring the `.docx`. - [professional-product-standard.md](references/professional-product-standard.md) before authoring the `.docx`. - [regulatory-and-content-rules.md](references/regulatory-and-content-rules.md) for contract-type, security, QASP, and language rules. - [handoff-specification.md](references/handoff-specification.md) before final chat output. - [runtime-adaptation.md](references/runtime-adaptation.md) for questions, document tools, TOC handling, and delivery. - [validation-gates.md](references/validation-gates.md) before delivery. ## Permanent correctness gates 1. **Separation:** The `.docx` contains no staffing handoff, FTE estimate, SOC code, IGCE content, CLIN table, pricing schedule, or skill-chain message. The two handoffs exist only in chat. 2. **No false FAR exception:** Do not cite FAR 37.102(d) as an hours prohibition. Use FAR 37.602(b)(1) for results-oriented PWS language. 3. **T&M/LH structure:** FAR 16.601(c)(2) requires fixed hourly rates by labor category, and FAR 16.601(d)(2) requires an overall ceiling price. Put the pricing schedule and ceiling in the Section B handoff, not the SOW/PWS body. 4. **Contract-type boundary:** Explain how a user-selected type changes the document, but do not originate the FAR Part 16 decision, T&M/LH Determination and Findings, commercial-item determination, or CPFF form selection. 5. **CPFF form:** Require the user to confirm Completion or Term under FAR 16.306(d). Do not default either form. Reflect the confirmed form in the document framework. 6. **Result-oriented PWS:** Organize a PWS around outcomes, measurable standards, AQLs, and assessment methods. Avoid prescribing staffing or contractor method unless a constraint is genuinely Government-controlled. 7. **QASP and CPARS:** QASP payment or administrative consequences must not pre-commit CPARS ratings. Never map an AQL threshold to `Satisfactory`, `Very Good`, `Exceptional`, `Marginal`, or `Unsatisfactory` CPARS ratings. 8. **Key personnel:** Name roles and qualifications only. Do not quantify staff. Never cite FAR 52.237-2 as a key-personnel substitution clause. 9. **Coverage derivation:** For the chat-only staffing handoff, derive coverage as annual coverage hours divided by productive hours. At 1,880 hours, one 24x7x365 seat is 4.6596 FTE, not three and not 4.2. 10. **Tier 1 mapping:** Preserve the project decision that Tier 1 help desk or contact-center agents map to SOC 43-4051. Do not change it during modernization. 11. **Decision gates:** Do not self-approve staffing or the Phase 2 Decision Summary. End each gate response at its confirmation question and wait. 12. **DOCX validation:** Use real heading styles, render every page to images, inspect every page, audit text and OOXML, and do not claim Word compatibility from a fallback renderer alone. ## Workflow selection ### Workflow A: full build Use for a concept, rough requirements, or a build from scratch. Run intake and all phases. ### Workflow B: SOO conversion Extract settled objectives, constraints, location, period, systems, volumes, security, and other facts from the SOO. Present the gaps that must be decided, then run the remaining phases without re-asking settled facts. ### Workflow C: scope reduction Use an existing SOW/PWS and, when available, IGCE cost drivers to present capability and coverage tradeoffs. The user selects reductions. Revise affected requirements and handoffs without placing staff counts in the document. ## Runtime pre-flight When this skill is entered immediately after a numbered Pre-Award Agent selection and the current assistant response has not already shown the orchestrator's outcome preview, emit these exact four lines before any intake or capability check: Begin line 1 with `Recommended outcome:`. Do not precede the block with a heading, acknowledgement, selection recap, routing narration, or code fence. ```text Recommended outcome: Validated SOW/PWS `.docx` plus two chat-only handoffs Includes: an executable work statement, measurable standards, a staffing handoff, and a Section B handoff Boundary/default: recommend PWS for performance-based services when the requirement supports it; the user or Contracting Officer retains contract type, commerciality, and other reserved decisions Next: collect the current requirement or source material and missing acquisition-strategy facts ``` This is a routing fallback, not a second preview. Do not repeat it when the orchestrator already rendered the four lines in the current assistant response, and never replace it with component intake. After the orchestrator's outcome preview, begin acquisition-strategy intake and reuse every supplied fact. Do not make document-authoring or rendering capability the first question or first action after workflow selection. A read-only or artifact-limited session may still inspect supplied material, identify gaps, and complete useful scope intake. Before promising or beginning the validated `.docx` build: 1. Confirm the host can read inputs and create `.docx` files. 2. Confirm a DOCX render path is available, preferably LibreOffice through the host's document workflow. 3. Confirm Python and `python-docx` or an equivalent OOXML authoring capability for `scripts/validate_docx.py`. 4. Match capabilities semantically. Do not depend on `/mnt` paths, a named client tool, or a generated namespace. 5. If document authoring is unavailable, stop at the artifact boundary and report the missing capability. Preserve the intake already completed and explain what remains needed to resume. Do not silently substitute Markdown or HTML for the contract-file deliverable. ## Acquisition strategy intake Collect framing decisions in one pass, skipping facts already supplied: 1. SOW or PWS. 2. User-confirmed contract type: FFP, T&M, LH, CR subtype, or hybrid by CLIN. 3. Commercial, noncommercial, or pending Contracting Officer determination. 4. Agency and solicitation format, if one controls. 5. Purpose, audience, acquisition stage, and desired file name. If the user is unsure about SOW versus PWS, explain that PWS is outcome-oriented and preferred for performance-based services to the maximum extent practicable. A recommendation about document form is permitted. Do not turn it into the contract-type or commerciality decision. If T&M or LH is selected, flag the D&F and overall ceiling-price requirement for Contracting Officer action. If CPFF is selected, require Completion or Term. If commerciality is unsettled, record `[DEFAULT: Commerciality determination pending]` in Constraints and Assumptions rather than deciding it. ## Phase 0: SOO or source intake For a supplied source document: 1. Extract background, purpose, objectives, performance constraints, location, period, systems, volumes, security, Government-furnished resources, and named standards. 2. Distinguish explicit facts from derived assumptions. 3. Present gaps as the questions needed for executable requirements. 4. Carry settled facts forward without asking them again. If the source is too thin to support task or objective decomposition, state that it can serve as background but additional scope decisions are required. ## Phase 1: scope decision tree Use [question-blocks.md](references/question-blocks.md). Prefer structured choices when the host supports them; otherwise use numbered choices with a free-text escape. Batch three to four related questions. Ask open text only for values such as system names, volume, or incumbent facts. Sequence: 1. Mission and service model. 2. Technical scope. 3. Scale and volume. 4. Organizational scope and phasing. 5. Period, transition, and Section B preferences. 6. Quality, oversight, security, and Government-furnished resources. ### Staffing derivation for the internal handoff Derive a candidate staffing table only after scope decisions. Every row must cite a basis: workload, coverage hours, system/module count, site count, delivery cadence, or an explicit user override. Heuristics are starting points, not facts. - Continuous coverage: covered seats times annual coverage hours divided by productive hours per FTE. - Development: derive by system/module count and team design; do not use an unexplained fixed ratio. - Contact center: annual contacts divided by workdays and contacts per agent-day, adjusted for service level and productive hours. - Transition and surge: separate time-limited rows. - O&M: derive from named workload, SLA, or environment size; never assume a fixed percentage of development staffing without user confirmation. Present the candidate staffing table in chat and ask the user to confirm, amend, or override it. End immediately after that question and wait. Preserve override reasons. ## Phase 2: Decision Summary gate After staffing approval and before authoring, present a concise Decision Summary containing: - SOW/PWS, contract type and subtype, CPFF form when relevant, and commerciality status. - Agency/template and solicitation-stage context. - All defaults or derived assumptions with one-line rationale. - Section 3 task areas for a SOW or performance objectives for a PWS. - Applicable deliverables, QASP or inspection approach, security, period, location, and transition. - Confirmation that staffing, SOCs, CLINs, and pricing will remain outside the document. The final sentence must ask the user to proceed or correct the summary. Stop immediately after the question. Do not author the `.docx` in the same response. ## Phase 2: document assembly After explicit approval: 1. Read [professional-product-standard.md](references/professional-product-standard.md), [document-specification.md](references/document-specification.md), and [regulatory-and-content-rules.md](references/regulatory-and-content-rules.md). Formal requirement controls remain binding, but the document should still read as deliberate professional work rather than a generated template. 2. Use the host's document-authoring workflow and a formal business-document design system. Preserve a supplied agency template when one exists. 3. Use real `Heading 1`, `Heading 2`, and `Heading 3` styles, real numbered or bulleted lists, explicit table geometry, page numbers, and accessible repeating table headers. 4. For more than eight main sections, include a dynamic TOC field. Set the document to update fields on open. If Word is available, update and save the field before delivery. Otherwise disclose the exact refresh step in chat. Do not put the refresh instruction inside the document. 5. Run the render, text, structure, and separation gates in [validation-gates.md](references/validation-gates.md). 6. Fix defects and repeat rendering until every page passes. Do not generate a second DOCX for either handoff. ## Phase 3: final validation and handoff Before delivery: 1. Confirm every requirement maps to a deliverable or measurable standard. 2. Confirm every deliverable has format, trigger or due date, and acceptance criteria. 3. Confirm document type, contract framework, security, location, period, and transition are consistent. 4. Confirm QASP consequences contain no CPARS rating commitments. 5. Confirm the document contains no staffing, SOC, IGCE, CLIN, pricing, or skill-chain leakage. 6. Run `scripts/validate_docx.py <document> --document-type <sow|pws>`. 7. Render and inspect every page. Deliver the `.docx`, then present both chat-only handoffs using [handoff-specification.md](references/handoff-specification.md). The pricing skill consumes the approved staffing table without repeating decomposition. The final message must state that the SOW/PWS is the contract-file artifact and the handoffs are internal workpapers outside it. When a dynamic TOC has not been updated in Word, end with: `Open the document in Word, press Ctrl+A (Cmd+A on Mac), then F9 to populate or refresh the Table of Contents, and save.` ## Out of scope - IGCE rates, wraps, fee, or cost calculations. - Acquisition plan, market research report, source-selection plan, J&A, or clause matrix. - Final contract-type, commerciality, D&F, or inherently-governmental determination. - Full standalone QASP beyond the embedded summary. - Section I or L clause selection. --- *MIT © James Jenrette / 1102tools. Source: github.com/1102tools-dev/federal-contracting-skills* -
test.md 3.3 KB
# SOW/PWS Builder Modernization Test Record Tested August 21, 2026 against the modernized skill. The historical April test record remains in `testing.md`. ## Automated validation - `quick_validate.py` passed the skill directory. - `scripts/validate_docx.py --help` and Python compilation passed. - A representative eight-page Enterprise Service Desk PWS fixture passed the DOCX ZIP, heading-order, QASP, separation, dynamic-TOC, and update-fields-on-open checks. - The fixture used 13 ordered Heading 1 sections, six structured tables, repeating table headers, explicit table geometry, a page-number field, and a dynamic TOC field. - Six fault-injected copies were rejected for the expected reasons: staffing handoff leakage, SOC/FTE leakage, CLIN handoff leakage, the incorrect FAR 37.102(d) citation, the incorrect FAR 52.237-2 key-personnel citation, and a missing QASP Summary. ## Render and visual review - Rendered with LibreOffice through the Codex document renderer to eight page PNGs and a PDF. - Inspected every rendered page. The first pass exposed table rows splitting across pages. The fixture was rebuilt with non-splitting rows and repeating headers, then rerendered and re-inspected. - The final render had no clipped text, overlapping content, broken glyphs, accidental blank pages, or orphaned Heading 1 sections. - The TOC remained a valid dynamic field with a placeholder in the LibreOffice render. The skill correctly requires a Word field refresh disclosure when Word has not populated it. - Microsoft Word was not retested in this pass. LibreOffice rendering is not recorded as proof of Word behavior. ## Claude behavior tests Surface: Claude Code CLI 2.1.238, `claude-opus-5`, explicit `/sow-pws-builder` invocation. 1. A rich FFP help-desk prompt with one 24x7x365 Tier 1 seat and 1,880 productive hours produced 4.6596 FTE for the coverage floor, preserved SOC 43-4051, kept staffing in chat, did not create a file, and ended at the staffing confirmation question. 2. A CPFF research-support prompt that asked the model to choose the form was stopped at the FAR 16.306(d) gate. Claude explained Completion and Term, refused to originate the choice, and ended by asking the user to confirm one form. ## Codex behavior test Surface: Codex CLI 0.149.0-alpha.4, GPT-5.6 Sol, extra-high reasoning, explicit `$sow-pws-builder` invocation. - The CPFF self-selection prompt was stopped at the required gate. Codex presented Completion and Term without choosing either and ended at the confirmation question. ## Confirmed gates - Staffing, SOC, IGCE, CLIN, and pricing content stays outside the contract-file artifact. - FAR 37.602(b)(1), not FAR 37.102(d), is used for results-oriented PWS language. - T&M/LH labor-category rates, estimated hours, and overall ceiling inputs remain in the chat-only Section B handoff. - CPFF Completion versus Term is user-confirmed. - A 24x7x365 seat at 1,880 productive hours is 4.6596 FTE. - Tier 1 help-desk intake remains SOC 43-4051. - Both the staffing gate and Decision Summary gate stop and wait for approval. ## Open coverage - Claude web, Claude Code after automatic compaction, Codex Desktop UI, and Microsoft Word were not rerun for this skill in this pass. - Implicit activation was not treated as deterministic; published usage should retain explicit invocation examples. -
testing.md 24.3 KB
# SOW/PWS Builder: Testing Record # Part 1: For Federal Acquisition Users ## The bottom line Two waves of independent testing in April 2026 (14 end-to-end runs, 168 binary assertions graded) show the SOW/PWS Builder reliably produces FAR-compliant Statements of Work and Performance Work Statements across six contract-type and workflow combinations on both Claude Opus 4.7 and Claude Sonnet 4.6. Five assertions failed across the two waves. Four of the five point to specific places where the skill's output needs a manual pass before solicitation release. None produced a document that was unsalvageable. ## Contract types tested and how reliably they work | Contract type and workflow | Models | Result | |---|---|---| | T&M Statement of Work, non-commercial | Opus, Sonnet | Reliable | | FFP PWS (commercial, full build) | Opus, Sonnet | Reliable | | FFP PWS (commercial, SOO conversion) | Opus, Sonnet | Reliable | | FFP PWS scope reduction (Workflow C) | Opus, Sonnet | Reliable after Wave 1 patch | | CPFF R&D PWS, non-commercial | Opus, Sonnet | Reliable; verify CPFF form (see checklist) | | Labor-Hour PWS, no materials, non-commercial | Opus, Sonnet | Reliable; verify DD Form 254 (see checklist) | ## Manual-verification checklist The Wave 2 testing surfaced four specific gaps in the skill's output. **All four were patched after Wave 2.** The patches have not yet been regression-tested (Wave 3 will do that), so until the next wave confirms the fixes hold, scan every output for these items. **1. Key Personnel substitution clause.** Historical gap: the skill emitted "FAR 52.237-2" as the Key Personnel clause on 5 of 6 Wave 2 runs. FAR 52.237-2 is actually "Protection of Government Buildings, Equipment, and Vegetation" — the wrong clause. The Wave 2 patch removed that default and added agency-aware guidance: NFS 1852.237-72 for NASA, HSAR 3052.237-72 for DHS, HHSAR 352.237-75 for HHS, DFARS where applicable for DoD, and a generic fallback that avoids FAR 52.237-2 entirely. Verify the citation still looks right for your agency. **2. CPFF form selection.** Historical gap: on both Wave 2 CPFF runs, the skill cited FAR 16.306 without selecting completion form or term form. The Wave 2 patch added a Phase 2 rule requiring the skill to explicitly name completion form (FAR 16.306(d)(1)) or term form (FAR 16.306(d)(2)) wherever the contract framework first appears. Verify your CPFF PWS says which form it is. The difference matters: completion form requires delivery of a specified end product before the full fixed fee is earned; term form obligates a stated level of effort over a stated period. Term form is usually the right default for R&D where technical outcomes are uncertain. **3. Classified security block.** Historical gap: on both Wave 2 Labor-Hour TS/SCI runs, DD Form 254 and Security Classification Guide references were missing. The Wave 2 patch added required DD 254 and SCG references in Section 9 whenever any clearance at Confidential or higher is called out. Verify the security block for any classified requirement. **4. Section ordering.** Historical gap: workers occasionally swapped Section 11 (Reporting) and Section 12 (QASP) across runs. The Wave 2 patch added prescriptive language fixing Sections 11 through 14 in their correct order and explicitly disallowing merging or reordering. Low severity. A visual scan is enough. ## Choosing between Opus 4.7 and Sonnet 4.6 Short answer: use Opus 4.7 as the default. **Opus 4.7** asks multi-choice clarification questions when the input has real ambiguity, waits for your explicit "proceed" before generating the document, and produces slightly more thorough clause citations. Two caveats: - On claude.ai web chat, the Opus safety classifier sometimes blocks legitimate biodefense, vaccine, or pathogen-adjacent scenarios. If you hit a block ("Chat paused"), switch to Sonnet 4.6 or use the Claude API directly. This happened once in testing on an NIH mRNA vaccine R&D prompt. - Opus consistently cites the wrong Key Personnel clause (see checklist item 1). **Sonnet 4.6** is faster and works well on richly specified prompts. Two caveats: - Sonnet tends to skip the clarification-question phase even when defaults should have been reviewed. On the DHS Labor-Hour test, Sonnet self-approved six applied defaults without waiting for user confirmation. Read the Decision Summary carefully and interrupt if you see something that needs adjusting. - On very large documents (roughly 30 KB and up), Sonnet on claude.ai web chat sometimes hits an output truncation limit mid-generation. A fresh chat retry usually completes. For a critical document, run Opus first. If Opus is blocked or unavailable, switch to Sonnet. ## What the skill does not do - **It does not produce Independent Government Cost Estimates.** It emits a staffing handoff table in chat only, for the IGCE Builder skill to consume as a separate step. - **It does not produce CLIN structures inside the PWS body.** CLINs go in Section B of the solicitation, and the skill correctly emits them as a chat-only handoff. - **It does not handle classified data.** It produces unclassified acquisition documents that describe requirements for classified work. No actual classified information has been processed in testing. - **It has not been tested on** Special Access Program or Special Access Required variants; commercial CPFF or commercial Labor-Hour (which require FAR 12.207(b) justifications); hybrid contract structures such as FFP plus T&M CLINs; multi-year ramping AQL structures where targets escalate across option years; or true document-to-document scope reduction using an actual prior .docx as input. ## Environmental gotchas on claude.ai web chat | Gotcha | What happens | Workaround | |---|---|---| | Opus 4.7 biodefense classifier | Chat pauses mid-intake on vaccine, pathogen, or similar scenarios | Switch to Sonnet 4.6 or use Claude API | | Sonnet 4.6 output truncation | "Claude's response could not be fully generated" on very large PWS builds | Fresh chat retry; optionally add size-constraint language to the prompt | | Sonnet 4.6 silent self-approval | Skipped "proceed" gate on one run; applied defaults without confirmation | Read the Decision Summary carefully; interrupt to correct if needed | # Part 2: For Developers and Technical Reviewers ## Testing methodology Two testing waves, both in April 2026. Same methodology both times. - Worker sessions ran end-to-end in fresh claude.ai web chat, the same environment the skill's end users run in. - A separate Claude Code Opus 4.7 instance (1M context window, max effort mode, Claude Max 20x subscription) graded each run independently. - The grader had no access to the worker's drafting conversation or self-assessment until after grading was complete. This separation distinguishes "observed behavior" from "claimed behavior." - Each run was graded against a 14-point binary assertion matrix (8 general assertions plus 6 scenario-specific assertions). - All assertions were committed before seeing any worker output, to prevent the grading standard from drifting to accommodate whatever the worker happened to produce (same discipline as pre-registering a study). ## Grading methodology Each assertion received a binary pass or fail plus a one-line note on anything suspicious. **General assertions (8, unchanged across both waves):** - G1 Acquisition Strategy Intake collected before drafting - G2 Phase 1 decision tree executed with visible Decision Summary before docx generation - G3 .docx produced - G4 No FTE counts, SOC codes, or staffing tables in document body (FAR 37.102(d)) - G5 No staffing or IGCE appendix in document - G6 Staffing handoff chat-only, not saved as a file - G7 Section numbering sequential - G8 Key Personnel by role and qualifications, not by headcount **Scenario-specific assertions (6 per scenario):** tailored to the contract-type and workflow behaviors being exercised. ## Wave 1 (initial, pre-publication testing) Six canonical scenarios plus two AskUserQuestion stress tests. ### Wave 1 scenarios - **Scenario 1:** Non-commercial T&M SOW, Workflow A, AFRL test range instrumentation under FAR Part 15 and FAR 16.601. Exercised the T&M Labor Category Ceiling Hours exception under FAR 16.601(c)(2). - **Scenario 2:** Commercial FFP PWS, Workflow B (SOO conversion), Treasury Digital Customer Experience Modernization. Exercised SOO parsing, gap identification, and performance-based outcome framing. - **Scenario 3:** FFP PWS Workflow C scope reduction, cybersecurity incident response and digital forensics. Reduced from $62M per year to $40M per year. Exercised trade-off menu generation, user validation gate, section regeneration, updated staffing handoff, and Section 14 cut documentation. - **AskUserQuestion stress test:** One-sentence lazy prompt ("Need a PWS for help desk services for my agency. Can you help me write one?") to test whether the skill drives structured multi-choice intake when a non-expert user provides minimal context. ### Wave 1 results #### Opus 4.7 worker results | Scenario | Workflow | Assertions | Result | |---|---|---|---| | Non-commercial T&M SOW | A | 14 of 14 pass | PASS | | SOO-to-PWS Conversion | B | 14 of 14 pass | PASS | | Scope Reduction | C | 14 of 14 pass | PASS | #### Sonnet 4.6 worker results | Scenario | Workflow | Assertions | Result | |---|---|---|---| | Non-commercial T&M SOW | A | 14 of 14 pass | PASS | | SOO-to-PWS Conversion | B | 14 of 14 pass | PASS | | Scope Reduction | C | 13 of 14 pass | FAIL (G4) | #### AskUserQuestion feature test | Model | Behavior observed | |---|---| | Opus 4.7 | Defaulted to prose questions on first pass. Reached for the structured multi-choice tool only when explicitly prompted by the user. | | Sonnet 4.6 | Auto-triggered the structured multi-choice tool on the first message. Batched 12 questions across 4 groups of 3. Included "Not sure" and "Something else" escape hatches. Produced a 422-paragraph contract-file-ready PWS after the user selected option one for every question. | ### The one Wave 1 bug **Sonnet Scenario 3 G4 failure.** The Workflow C scope-reduction documentation in Section 14 contained specific staffing counts ("3-4 FTE examiners," "2-3 platform engineers," "8 Tier 2 analysts") as prior-state descriptors. Document bodies remain subject to FAR 37.102(d) even when describing historical or prior state. Opus passed the same scenario using capability-framed language ("night-shift onsite headcount eliminated," "no dedicated playbook engineer") without specific counts. Root cause: the skill's FAR 37.102(d) enforcement block was focused on staffing tables and the Phase 3 handoff. It did not explicitly address Section 14 scope-reduction narrative. Workers interpreted "describe what was cut" as license to quote the prior staffing plan. Fix: added explicit guidance in the Section 14 block requiring capability-and-coverage language, not staffing counts, for both current-state and historical-state descriptions. Included compliant and non-compliant examples. Prescribed the five-field documentation structure (Prior Scope, Revised Scope, Estimated Annual Savings, Rationale, Residual Risk). ## Wave 2 (follow-up, post-publication testing) Three contract-type and workflow combinations not exercised in Wave 1. Each run once on Opus 4.7 and once on Sonnet 4.6, for six runs total. ### Wave 2 scenarios - **Scenario 4:** Commercial FFP PWS, Workflow A, GSA PBS enterprise IT help desk for 25,000 users across 11 regions, $15M per year, 24x7x365, NIST 800-171 and CMMC Level 2 required. - **Scenario 5:** Non-commercial CPFF R&D PWS, Workflow A, NASA Glenn Research Center cryogenic fluid management research for in-space propellant transfer and long-duration storage, $45M with 8% fixed fee, 5-year term, Principal Investigator model, NASA FAR Supplement applies. - **Scenario 6:** Non-commercial Labor-Hour PWS (no materials), Workflow A, DHS CISA cyber threat analyst advisory support to the Cyber Threat Intelligence Integration Office, $8M NTE over 3 years, TS/SCI for all positions, 5 labor categories. Scenario 5 was originally planned as NIH/NIAID mRNA vaccine platform R&D. The Opus 4.7 safety classifier on claude.ai web chat paused the worker mid-intake. The scenario was swapped to NASA Glenn CFM (same contract mechanics, no biodefense trigger words) to preserve apples-to-apples grading across models. ### Wave 2 results | Run | Model | Scenario | Result | Failed assertion | |---|---|---|---|---| | 1 | Opus 4.7 | Commercial FFP (GSA PBS) | 14 of 14 PASS | — | | 2 | Sonnet 4.6 | Commercial FFP (GSA PBS) | 14 of 14 PASS | — | | 3 | Opus 4.7 | CPFF R&D (NASA Glenn) | 13 of 14 PASS | S5 | | 4 | Sonnet 4.6 | CPFF R&D (NASA Glenn) | 13 of 14 PASS | S5 | | 5 | Opus 4.7 | Labor-Hour (DHS CISA) | 13 of 14 PASS | S6 | | 6 | Sonnet 4.6 | Labor-Hour (DHS CISA) | 13 of 14 PASS | S6 | ### The two Wave 2 failure modes **Failure mode A: S5 on CPFF, both models.** Neither worker explicitly committed the document to completion form (FAR 16.306(d)(1)) or term form (FAR 16.306(d)(2)). Both documents cited FAR 16.306 multiple times but never made the form selection explicit. The distinction governs how fixed fee is earned: completion form ties earning to delivery of a specified end product; term form ties earning to level of effort over a stated time period. For R&D where technical success is uncertain, term form is usually more appropriate. Root cause: skill-template gap. The Phase 2 CPFF block does not currently require or emit an explicit form selection. Candidate patch: add a Phase 2 rule requiring the document to name completion form or term form wherever the contract framework first appears. **Failure mode B: S6 on Labor-Hour, both models.** Neither worker referenced DD Form 254 in the security section of a TS/SCI Labor-Hour PWS. Both covered clearance levels, SCIF requirements per ICD 705, and derivative classification, but DD Form 254 and Security Classification Guide references were missing. Root cause: skill-template gap. The Phase 2 security block does not currently emit DD Form 254 or SCG references for classified requirements. Candidate patch: when clearance above Confidential is called out, the Phase 2 security block must reference DD Form 254 and the applicable SCG. ### Behavioral observations from Wave 2 (not assertion failures) 1. **FAR 52.237-2 misreference for Key Personnel substitution.** Appeared on 5 of 6 Wave 2 runs. Sonnet corrected it to NFS 1852.237-72 on the NASA CPFF scenario but reverted to the template default on the other runs. Skill-template bug, not a worker variance. 2. **Sonnet skips AskUserQuestion intake on medium-ambiguity prompts.** Sonnet triggered AskUserQuestion on 0 of 3 Wave 2 scenarios. Opus triggered on 2 of 3. This inverts the Wave 1 finding, where Sonnet was the auto-trigger model and Opus needed explicit prompting. The Wave 1 patch fixed Opus; Sonnet has now drifted in the other direction. 3. **Sonnet self-approved Phase 1 Decision Summary on Scenario 6.** Said "Proceeding to document assembly now" immediately after presenting the summary, without waiting for user "proceed." This denies the user a chance to override applied defaults before generation. The G2 assertion (visible Decision Summary before docx) passed on visibility, but the user-gate is a separate concern. 4. **Section ordering drift.** QASP and Transition sections swap positions (Section 11 versus Section 12) across and within models. Terminology is always correct; only position drifts. Low severity. 5. **Opus 4.7 biodefense classifier block.** Environmental observation on claude.ai web chat, not a skill defect. The classifier paused the worker mid-intake on an NIH/NIAID mRNA vaccine R&D scenario. Federal acquisition users working on biodefense, vaccine, or pathogen-related contracts should expect potential classifier blocks on Opus web chat and plan to use Sonnet or the Claude API. 6. **Sonnet 4.6 output truncation on large documents.** Environmental observation. First attempt on Scenario 4 Commercial FFP PWS hit "Claude's response could not be fully generated" mid-.docx-generation. Fresh chat retry completed cleanly. 7. **Sonnet domain reasoning on CPFF was notably sharp.** Sonnet on Scenario 5 caught that CPFF fee cannot be administratively reduced under FAR 16.306 and flagged the skill's QASP fee-deduction language as legally problematic, recommending either repositioning as a cure-notice tool or conversion to CPAF/CPIF. That is substantive acquisition-law reasoning, not template output. Opus did not surface this on its parallel CPFF run. ## Cumulative results across both waves | | Wave 1 | Wave 2 | Total | |---|---|---|---| | Worker runs | 8 (6 canonical + 2 AskUserQuestion) | 6 | 14 | | Scenario assertions graded | 84 | 84 | 168 | | Passed | 83 | 80 | 163 | | Failed | 1 (fixed in patch) | 4 (cross-model, patch candidates) | 5 | ## Skill patches: shipped and candidate ### Shipped after Wave 1 (9 patches) | Patch | Section affected | Trigger | |---|---|---| | Unconditional handoff rule moved to top of Phase 3 with emphatic language | Phase 3 opening | 3 of 7 runs skipped handoff emission | | SOW "Inspection and Acceptance" versus PWS "QASP" label differentiation | Phase 2 Section 12 block | Scenario 1 both models | | Table of Contents instruction for documents exceeding 8 sections | Phase 2 opening | All runs produced 300-700 paragraph docs without TOC | | SOO-implied objectives rule | Phase 0 | Scenario 2 both models added unstated objectives legitimately | | Phase 2 Invocation Gate requiring Phase 1 Decision Summary before docx | Phase 2 opening | Cross-model Phase 1 compression plus network-blip resilience | | Section 14 assumption format template (4-column) | Phase 2 Section 14 block | Workers independently invented divergent structures | | Workflow C Section 14 compliance rule (no FTE counts in cut descriptions) | Phase 2 Section 14 block | The Wave 1 G4 failure | | AskUserQuestion tool usage instruction | Phase 1 opening | Opus defaulted to prose without explicit instruction | | Anti-redundancy rule | Phase 1 opening | Both models re-asked explicit answers | Skill version lines: 361 before Wave 1 patches, 380 after. Ceiling: 1,000. ### Shipped after Wave 2 (5 patches) | Patch | Section affected | Trigger | |---|---|---| | Agency-aware Key Personnel substitution clause selection (removed the wrong FAR 52.237-2 default; added agency-specific guidance for NASA, DHS, HHS, DoD, and a generic fallback) | Phase 2 Key Personnel block | FAR 52.237-2 emitted on 5 of 6 Wave 2 runs | | Explicit CPFF form commitment required (skill must name completion form per FAR 16.306(d)(1) or term form per FAR 16.306(d)(2) wherever contract framework first appears) | Phase 2 Section 5 block | S5 failure on both Wave 2 CPFF runs | | DD Form 254 and SCG reference added to Section 9 for any classified requirement at Confidential or higher | Phase 2 Section 9 block | S6 failure on both Wave 2 LH runs | | Strengthened Phase 1 proceed gate; explicit "DO NOT self-approve" rule with requirement that the response must END after presenting the Decision Summary | Phase 2 Invocation Gate | Sonnet self-approved on Scenario 6 | | Fixed section ordering made explicit and prescriptive (Section 11 Reporting, Section 12 QASP, Section 13 Transition, Section 14 Constraints — do not merge, combine, swap, or rename) | Phase 2 Section Structure header | QASP and Transition position drift across runs | Skill version lines: 426 before Wave 2 patches, 450 after. Ceiling remains 1,000. These patches are shipped in the current skill but have not yet been validated against a fresh test wave. Wave 3 will regression-test the same six scenarios against the patched skill to confirm the Wave 2 failure modes no longer reproduce. ### Post-Wave 2 bloat trim (5 cuts, no rules removed) Following a discipline pass on April 23, 2026, five pieces of non-load-bearing prose were removed from the skill without altering any behavioral rule. Every rule that was enforced before the trim is still enforced. | Cut | Section | What was removed | |---|---|---| | Section-ordering trailing anecdote | Phase 2 Section Structure | "Workers have been observed placing QASP at Section 11 and Transition at Section 12..." — testing history narrated as a rule. The preceding prescriptive sentence already covers it. | | Forbidden-appendix bullet enumeration | Phase 2 Appendices | Six specific forbidden appendix titles condensed into one inline example list. Universal rule is "no staffing-related appendix"; exhaustive enumeration was bloat. | | Staffing Handoff "Why this rule exists" paragraph | Phase 3 Staffing Handoff | Self-referential history about prior skill versions compressed to two sentences. The rule above is the rule; the why just needs to support judgment. | | CLIN Handoff DO NOT block duplication | Phase 3 CLIN Handoff | The DO NOT list repeated the staffing handoff pattern verbatim. Replaced with "Same DO NOT rules apply as for the Staffing Handoff above" plus a one-line variant summary. | | "Experienced 1102s flag this immediately" filler | Phase 3 CLIN Handoff | Rhetorical filler with no rule content. | Skill version lines: 450 before trim, 437 after. Ceiling remains 1,000. ## What was not tested - **Special Access Program and Special Access Required variants.** Classified coverage was limited to collateral Secret, Top Secret, and TS/SCI contexts in notional scenarios. - **Commercial CPFF or commercial Labor-Hour.** Both require FAR 12.207(b) written Determinations and Findings and are edge cases on top of edge cases. - **Hybrid contract structures** (FFP plus T&M or LH CLINs; FFP with cost-reimbursable travel CLINs was exercised only incidentally in the Wave 2 handoff tables). - **True document-to-document scope reduction** using an actual prior .docx as Workflow C input. The Wave 1 test used a bulleted summary of the prior PWS rather than an actual document. - **Multi-year ramping AQL structures** where performance targets escalate across option years. Exercised informally in Wave 1, not systematically. - **Hardware-intensive Test and Evaluation domains** beyond the single Wave 1 Scenario 1 (AFRL test range) and Wave 2 Scenario 5 (NASA CFM). - **Grants, cooperative agreements, and Other Transaction agreements** are out of scope for this skill. Separate skills cover those. Users working in these contexts should expect to validate outputs more carefully and may encounter edge cases that the test waves did not surface. ## Note on classified scenario testing Wave 1 Scenario 1 used a notional classified-context prompt (AFRL test range, Secret and Top Secret collateral clearances). The test verified only that the skill produces proper classified-contract boilerplate: clearance-by-position requirements, Security Classification Guide references, OPSEC language, facility clearance statements, and DD Form 254 references at the prompt level. All scenario inputs were fictional. No actual classified information was processed, transmitted, or stored at any point during testing. The skill itself does not and cannot handle classified data; it produces unclassified acquisition documents that describe requirements for classified work, which is standard practice for any SOW or PWS. Wave 2 Scenario 6 (DHS CISA TS/SCI) extended classified-context coverage to a TS/SCI Labor-Hour PWS using a fictional CISA scope. Same constraint: no real classified information was ever handled. Coverage was limited to collateral Secret, Top Secret, and TS/SCI contexts. Special Access Program and Special Access Required variants were not exercised. ## Independent grading Every assertion was graded by a separate Claude instance reading only the worker's final output and the chat transcript. The grader had no access to the worker's internal reasoning and did not see the worker's self-assessment until after grading was complete. This separation is the load-bearing credibility claim of this testing program: it distinguishes "the worker claimed to do X" from "the output demonstrates X." --- **Testing Methodology** Evaluators: James Jenrette (1102tools) and Claude Code Opus 4.7 (1M context window, max effort mode, Claude Max 20x subscription). Worker models tested: Claude Opus 4.7 and Claude Sonnet 4.6 on claude.ai web chat, the same environment the skill's end users run in. Wave 1: 8 runs, 84 assertions, 83 passes, 1 confirmed failure (fixed in patch). Wave 2: 6 runs, 84 assertions, 80 passes, 4 confirmed failures (cross-model, patch candidates identified). Cumulative: 14 runs, 168 assertions, 163 passes. Date: April 2026 (both waves). Skill: sow-pws-builder. Source: github.com/1102tools-dev/federal-contracting-skills. License: MIT.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.