Claude Skill

docs

Imported from gal-a/qikly/docs.

LLM Mart · 0 points · 1 views 4 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download gal-a-qikly-docs-30b587b.zip · 2704 KB
gal-a/qikly 15 0 forks Apache-2.0 Updated 1d ago
Part of gal-a/qikly — 2 skills

Install

skills CLI npx skills add https://github.com/gal-a/qikly/tree/main/docs
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install gal-a-qikly@llmmart
Git git clone https://github.com/gal-a/qikly.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole gal-a/qikly collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

qikly as an agent Skill

What this gets you: your coding agent stops needing to be told about qikly. Ask it for tests you can trust and it reaches for the tool on its own, and it knows the part that is hard to guess, which line of your specification is a decision the coder needs and which is a consequence to withhold.

A Skill is a folder of Markdown your agent reads when what you are asking matches what the Skill says it is for. No server, no configuration, no process to keep running.

Install it

First, make sure the qikly you are about to run is the current one. The Skill ships inside the package, so an old qikly writes an old Skill, and it then describes commands that do not exist yet.

pip uninstall qikly -y
pip install qikly
qikly --version

Uninstall first rather than --upgrade: an interrupted or repeated upgrade can leave more than one version's metadata behind, and pip then reports one version while the files on disk are another's. The clean pair takes seconds and removes the question. Compare what --version prints against the latest release.

Then, from the directory of the project you want the Skill in:

qikly --install-skill

That writes .claude/skills/qikly/, which is where agents look. Then open your agent in that directory and ask for something ordinary, without mentioning qikly:

Write tests for src/following_distance.py that would actually catch a bug.

src/following_distance.py stands in for a module you actually have; name a real one, because an agent asked about a file that does not exist will spend its answer asking you which file you meant. Every example on this page uses following distance from radar samples, which is the worked case in the Skill itself, so the two read together.

It should reach for qikly by itself, and it does: first time in Claude Code, Gemini CLI, Codex and Cursor, in a project with a rival testing skill installed beside this one and a request that never mentioned qikly. That is the hard version of the test, because the agent had a competing option and no hint. Your own project has more skills in it than that one did.

If it does not, see if your agent does not pick it up.

Using GitHub Copilot in VS Code? Copilot reads none of the skill folders above. Run qikly --install-skill copilot, which writes the Skill into .github/instructions/ along with the qikly.instructions.md file Copilot actually opens, then use Copilot Chat in Agent mode. That path has not been watched loading yet, so tell us if it works for you. The MCP server is the other way in and is tested there: docs/mcp.md.

A few options, none of them needed the first time. --dry-run shows what it would write. --force replaces a Skill you have already edited, keeping a timestamped copy of the old one. And --install-skill agents, cursor or gemini write where those tools document their own skill folders, rather than Claude's.

What is in it

  • The one question: given only the requirements, could two competent developers legitimately disagree about this line? Yes means it is a decision and belongs in requirements; no means it is a consequence and belongs in acceptance_criteria.
  • Both mistakes and their signatures, so an agent can recognise a stalled loop as a spec problem rather than a code problem.
  • How to write a criterion that can be tested: name values not adjectives, name both sides of a boundary, and make sure the data contains what you name.
  • The five steps on your own module, and which commands are free.
  • What to do when a run does not converge, and the one thing never to do, which is loosen a criterion to get green.
  • The honest limit, so an agent does not oversell it on your behalf.
  • Two bundled references: the task file reference and the troubleshooting guide, complete, so the Skill works with no network.

Skill or MCP server, and why qikly has both

Skip this if you are not using the MCP server. The Skill works on its own. This is here because the two look interchangeable and are not: they answer different halves of the same problem, and neither replaces the other.

MCP server Skill
What it ships a running process exposing typed tools a folder of Markdown
What it gives the model the ability to call qikly_scaffold, qikly_validate, qikly_run, qikly_explain the judgement around those calls
Setup qikly --install-mcp, then host configuration qikly --install-skill
Can enforce a rule yes, in server code no, it is instructions
Works when the agent has no shell yes no

The MCP server is the one that can enforce things. qikly_run is built so it cannot return the acceptance criteria, and a test in qikly's own suite fails the build if any code path lets a criterion through. That guarantee lives in code, and it is why the server exists.

The Skill is the one that can teach. Which line of your spec is a decision and which is a consequence; that a repeating identical patch means a decision is in the wrong half; that a criterion naming a value your data never holds produces a test that passes whatever the code does. None of that is a function call, and an agent that does not know it will use qikly and get less out of it.

The guarantee does not rest on the Skill. A Skill is text in a context window, so it can inform but not enforce. The withholding is enforced by the tool, in code, whether or not this Skill is ever installed. The Skill's own text says so, and a test asserts that it still does.

If your agent does not pick it up

Three causes, in the order to check them.

1. Your agent is not in that directory. The Skill is per project. It is files on disk, so an agent started in another folder, or running in a browser with its own sandbox rather than on your machine, cannot see them. Start the agent in the directory you installed into. This is the commonest cause by some distance.

2. Just name it. This always works, because it does not depend on the agent being told what the Skill is for:

Use the qikly skill to write tests for src/following_distance.py.

If naming it works and the neutral request did not, the Skill is fine and the problem is discovery.

3. The host did not pass the description along. Whether an agent reaches for a Skill unprompted depends on how much it explores before it starts typing, and on what the host told it. Claude Code reserves a fraction of the context window for the whole skill listing, 1% by default; when the listing does not fit, Anthropic's own bundled skills keep their descriptions and everything else is ranked by how often you have used it. So a skill you have never invoked can arrive as a bare name with nothing to match against. Raising skillListingBudgetFraction in your Claude Code settings gives the listing more room, and that single change turned four failed routing attempts into a clean one during testing.

Which is why naming it once is more than a workaround. That ranking is by how often you have used each skill, decaying over about a week, so a skill you have never invoked sorts below every skill you have. Name it in one request and it moves above them, and the next neutral request has a much better chance of finding it by itself. The first invocation is the only hard one.

Checking that it works

A Skill either loads or it does not, and it never tells you which, so it is worth five minutes once. If you only do one of these, do number four: the others check that the Skill arrived, and that one checks that it is right.

1. The files are where your agent looks for them. In the project you ran qikly --install-skill in:

ls .claude/skills/qikly/SKILL.md        # macOS, Linux
dir .claude\skills\qikly\SKILL.md       # Windows PowerShell

And start your agent in that same directory. The Skill is per project, not per machine: an agent started somewhere else, or running in a browser with its own sandbox rather than on your computer, cannot see these files and will never load them. That is the commonest reason a correctly installed Skill appears to do nothing.

2. It loads on a request that should trigger it. Start a fresh session and ask for something in its territory, naming a real module of your own and not mentioning qikly:

Write tests for src/following_distance.py that would actually catch a bug in it.

The agent should mention qikly, or the decisions-and-consequences split, unprompted. If it does not, see when an agent does not pick it up below before changing anything.

3. It does not load when it should not. Ask something unrelated, such as "rename this variable everywhere", and it should stay quiet. A Skill that loads for everything costs context on every request.

4. It gives the right answer to the question that matters. This is the one to do if you do only one. Ask:

My spec says "warn when following distance breaks the two-second rule", and the acceptance criteria say a headway of exactly 2.00 s does not warn. Is that the right split?

The answer should be no, and the reason should be that "breaks" can be read as "below" or "at or below", so the boundary is a decision and belongs in the requirements. That is the single most valuable thing in the Skill, and if it comes back wrong the rest does not matter much.

5. It does not claim more than the tool does. Ask what guarantees the coding agent never sees the criteria. The answer should point at qikly's code and its build-failing test, not at the Skill.

Keeping it current

An installed Skill does not update itself, and until 0.5.4 nothing told you. pip install --upgrade qikly replaces the package; it cannot touch a folder copied into your project, so after an upgrade you can be following instructions that name a different set of commands. From 0.5.4 any qikly command says so when it notices:

note: the qikly Skill in .claude\skills\qikly is older than this qikly, so
it describes a different set of commands. `qikly --install-skill --force`
replaces it and keeps a copy of the old one.

The path is printed with your platform's own separator, so it reads with forward slashes on macOS and Linux.

--force keeps a timestamped copy of what it replaces, so a Skill you have edited is recoverable. That is the whole update mechanism: qikly tells you, and you run one command. There is no background process and nothing phones home; the check is two version strings read from two files on your disk.

The Skill's version is its own and moves when its instructions move, not when qikly releases, so a release that does not touch it produces no notice.


The Skill itself lives at src/qikly/skills/qikly/ and ships inside the installed package, so --install-skill works from pip install qikly as well as from a clone.

A stale Skill is worse than no Skill, because an agent quotes it with confidence and the reader has no way to tell. So it is updated whenever qikly changes in a way it describes: a new or renamed flag, a change to which commands are free, a re-measured convergence figure, a new failure mode worth carrying.

Half of it cannot go stale on its own. The two bundled references are asserted byte-identical to docs/TASK_FILE_REFERENCE.md and docs/TROUBLESHOOTING.md by a test, so a drifted copy fails the build. The prose in SKILL.md has no such test, which is why it is on the release checklist instead.

Files (qikly)
  • images
    • qikly_demo.gif 972.5 KB · in bundle
    • qikly_flow.png 203.1 KB · in bundle
    • qikly_hero.png 1.1 MB · in bundle
    • qikly_hero.svg 8.1 KB · in bundle
    • qikly_icon.png 6.2 KB · in bundle
    • qikly_icon.svg 1.7 KB · in bundle
    • qikly_icon_light.png 5.9 KB · in bundle
    • qikly_icon_light.svg 2 KB · in bundle
    • qikly_social.png 359.8 KB · in bundle
    • qikly_wordmark.svg 3.2 KB · in bundle
  • .nojekyll 0 B · in bundle
  • apple-touch-icon.png 17.4 KB · in bundle
  • CNAME 15 B · in bundle
  • CONFIGURATION.md 6.8 KB
    # Configuration and LLM provider
    
    Settings, environment variables and provider setup. For getting a key onto a
    machine, into CI, or diagnosing one that is wrong, see
    [PROVIDER_KEY_SETUP.md](https://github.com/gal-a/qikly/blob/main/docs/PROVIDER_KEY_SETUP.md).
    
    ## Configuration
    
    ### Timeouts
    
    Every model call has a deadline of **300 seconds**, set by
    `QIKLY_REQUEST_TIMEOUT` in seconds. `0` waits forever, which is what provider
    SDKs do by default and is why the setting exists: a stalled connection blocks a
    call that never raises, so nothing downstream can react to it. With a deadline
    the same stall becomes an ordinary transient error and is retried with backoff.
    
    A task process prints `still running, N minutes elapsed` every five minutes, so
    that "not answering" is visible rather than inferred.
    
    
    `config/settings.yaml`, shared across all tasks:
    
    | Key | Meaning |
    |---|---|
    | `orchestrator.max_retries_per_stage` | Attempt budget per stage before the run raises. If you see the exact same patch content repeating verbatim, that's usually a requirement fighting the model's real-world prior (see below) rather than a budget problem. If instead each attempt is a *different* patch that never resolves the same failing test, that's a different signal: a bug that needs more than the failure text to resolve, rather than an artificial requirement; see [Where it fits today](https://github.com/gal-a/qikly/blob/main/README.md#where-it-fits-today) on telling the two signatures apart. |
    | `orchestrator.test_order` | Stage order; `unit` is always forced last (it's generated from the implementation, which doesn't exist yet during integration/system). |
    | `agent.max_patch_size` | Rejects an oversized PATCH and asks the model to retry smaller. Tune per task if a bigger implementation needs more room. |
    | `logging.save_transactions` | Turns off `transactions_*.jsonl` logging entirely; also disables the HTML report, which reads that log. |
    
    ## LLM provider
    
    Every call in a run goes to one provider. One provider per run; mixing them
    per agent role is not supported.
    
    ### Getting a key
    
    | Provider | Where the key comes from | Install |
    |---|---|---|
    | **Gemini** (default) | [aistudio.google.com/apikey](https://aistudio.google.com/apikey) | included |
    | **OpenAI** | [platform.openai.com/api-keys](https://platform.openai.com/api-keys) | `pip install "qikly[openai]"` |
    | **Anthropic** | [console.anthropic.com/settings/keys](https://console.anthropic.com/settings/keys) | `pip install "qikly[anthropic]"` |
    
    Then export the key and pick the provider:
    
    ```bash
    # Gemini, the default. Nothing else needed.
    export GEMINI_API_KEY=...
    qikly --demo
    
    # OpenAI
    pip install "qikly[openai]"
    export OPENAI_API_KEY=...
    export LLM_PROVIDER=openai
    qikly --demo
    
    # Anthropic
    pip install "qikly[anthropic]"
    export ANTHROPIC_API_KEY=...
    export LLM_PROVIDER=anthropic
    qikly --demo
    ```
    
    On Windows PowerShell, `$env:OPENAI_API_KEY = "..."` instead of `export`.
    
    `API_KEY` works for any of them, and each provider's own conventional variable
    is accepted too, so a machine already configured for one needs nothing extra.
    
    **A note on which provider to start with.** Every convergence figure in this
    README was measured on `gemini-3.5-flash-lite`, over hundreds of runs, and the
    bundled demo is tuned to that path. The other providers work and are far less
    travelled here, and an entry-level model on any of them may stall on tasks that
    the measured path clears. If a provider you have chosen converges poorly, reach
    for a stronger model on it before concluding anything about the tool: model
    choice moves convergence more than any setting in this file.
    
    **And which model.** Prefer a small fast one: a run makes one call per stage
    per iteration, so a model that reasons before it answers turns minutes into
    tens of minutes. `gemini-3.5-flash-lite` is the floor the published figures
    were measured on, not a best case, and a later flash model should clear it.
    The table of what to start with per provider, and when a thinking model is
    worth the wait, is in
    [docs/PROVIDER_KEY_SETUP.md](https://github.com/gal-a/qikly/blob/main/docs/PROVIDER_KEY_SETUP.md#which-model).
    
    **PowerShell, CI, persisting a key, restricting one, and what a wrong key or
    a wrong model looks like:** [docs/PROVIDER_KEY_SETUP.md](https://github.com/gal-a/qikly/blob/main/docs/PROVIDER_KEY_SETUP.md).
    
    ### Checking a key works, for about a cent
    
    ```bash
    qikly --validate                       # free: does not touch the network
    LLM_PROVIDER=openai qikly --demo       # one task, about a cent
    ```
    
    `--demo` is the real test. It makes actual calls, writes to a throwaway folder,
    and reports the model, the estimated cost and whether it converged. A wrong key
    fails on the first call with a message naming what to check.
    
    `--validate` will not catch a bad key, because it never opens a socket. That is
    the point of it.
    
    ### The variables
    
    | Variable | Meaning |
    |---|---|
    | `LLM_PROVIDER` | `gemini` (default), `openai` or `anthropic` |
    | `LLM_MODEL` | Overrides the provider default: `gemini-3.5-flash-lite`, `gpt-4o`, `claude-sonnet-5` |
    | `API_KEY` | The key. Provider-specific names above are accepted too |
    | `QIKLY_REQUEST_TIMEOUT` | Seconds per call, default 300. `0` waits forever |
    | `QIKLY_MAX_CALLS` | Hard stop after N model calls, for an unattended run |
    
    One provider per run. Mixing them per agent role is not supported.
    
    ### Determinism is best-effort, and uneven
    
    With a seed set, Gemini and OpenAI are called at `temperature=0` and are given
    the seed itself. **Anthropic gets neither.** Its Messages API has never had a
    seed, and SDK 1.x removed `temperature` from `messages.create()` entirely, so
    there is no sampling lever left to pull.
    
    That matters if you compare rates across providers: an Anthropic figure carries
    more run-to-run variance than the others by construction. No provider promises
    identical output either way, so treat all of this as reduced drift rather than
    reproducibility.
    
    ### Keeping providers working
    
    The SDKs are other people's code on other people's release schedules, and this
    is the part of qikly most likely to break without you touching it. Two habits
    cover it:
    
    ```bash
    pip install "qikly[all-providers]"
    python -m pytest tests/test_provider_signatures.py -v
    ```
    
    That reads the signature of every SDK you have installed and compares it
    against what qikly sends, so a removed or renamed parameter fails a test rather
    than a user's first run. It is how the Anthropic `temperature` break was found.
    
    What it cannot catch is a parameter that still exists and now means something
    different, or a model name retired server-side. **One `--demo` per provider
    before each release** covers that, costs a few cents, and is the only check
    that exercises the real API.
    
    Dependencies carry upper bounds for the same reason. Raising one after testing
    is a two-line change; not having one lets a major version arrive unannounced.
    
    
  • design_1_case_study.md 16.2 KB
    <!--
    One of three. See MAINTAIN.md: these three files, README.md, docs/index.html
    and the investor deck move together. A number changed here has to change in
    all of them.
    -->
    
    # Engineering a Better System for Testing AI-Generated Code
    
    **Your AI writes both the code and its tests. How do you know the tests are really valid?**
    
    **The solution: two agents.** One turns the acceptance criteria into tests.
    The other writes the code and **never sees the acceptance criteria.**
    
    ![Qikly: automated code and test generation, kept apart](images/qikly_hero.png)
    
    Imagine a student who writes the exam paper, writes the answer key, and then sits the exam. They pass. Obviously they pass. Nobody would accept that as evidence the student knows the material, and nobody should.
    
    That is what happens when one model is handed a specification containing the acceptance criteria and asked to produce both the implementation and the suite that checks it. It reads the spec, resolves every ambiguity in it, writes code according to those resolutions, and then writes tests according to *the same resolutions*. Where the spec said "reject malformed rows" and left "malformed" undefined, the agent picked a definition, implemented it, and then tested that definition. Everything goes green. It was always going to.
    
    Run one model at temperature zero on both jobs and the two resolutions are almost identical *by construction*: the verification step burns compute and returns a tick that carries no information about whether the code is correct. That effect should get worse as models improve, since every gain in determinism tightens the agreement between the code and the tests that judge it. That last sentence is a working assumption rather than a measurement, and nothing here tests it.
    
    The fix is to take the answer key away from the student. Test generation gets the acceptance criteria in full. The coding agent gets the same specification with that section cut out, and when a test fails it sees pytest's report of that failure, never the specification section it was cut out of. Now a green suite means something happened: code written by someone who could not read the standard nevertheless satisfies it. This post is about a tool I built to do exactly that, and part 1 shows what it is and one real repair followed end to end, so you can judge the idea on something concrete rather than on a claim.
    
    ---
    
    **This is part 1 of three.**
    
    | | | |
    |---|---|---|
    | **1. The case** (you are here) | [2. How well it works](design_2_performance.md) | [3. How it is built](design_3_mechanism.md) |
    
    ## How it works, in one diagram
    
    ```mermaid
    flowchart TD
        SPEC["<b>Full specification</b><br/>task.yaml<br/>requirements + interface<br/>acceptance_criteria"]
        REQ["requirements<br/>+ interface"]
        AC["acceptance_criteria"]
        CODE["<b>Coding agent</b><br/>writes the implementation<br/>FIX then PATCH on failure"]
        TEST["<b>Test-writing agent</b><br/>writes the suite"]
        IMPL["Implementation"]
        SUITE["<b>pytest suite</b><br/>Tests for:<br/>1 integration, 2 system,<br/>then 3 unit"]
        RUN{"Run the suite"}
        FAIL["<b>pytest output</b> only<br/>no acceptance criteria"]
        OUT["Converged<br/><b>outputs:</b> code + suite<br/>+ audit trail"]
        STALL["Did not converge<br/><b>failure errors and audit trail</b><br/>exits non-zero, ships nothing"]
    
        SPEC --> REQ
        SPEC --> AC
        AC -. "never reaches" .-x CODE
        REQ --> CODE
        REQ --> TEST
        AC --> TEST
        CODE --> IMPL
        TEST --> SUITE
        IMPL --> RUN
        SUITE --> RUN
        RUN -- pass --> OUT
        RUN -- fail --> FAIL
        FAIL -- "repair loop:<br/>FIX, then PATCH" --> CODE
        FAIL -- "retry budget spent" --> STALL
        IMPL -. "unit stage only:<br/>written last, from the code" .-> TEST
    
        classDef codeView fill:#f3e8ff,stroke:#7e22ce,color:#4c1d95
        classDef standardView fill:#d9ebea,stroke:#0e6a70,color:#0b3d40
        classDef converged fill:#dcfce7,stroke:#15803d,color:#14532d
        classDef stalled fill:#fdf0d5,stroke:#b45309,color:#78350f
        class CODE,IMPL,FAIL codeView
        class AC,TEST,SUITE standardView
        class OUT converged
        class STALL stalled
        linkStyle 10 stroke:#15803d,stroke-width:2px
        linkStyle 11,12 stroke:#7e22ce,stroke-width:2px
        linkStyle 13 stroke:#b45309,stroke-width:2px
    ```
    
    **Purple is what the coding agent can see. Teal is what the standard is
    written from.** They never touch. The purple arrows are the repair loop, and
    that is where almost all of a run happens. A run that never converges is still worth having: it exits
    non-zero, names the blocking tests, and keeps the same complete record. A failing suite sends the coding agent the failure text and nothing
    else, it produces a fix and a patch, and the suite runs again. That cycle
    repeats until the stage passes or the retry budget runs out, and each stage
    clears before the next is generated.
    
    | | Receives |
    |---|---|
    | **The test generation agent** | The requirements, the interface contract, and every acceptance criterion in full. It writes integration, system and unit tests against the standard. |
    | **The coding agent** | The same file with the criteria section removed, plus the text of whatever test just failed. The same vague brief a developer usually works from. |
    
    Two dotted lines in the diagram, pointing opposite ways. The first is the whole
    idea: the coding agent is never told the acceptance criteria it is judged
    against, only the requirements and the interface, so when a test fails it has to
    reason from behaviour rather than recall an answer it was given. The second is the one exception, and it runs the other way: unit tests
    are written last, from the code, because they have to name real functions.
    
    ---
    
    ## One simple end to end repair loop, for a code implementation that did not originally enforce a maximum value for the tax rate (100%)
    
    Everything below is lifted from a single real run of the bundled `CALC_TAX`
    task, on `gemini-3.5-flash-lite`. It is abbreviated, not
    invented: the prompt fragments, the failure text, the reasoning and the diff
    are what actually passed through the loop.
    
    ### 1. What the coding agent is given
    
    The task file has three parts. The coding agent receives two of them.
    
    First the **requirements**, in the words a person would use:
    
    ```yaml
    requirements:
      - "Read both CSV files listed in this task's inputs (input_01.csv and
         input_02.csv), and combine their rows into a single list before
         validation"
      - "Validate each row: order_id, item_price, quantity, tax_rate
         (a percentage, e.g. 8.25 means 8.25%)"
      - "Apply reasonable, strict real-world data-quality validation ...
         reject anything malformed, implausible, or out of range"
    ```
    
    The task reads **two** CSV files rather than one because the orders arrive as
    two separate batches, and combining them before validation is itself part of
    what the tests check. Nothing about the tax rule depends on there being two;
    it is there so the pipeline has a real extract step to get wrong.
    
    Note what "reasonable" and "out of range" do not say. They do not say where the
    range ends. A developer handed this would have to decide, and so does the
    agent.
    
    Then the **interface**, the contract as a description rather than as code:
    
    ```yaml
    interface:
      module: "outputs.agent_src.code.CALC_TAX.calc"
      integration_functions:
        - "extract(input_path) -> list[dict]"
        - "transform(rows) -> dict   # {\"accepted\": [...], \"rejected\": [...]}"
        - "load(data, output_path) -> None"
    ```
    
    ### 2. What it is not given
    
    The same file carries eleven acceptance criteria. This is one of them:
    
    ```yaml
      - "tax_rate is only valid if it is a non-negative number -- a negative
         tax_rate is invalid. A tax_rate representing more than 100% is also
         invalid"
    ```
    
    That criterion goes to the test-writing agent in full. It is cut out of the
    file before the coding agent is handed it, and `qikly --explain CALC_TAX`
    prints both halves so you can see the cut.
    
    ### 3. The first code implementation, and the test it fails
    
    The agent writes `calc.py` from the spec above. Its validation reads:
    
    ```python
    tax_rate = _parse_decimal(rate_str)
    if tax_rate is None or tax_rate < 0:
        # reject
    ```
    
    The coding agent only wrote half the rule. It has the half that rejects a
    negative rate, and it is missing the half that asserts the value does not
    exceed 100%. Therefore the test fails. The suite, written from the full
    criteria, contains a test the agent has never seen:
    
    ```
    outputs/tests/CALC_TAX/integration/test_integration.py::test_tax_rate_validation_rules FAILED
    
        for row in result["accepted"]:
            rate = float(row["tax_rate"])
    >       assert 0 <= rate <= 100
    E       assert 150.0 <= 100
    ```
    
    8 of 9 integration tests pass. This one does not.
    
    Here is the whole test, because how little there is of it is the part worth
    reading:
    
    ```python
    def test_tax_rate_validation_rules():
        rows_1 = extract(INPUT_01)
        rows_2 = extract(INPUT_02)
        result = transform(rows_1 + rows_2)
    
        for row in result["accepted"]:
            rate = float(row["tax_rate"])
            assert 0 <= rate <= 100
    ```
    
    **It checks the rule in one direction only**, and that is worth noticing rather
    than glossing. Every accepted row must be in range, which is the assertion that
    failed. Nothing here requires that a row rejected *for* `tax_rate` was actually
    out of range, so an implementation that rejected every row would satisfy this
    test. The criterion rules that out; this test does not enforce that half of it.
    
    That is the honest state of a generated suite, and it is the reason the tool
    reports which criteria a suite traces to rather than asking you to trust that
    the coverage is complete.
    
    **It constructs no input.** It reads whatever the fixture files hold and
    asserts a property of the output. That is forced by when it was written, which
    is the next point, and it is why the failing value is 150.0 rather than
    something the test chose: the bad row was already in the data.
    
    **This run predates criterion traceability**, which is why the test above
    carries no marker. Current runs put the criteria a test came from on the last
    line inside its docstring, as `# Criteria: 7`, so a reader can follow any test
    back to the line of the specification it enforces. Test generation writes that;
    `qikly --explain CALC_TAX` prints the criteria in the same order. The quick
    start shows the current shape.
    
    **Why an integration test and not a unit test?** Because of when it was
    written. At that point no implementation existed, so the only names the
    test-writing agent could use were the ones the `interface` section promised:
    `extract`, `transform`, `load`.
    
    This test calls **`extract` then `transform`**, feeding the first one's output
    into the second, and stops there: `load` only writes the result to disk, and the
    tax rule is already visible in what `transform` returns.
    
    Integration tests are not every pair of functions. They follow **the chain the
    requirements describe**, `extract` to `transform` to `load`, and each test walks
    as much of that chain as its property needs. There is one pipeline, and the
    number of tests is set by how many properties the acceptance criteria assert
    about it. Nine here, for eleven criteria. The single end-to-end entry point, `run_calc`, is tested separately at
    the system stage.
    
    That constraint shapes the assertion. It cannot feed the code a tax rate of 150
    directly, because it does not choose the input; it reads whatever the fixture
    files contain. So it asserts an **invariant over the output**: whatever ends up
    accepted must have a rate between 0 and 100. The bad row was already sitting in
    the fixture data, and the invariant caught it.
    
    The unit stage, generated later from the finished code, writes the same rule the
    other way round, because by then it can import internal helpers and construct
    inputs:
    
    ```python
    invalid_rates = ["-0.01", "-5", "100.01", "150", "abc", ""]
    for tr in invalid_rates:
        res = transform([{"order_id": "F1", "item_price": "10.00",
                          "quantity": "1", "tax_rate": tr}])
        assert len(res["accepted"]) == 0
    ```
    
    Same criterion, two stages, two styles: an invariant over real data first, then
    the constructed edge cases once there is code to point at.
    
    ### 4. The FIX, what goes back to the coding agent
    
    Pytest's own output for the tests that are failing right now, and nothing else.
    Be precise about what that contains, because it is more than one line: the test
    name, its source from `def` down to the failing statement, its docstring and
    comments if it has any, the assertion that failed, and the intermediate values
    pytest prints underneath it. Source after the failing line is not shown. Tests
    that are currently passing give up their names, because pytest lists everything
    it collected, but not their bodies, docstrings or assertions.
    
    What it never contains is the specification. The acceptance criteria are
    stripped from the task before the coding agent sees it, and no part of the
    criteria text is ever placed in a FIX or PATCH prompt. So the agent can read
    the one case it just failed, in the test author's words, and cannot read the
    rule that case came from, the other criteria, or the tests it has not failed
    yet.
    
    That distinction is the whole design, and it is narrower than "the agent is
    blind". A developer handed a failing test sees the same thing. What neither of
    them can do is change a test that was written before the code existed.
    
    From that input it produces a **FIX**, which is **reasoning rather than code**:
    
    ```
    failure_summary: test_tax_rate_validation_rules failed because a tax rate
                     greater than 100 was incorrectly accepted.
    root_cause:      The tax rate validation logic lacks an upper bound check,
                     allowing rates above 100% (e.g. 150).
    plan:
      - Add a validation rule to ensure tax_rate is less than or equal to 100
        (and non-negative).
    target_files:
      - outputs/agent_src/code/CALC_TAX/calc.py
    ```
    
    It has reconstructed the withheld rule from one failing assertion.
    
    ### 5. The PATCH
    
    The FIX names the files; a second call produces a unified diff of only those
    files, applied all or nothing:
    
    ```diff
    --- a/outputs/agent_src/code/CALC_TAX/calc.py
    +++ b/outputs/agent_src/code/CALC_TAX/calc.py
    @@ -52,3 +52,3 @@
             tax_rate = _parse_decimal(rate_str)
    -        if tax_rate is None or tax_rate < 0:
    +        if tax_rate is None or tax_rate < 0 or tax_rate > 100:
                 rej = dict(row)
    ```
    
    ### 6. The stage clears, and the earlier ones are re-checked
    
    A run works through three stages, and the console names each one by number.
    **Stage 1 is the integration tests** and **stage 2 the system tests**, both
    written from the specification before any code exists; stage 2 drives the
    single end-to-end entry point, `run_calc`, rather than the individual
    functions. **Stage 3 is the unit tests**, generated last, once there is code
    whose internal helpers they can import and call. Every stage has to pass before
    the next one is generated.
    
    ```
    [stage 1/3] [iteration 3] 9/9 passed
    [stage 2/3] [iteration 1] 5/5 passed
    [stage 1/3] [iteration 1-regcheck] 9/9 passed
    ```
    
    Then the unit stage is generated, from the code that now exists, and the same
    loop runs again. That run converged in 38 seconds and 11 model calls, for about
    \$0.006.
    
    ### What this example shows
    
    The coding agent eventually wrote `tax_rate > 100`, and only because a test
    told it 150 was wrong, not because a criterion told it 100 was the limit. Had
    it been handed the criteria, it would have written that bound on the first
    attempt, the test would have passed immediately, and the green would have meant
    only that one model agreed with itself twice.
    
    Note that in real scenarios it is often impractical to hand over all the
    criteria in advance, for example when many edge cases exist.
    
    The failure is the evidence. It is what a test is capable of producing, unlike
    a shared context, which cannot do that.
    
    ---
    
    ## Where to go next
    
    That is the whole idea, demonstrated once end to end. Two questions follow
    naturally, and each has its own part.
    
    **Does it actually work, and how often?** Three sweeps, 967 runs, what
    reproduced and what did not, including the results this project measured and
    then withdrew. [Part 2: How well it works](design_2_performance.md).
    
    **How is it built, and how do I use it on my own specification?** The five
    agents, the FIX and PATCH separation, the stage ordering, and the command to
    run for whichever parts of a task file you already have.
    [Part 3: How it is built](design_3_mechanism.md).
    
  • design_2_performance.md 18.6 KB
    <!--
    One of three. See MAINTAIN.md: these three files, README.md, docs/index.html
    and the investor deck move together. A number changed here has to change in
    all of them.
    -->
    
    # How well it works: three sweeps, and the results we withdrew
    
    **Part 2 of three.** [Part 1](design_1_case_study.md) makes the case that the
    agent writing the code should never see the acceptance criteria, and shows one
    repair end to end. This part asks the harder question: does it work, how often,
    and how much of that is measurable?
    
    Everything below is from runs on `gemini-3.5-flash-lite`, a small cheap model
    chosen so that repeated sweeps were affordable. Treat the figures as a floor
    rather than a ceiling. Where a measurement did not survive scrutiny, it is
    reported as withdrawn rather than removed.
    
    
    ---
    
    **This is part 2 of three.**
    
    | | | |
    |---|---|---|
    | [1. The case](design_1_case_study.md) | **2. How well it works** (you are here) | [3. How it is built](design_3_mechanism.md) |
    
    ## How well it works
    
    Three claims, in descending order of how much weight they can carry.
    
    ### One you can check yourself, with no statistics at all
    
    **The coding agent never receives the acceptance criteria.** Not "is instructed
    not to look at them", and not a convention someone has to remember: the criteria
    are removed before the FIX and PATCH prompts are assembled, and a test in the
    repository fails the build if any call site lets one through.
    
    ```bash
    qikly --explain CALC_TAX
    ```
    
    It prints the task file twice, once as each agent receives it, and the
    difference between them: all eleven acceptance criteria go to one side and are
    cut before the other side is handed the file. No API key, no model call, about
    a second, and it works from a plain `pip install`.
    
    That shows the criteria being removed once, for one task. The test suite
    checks something stronger: that no call site anywhere in the codebase can pass
    a criterion to the coding agent, so the removal cannot be undone by a future
    change. If you have cloned the repository rather than installed the package,
    you can run it:
    
    ```bash
    python -m pytest tests/test_withholding.py -v
    ```
    
    That takes a few seconds and needs no API key, no sample size and no
    confidence interval. It is the strongest claim here precisely because it is not
    a measurement: it is a property of the code, and it cannot rot.
    
    ### It converges, and the rate reproduces
    
    Measured on `gemini-3.5-flash-lite`, a small cheap model chosen to make repeated
    sweeps affordable, so treat these as a floor rather than a ceiling.
    
    **Roughly 8 runs in 10 finish with the code passing every integration and system
    test. Roughly 6 in 10 pass everything including unit tests.**
    
    The reason those are round numbers is that they were measured three times, over
    967 runs, and then re-measured on a changed system.
    
    | Sweep | Runs | Integration + system | Integration, system and unit |
    |---|---|---|---|
    | 14 August | 427 | 80% | 59% |
    | 30 August, same tasks and settings | 140 | 87% | 67% |
    | 31 August, after correcting the benchmark | 400 | 83% | 64% |
    | 23 and 24 September, on 0.5.1 | 390 | 80% | 60% |
    
    **The last row is not a further draw from the same urn, and is deliberately
    not pooled with the three above it.** 0.5.1 hardened the process
    itself. Six changes, every one of them either narrowing what the coding agent
    receives or removing a way a run could stall on something no implementation
    could satisfy:
    
    - The criteria strip stopped being a regular expression and started asking the
      YAML parser where the section ends, closing two leaks a regex could not see:
      a comment written between two criteria, and a criterion wrapped onto a line
      beginning at column 0. Either handed the coding agent real criteria.
    - `ADAS_HEADWAY`'s header comment stopped naming which way its own withheld
      boundary falls.
    - The `# Criteria: N` markers stopped reaching the coding agent.
    - `CALC_TAX` shed half its commentary, which the agents had been reading on
      every call.
    - Integration tests are told to name the seam they cross.
    - Test generation is told to call the function under test once per scenario,
      after a run stalled for eleven iterations against `transform(transform(rows))`,
      a test no implementation can satisfy.
    
    **None of it makes the task easier, and two items make `ADAS_HEADWAY` strictly
    harder**, since that task had been telling the coding agent the answer it was
    about to be tested on. So the right reading of these rows is not "the rate held
    while we changed things" but "the rate held while the bar was tightened".
    
    Pooling them with the 967 would produce one number describing two systems,
    which is the single reason every positive result this project has withdrawn was
    withdrawn. Reported separately they say something a larger sample could not:
    80% and 60% land inside the spread of the three earlier sweeps, 80 to 87 and 59
    to 67, so none of this moved convergence by anything a sample this size can
    detect.
    
    The 0.5.1 row is three sweeps of 130 runs pooled, taken over two days on the
    same tasks and settings. Pooled, it converges **80% [76-84] and 60% [55-65]**,
    which is the tightest estimate on this page and the same 8 in 10 and 6 in 10
    the August sweeps found before any of the hardening above existed.
    
    The third sweep is the interesting one, and the reason is what happened between
    the second and the third. Asking a separate agent which criteria no input row
    could trigger turned up unreachable rules in eight of the ten tasks, from one
    in `CALC_CALENDAR` to seven of thirteen in `MERGE_CONTACTS`. A criterion
    nothing can reach produces a test that passes whatever the code does, so part
    of every bar was not lower, it was absent. Thirty-one rows were added, and the
    whole sweep was taken again.
    
    The prediction was that convergence would fall, because the bar had genuinely
    become enforceable. **It did not move.** 64% sits between the two earlier
    figures and inside both intervals. Either the added rows exercise behaviour the
    implementations were already getting right, or the difference is smaller than
    this measurement can see, and separating those needs fault injection rather
    than convergence. That experiment is not run.
    
    Three independent sweeps agreeing is worth more than any one number's decimal
    places, so the decimal places are not quoted.
    
    One thing about a rate like this is worth internalising before you run anything:
    **a single run is an artifact, not a rate.** The same task with the same seed
    converges on some runs and exhausts its budget on others. To make a claim about
    how often anything converges, repeat the sweep and read the interval.
    
    ### Nearly the whole gap between those two numbers is the unit stage
    
    This is the most stable finding here and the most useful one, because it tells
    you where runs actually fail. It held at about twenty points across both sweeps.
    
    The reason is structural rather than mysterious. The unit suite is the largest,
    so there is more to satisfy. It runs last, when the earlier stages already pass
    and regression re-checks force them to keep passing, so a fix has the least room
    to move. And it is the one stage whose tests are written with sight of the
    implementation, so it can assert on incidental internal structure rather than on
    required behaviour.
    
    That last point deserves emphasis: **the only stage where the blindness is
    broken is also the hardest stage.** Whether that is cause or coincidence is
    testable, and untested.
    
    ### What explains the variation
    
    The variation in the convergence rates is not due to size because every task has 6 to 7 spec requirements, 10 to 13 criteria, 3 interface functions and 2 input files.
    
    Three things drive the difference:
    
    1. *Carried state of a task.* Whether row N's correctness depends on the rows before it. The three `CALC_*` tasks are row-independent and average 78%. The three `MERGE_*` tasks all carry state and average 46%.
    2. *Cross-file coupling.* Whether two input files can be concatenated or must genuinely be interleaved. `MERGE_STOCK` must merge-sort by timestamp before computing anything; concatenation produces plausible, wrong answers.
    3. *Conflict with the model's priors.* Where a criterion asks for a convention the LLM would otherwise resolve differently.
    
    ### When a run stalls
    
    A run that exhausts its budget exits non-zero, names the blocking tests, and
    keeps the complete record. **There is no path by which it reports success on
    code its own tests reject**, and in several hundred measured runs there has
    never been such a case.
    
    Stalls are usually near misses rather than wreckage: most of the blocking stage
    is already passing, and a large share fail exactly one test. They also cluster.
    The exact failing test names almost never repeat, but the *subjects* do, so
    what you get is a short list of nameable problems rather than a diffuse failure
    rate.
    
    **Three of the four causes are fixed by editing text, not by buying a bigger
    model.** They live in the specification or the bar rather than in the coding
    agent, which means most stalls are within your control and cheap to resolve.
    The signature of each, and what to do about it, is in
    [TROUBLESHOOTING.md](TROUBLESHOOTING.md).
    
    ## Is the bar any good?
    
    This is the question that matters, and the tool is built to answer it with
    evidence rather than assertion.
    
    A programme of work is running on exactly this: cross-testing implementations
    against other runs' suites, planting faults in code a suite has already accepted
    to see whether it notices, and asking whether automatically refined criteria
    catch more than a first draft. Those results will be published once substantial
    user data has accumulated and been carefully analysed, because a figure earned
    across many real specifications is worth far more than one earned across ten
    example tasks.
    
    One finding from that work is already solid enough to act on today, because it
    is a mechanism rather than a rate.
    
    **A suite that never tests the boundary cannot detect an error at the boundary**,
    no matter how many other cases it covers. Change a `>` to a `>=` in code a
    generated suite has already approved:
    
    ```python
    # the code the suite accepted
    if quantity > 100:
        reject(row, "quantity too large")
    
    # the same code with one character changed
    if quantity >= 100:
        reject(row, "quantity too large")
    ```
    
    Those two versions disagree about exactly one input, `quantity == 100`, and
    agree about every other input in the universe. Suites generated for this task
    tested 5, 50 and 250:
    
    | Test input | Original | Mutated | Suite can tell? |
    |---|---|---|---|
    | 5 | accept | accept | no |
    | 50 | accept | accept | no |
    | 250 | reject | reject | no |
    | **100** | **accept** | **reject** | **yes, but no test used 100** |
    
    Adding more tests at 5, 50 and 250 would not help. Only a test at exactly 100
    would. What this says is that the generated suites were testing that the logic
    works, not that it stops in the right place, which is the difference between a
    test written to demonstrate behaviour and a test written to catch a mistake.
    
    **And it tells you exactly what to fix, in your own file, today.** A criterion
    phrased "reject quantities above 100" invites a test at 250. A criterion phrased
    "100 is accepted and 101 is rejected" forces a test at the boundary. The blind
    spot is in how the criteria are worded, which you control, rather than in the
    model, which you do not.
    
    ## Where the bar comes from
    
    Everything above assumes a human wrote the criteria. The harder question is whether the tool can write them itself, because that is the expensive half of QA.
    
    The procedure goes as follows: draft initial criteria from spec requirements alone; converge a code implementation against that draft in an isolated workspace; then give a reviewing agent the specification, the current criteria, *and the finished implementation*, and ask what the code does that none of the criteria constrain?
    
    The reviewer is looking for a specific thing: the gap between "passes what is currently tested" and "actually correct". These are places where the validation is looser than it appears, where one code path applies a normalization and another does not, or where a boundary condition is handled correctly only by accident rather than by design. Each such finding becomes a new criterion, tagged with an appropriate category.
    
    Before trusting a drafted bar on a task where you have not written one, there is a way to find out what it would have cost you on a task where you have:
    
    ```bash
    qikly --compare-criteria <MY_TASK>
    ```
    
    It drafts criteria from that task's requirements alone, then reports them against the ones you wrote, treating yours as the ground truth. Output is covered, partial and missed, with every gap quoted in full and the drafted criteria that match nothing of yours listed separately, since those are either noise or a rule you know and never wrote down. The matching is a model's reading rather than a measurement, and says so: two criteria can mean the same thing in different words. It also says nothing about test quality, because two bars can describe the same rule and produce suites that catch different faults, which is the whole reason this project measures a bar by fault detection rather than by its text.
    
    *Currently, the tool only makes additions.* It appends criteria and never modifies or removes one.
    
    **This constrains the machine, not you.** The criteria live in your YAML file. You can rewrite one, delete one, or throw the whole set out, at any time, exactly as you would edit any other file you own. Nothing is locked. What the *system* cannot do is remove a criterion by itself.
    
    The reason is narrow and worth stating plainly. The tool's success condition is "all criteria satisfied". If the tool could also edit the criteria, then the cheapest way to satisfy a failing criterion is to delete it, and every run would converge by definition. Success would stop meaning anything. Append-only is what keeps a passing run evidence of something rather than evidence of nothing.
    
    So the answer to "we let it add criteria that may be wrong and can never be fixed?" is no, on two counts. A proposed criterion is never adopted automatically: a human reads it and accepts it, because a drafted bar is a proposal, not ground truth. And once adopted, it is yours to correct like any other line in the file.
    
    The extension not yet built is letting the tool *propose* a removal, with the evidence for it, for a human to accept or reject through a normal gate such as a pull request, keeping the superseded criterion in the history rather than erasing it. The line that matters is between proposing and enacting, not between adding and removing.
    
    ### Does refining the criteria make the suite catch more?
    
    **An open question, and an active one.** Refinement reliably *grows* the bar:
    54% more criteria, and `CALC_TAX` goes from 12 to 26 across three rounds. Growth
    and improvement are different things, so the question worth answering is whether
    the grown bar catches more real defects.
    
    Measuring that is a genuinely interesting experimental problem, and most of the
    work so far has gone into the apparatus rather than the answer. Comparing two
    bars means holding everything else constant: the code, the fault set, the suite
    size, and the substrate each suite is judged on. Each of those took a round to
    get right, and the current design controls all four.
    
    The remaining piece is fixture coverage. In a fault-injection comparison a
    planted fault can only be caught if some input row reaches the behaviour it
    changes, which is exactly what the [fixture proposal agent](design_3_mechanism.md#the-five-agents)
    was built to close. Results will follow once substantial user data has
    accumulated and been carefully analysed.
    
    Refinement is shipped and usable today. A quantified claim about how much it
    sharpens the bar is what the work above will produce.
    
    ## Should the specification iterate too?
    
    No, and that is a deliberate design decision worth explaining.
    
    The spec requirements are the fixed point everything is judged against. A system permitted to rewrite its own goal can always satisfy a failing test by weakening what was asked. That is the same failure mode append-only prevents on the criteria side, and automating both ends would remove the last thing making convergence meaningful.
    
    The legitimate version is different and worth building: when the review finds behaviour no criterion constrains, it is often because *the specification was ambiguous there*. Reporting "your spec does not say what happens when X, and the implementation chose Y" is a proposal to a human, not a self-edit. For a real team that may be the most valuable thing the tool produces.
    
    ## In summary
    
    - **An agent that writes its own tests is grading its own homework.** At temperature zero with identical inputs it is provably vacuous: the same model that wrote the bug writes the test that blesses it.
    - **The fix is structural, not procedural.** The coding agent never receives the acceptance criteria. Not "is told not to look", but never has them in its context.
    - **You can verify that in one command.** `qikly --explain <MY_TASK>` prints what each agent is given and the difference between them. No API key, no model call, no sample size. From a clone, `pytest tests/test_withholding.py` additionally proves no call site can leak one, including a call site added next year. Both are properties of the code rather than benchmark results, so neither can go stale.
    - **It converges, and the rate reproduces.** Roughly 8 runs in 10 pass every integration and system test, roughly 6 in 10 pass everything including unit tests, measured three times on a small cheap model over 967 runs, the third after correcting the benchmark itself, then re-measured on 0.5.1 over 390 runs after the process was hardened, and reproduced at 80% and 60%. Every run that does not converge exits non-zero and names its blocking tests, and none has ever reported success on code its own tests rejected.
    - **Nearly the whole gap between those two figures is the unit stage**, which is also the only stage whose tests are written with sight of the code.
    - **How you word a criterion decides how sharp the test is, and that is yours to control.** A suite tests the boundary when the criterion names the boundary: "100 is accepted and 101 is rejected" produces the test that "reject quantities above 100" leaves to chance. This is the highest-leverage thing you can do in your own file.
    - *Whether automatic criteria refinement produces a measurably sharper bar is the open question, and an active one.* It grows the bar reliably, and the experiment design to quantify the rest is built.
    
    The measurement programme continues alongside the tool, and its results will be published once substantial user data has accumulated and been carefully analysed. Figures earned across many real specifications, from many people, are worth considerably more than figures from ten example tasks. What is published here has already survived an independent re-measurement; the rest will meet the same bar before it joins it.
    
  • design_3_mechanism.md 20.1 KB
    <!--
    One of three. See MAINTAIN.md: these three files, README.md, docs/index.html
    and the investor deck move together. A number changed here has to change in
    all of them.
    -->
    
    # How it is built, and how to use it
    
    **Part 3 of three.** [Part 1](design_1_case_study.md) makes the case and shows
    one repair end to end. [Part 2](design_2_performance.md) reports what was
    measured. This part is the architecture: what separates it from the
    alternatives, the five agents, and the command to run for whichever parts of a
    task file you already have.
    
    
    ---
    
    **This is part 3 of three.**
    
    | | | |
    |---|---|---|
    | [1. The case](design_1_case_study.md) | [2. How well it works](design_2_performance.md) | **3. How it is built** (you are here) |
    
    ## What makes this different
    
    **The tests come from the standard, not from the code.** This is the one that matters most. Every other AI test generator in this space derives its tests from an implementation that already exists: it reads the code to decide what to assert, which produces an excellent regression harness that locks in current behaviour. qikly writes the integration and system suites from the acceptance criteria **before any implementation exists**, so there is nothing for the standard to be shaped by. That is why a green suite here carries information: the tests describe what the code should do, not what it already does.
    
    **The withholding is a mechanism you can watch, not a promise.** The criteria are cut out of the file in code, before the FIX and PATCH prompts are assembled, so no representation of them exists in the coding agent's context. `qikly --explain <MY_TASK>` prints what test generation receives, what the coding agent receives, and the difference between them, from a plain install with no API key. A test fails the build if any call site lets a criterion through, including one added next year by someone who has never read this page.
    
    **What you get is an executable suite you keep.** Real pytest files, plus JUnit XML for whatever tracks tests where you work. Read them, run them, put them in continuous integration, and when one fails in six months it fails for a reason you can inspect and argue with. A suite is a durable asset in a way a model's verdict is not: a verdict cannot be re-run against tomorrow's commit.
    
    **It helps with writing the standard, not just checking against it.** `--scaffold` turns code you already have into a task file, `--criteria-from` reads criteria straight out of the ticket that already holds them, `--generate-criteria` drafts a first bar from requirements alone, `--check-criteria` looks for two statements in the specification that no implementation could satisfy at once, criterion against criterion included, and the refinement loop reviews converged code to propose criteria the first draft missed. Whether that last step produces a measurably *sharper* bar is an open question this project is still working on.
    
    **Every run is reproducible, and the whole trail is kept.** A run records
    the provider, the model, the settings and the version that produced it, next to
    every failing test, every FIX with its stated root cause, and every PATCH as a
    diff. You can read back exactly why a line of code exists: which assertion
    forced it, what the model concluded, and what it changed. That record is what
    makes a convergence rate a measurement rather than an anecdote, it is what let
    three of this project's own positive results be withdrawn on inspection, and it
    survives a run that never converges. Nothing here asks you to take a number on
    faith.
    
    ### "Why not just use two different models?"
    
    It is the first thing most people ask, and it does help a little. It does not reach the underlying issue, though, because both models still read the same criteria and so both still write to them. Where the criteria say "reject malformed rows" and never define malformed, two models resolve that ambiguity from the same sentence, and the implementation is still built around the resolution the tests will check for. Two models is also a habit rather than a mechanism: nothing checks they stayed different, and a settings change a year from now undoes it with no test to notice.
    
    Withholding removes the channel rather than the coincidence, so it holds whichever model is writing. The two compose nicely, incidentally, since qikly picks a provider and model per agent role: you can withhold *and* use two models.
    
    ### "Why not just add a reviewer agent?"
    
    The newer form of the same question, and the one to answer carefully, because independent verification steps are now shipping in mainstream coding agents: a second agent, usually from a different model family, reviews what the first produced.
    
    It helps, and it stops in the same place. A reviewer handed the same specification has read the same acceptance criteria and resolves the same ambiguity the same way. It catches what is visible from that context: an inconsistency, a requirement plainly skipped, a bug that looks like a bug. It cannot catch the case this design exists for, where the code and the standard agree because both came from one reading of a line that admitted two. Relative to the shared interpretation nobody in that loop is wrong, which is precisely why they all agree.
    
    The problem was never that nothing was checking. It is that everything checking had already read the answer key. Model diversity varies who is looking; withholding varies what they were shown, which is the only one of the two that changes what can be found. And it is enforced by a test rather than by an arrangement someone has to remember to keep.
    
    One more difference, and it is the one that outlasts the run: a reviewer emits a verdict, and this emits a pytest suite that is still there in six months, running against tomorrow's commit.
    
    ## The mechanism
    
    A task file has **three** parts, and the cut runs between the third and the
    first two:
    
    1. **`requirements`** is the vague part, and it holds the decisions: what a real specification says before anyone sharpens it, including every choice that could have gone another way, such as a threshold, a unit or an exemption.
    2. **`interface`** is the contract as a description rather than code, the function signatures and where the module will live. Both agents read it.
    3. **`acceptance_criteria`** is the sharp part, and it holds the consequences: specific, objectively checkable statements of what must be true if those decisions were implemented correctly, including the boundary values and edge cases a careless reading gets wrong. A decision does not belong here: withheld from the one agent that needed it, it produces a stuck loop rather than a harder test.
    
    The test generation agent receives all three. The coding agent receives the
    first two, **with the acceptance criteria removed in code before the prompt is
    built**. It is not instructed to ignore them. It cannot be persuaded, prompted,
    or induced into seeing them, because no representation of them exists in its
    context.
    
    Everything else follows from that asymmetry. A run works through three stages,
    `integration` then `system` then `unit`:
    
    1. **Integration and system tests are generated first**, from the spec alone,
       before any code exists. They cannot see an implementation because there is
       not one yet.
    2. **The coding agent writes the implementation**, from the spec minus the
       criteria.
    3. **pytest runs.** On failure the model produces a **FIX** (failure summary,
       root cause, plan, and the files it intends to touch) and then a **PATCH**
       (a unified diff of only those files), applied all or nothing. Repeat until
       the stage passes or the budget is spent.
    4. **Unit tests are generated last**, once real code exists for them to name.
       This is the only stage allowed to see the implementation.
    5. **Clearing a stage re-runs the earlier ones**, so a later fix cannot
       silently break something that already passed.
    
    Two things are worth noticing about step 3. The reasoning and the diff are separate LLM calls, which makes the reasoning independently auditable and restricts the diff's context to the files the reasoning identified. And the FIX is generated from pytest's output alone: the failing test's name, its source down to the failing line, its docstring and the assertion, with no acceptance criteria. The same feedback a developer sees in their terminal, and no more.
    
    There is no separate "write the initial implementation" step anywhere. The first run happens against an empty source tree, fails on a collection error, and that failure feeds the same FIX and PATCH cycle as every later repair. The first line of code and the hundredth are produced by one mechanism.
    
    ### The five agents
    
    Everything above is done by five agents, each with its own definition file in
    `inputs_public/agent_defs/` and its own entry under `agents:` in
    `settings.yaml`, so any of them can be pointed at a different model without
    touching the others. Two of them exist because a single agent doing both jobs
    is the failure this project is about.
    
    | Agent | Settings key | Reads | Never reads | Produces |
    |---|---|---|---|---|
    | **Coding agent** | `code` | requirements, interface, and pytest's output for whatever test just failed | the acceptance criteria | a FIX, the reasoning and the files it intends to touch, then a PATCH, a unified diff of only those files |
    | **Test-writing agent** | `test`, or `test_integration` / `test_system` / `test_unit` | requirements, interface, and every acceptance criterion in full | the implementation, except at the unit stage, which is written last and exists to name real functions | one pytest file per stage |
    | **Criteria drafter** | `criteria` | the requirements alone | any implementation, and any existing criteria | a first draft of the bar, only when `--generate-criteria` asks for one |
    | **Criteria reviewer** | `review` | the specification, the current criteria, and a finished implementation | nothing withheld: this is the one agent shown everything at once | proposed additional criteria, each tagged with a category, for a human to accept or reject |
    | **Fixture proposer** | `fixtures` | the requirements, the criteria, and the fixture files as they stand | the implementation, and the criteria it is not asked about | for each criterion nothing currently reaches, one row that would, written to a proposal file for a human to accept |
    
    The coding agent runs far more often than the others, once per repair attempt,
    which is why it is the one where a cheaper model pays for itself. The reviewer
    is the one asked to find what nobody wrote down, which makes it the likeliest
    to be worth a stronger model. Both are one line in `settings.yaml`.
    
    The fixture proposer is the newest and the only one that suggests changing the
    *inputs* rather than the code or the bar, which is why it is the only one whose
    output never lands anywhere automatically. A criterion nothing can trigger
    produces a test that passes whatever the code does, and in this project's own
    measurements roughly two thirds of deliberately planted faults were missed by
    every suite for that reason: the bar was unmeasurable rather than wrong. But a
    row is harder to review than a sentence, because it is only right or wrong
    relative to the criterion it was proposed for, and a fixture set that grows in
    whatever direction a model finds interesting stops resembling the data you
    actually process. So proposals go to a file, capped per round, each row printed
    under the criterion it exists to reach, and the file reports what share of your
    data a machine has written so the drift is visible in aggregate rather than one
    plausible row at a time.
    
    ## Using it
    
    A word first on how this was built, because it is not incidental to the subject. I architected the tool's objectives and its orchestration, and I guided the research: which experiments to run, which results to believe, and which of my own claims to discard when the numbers did not support them. **The code itself was largely written and tested by Claude Opus 5.0, across many iterations of review, correction and rework.**
    
    That feels worth stating plainly in a post about not letting one model mark its own homework. The separation this tool enforces is the same separation I relied on while building it: the measurements decided what was true, not the author of the code, and several of the conclusions below are ones I did not want.
    
    ```bash
    pip install qikly
    export GEMINI_API_KEY=...
    qikly --demo
    ```
    
    The demo runs one task end to end in an output directory and prints what it built. That takes about thirty seconds. It writes nothing outside that directory.
    
    ### What you already have decides how you use it
    
    A task file has the three parts described under [The mechanism](#the-mechanism),
    and which of them you already have decides both what qikly does for you and
    which command you run. The short version: `qikly --explain <MY_TASK>` to see the
    split for yourself, `qikly --demo` to watch a whole run, `qikly --init` to start
    a task from nothing, and `qikly --scaffold <FILE>.py` to start one from code you
    already have. The full table, with the exact command for each starting point, is
    in
    [QUICK_START_ON_YOUR_OWN_DATA.md](https://github.com/gal-a/qikly/blob/main/docs/TASK_FILE_REFERENCE.md#which-command-depends-on-which-parts-you-already-have).
    The four that matter most, most valuable first:
    
    1. **You have #2 and code somebody else wrote, and you want that code
       verified.** Something else produced the implementation. Point `--scaffold`
       at it and it reads the real signatures into the interface for you, leaving
       the requirements and the criteria yours to write. The suite is then written
       from criteria the implementation's author never saw. A suite generated from
       the same context as the code is a model agreeing with itself, and agreement
       is not evidence. This is the use no other tool in this space covers.
    
       One honest caveat comes with it, and the run states it rather than leaving
       you to find it. qikly did not write that implementation, so it cannot know
       what its author read: the separation is evidenced here rather than enforced.
       The run asks git whether the task file was last changed before the
       implementation's first commit and prints the answer next to what that answer
       does not show. A failing test is a real finding either way, since it was
       written from the criteria without reading the code. It is the passes the
       evidence qualifies.
    
    2. **You have #1, #2 and #3, and no code.** A specification exists and an
       implementation does not. You get a first implementation plus the suite that
       justifies it, and nothing is drafted on your behalf. Everything lands in
       `outputs/` for review; nothing is written to your source tree.
    
    3. **You have #1 and #2, and want a first draft of the bar.**
       `--generate-criteria` drafts #3 from the requirements alone, or
       `--criteria-from` lifts it out of the ticket where you already wrote it.
       Then the run proceeds as above, and the draft is yours to correct.
    
    4. **You have everything and want it run unattended.** Non-zero exit on any
       non-convergence, so a scheduler or CI job can run many specifications and
       keep the reports as artifacts. There is a GitHub Action, and JUnit XML for
       whatever tracks tests where you work.
    
    ## Now it's your turn to test it on your specification
    
    The engine is open source under Apache 2.0, because the central claim is one nobody should take on faith. The whole argument rests on the coding agent genuinely never seeing the bar, and that is something you can test and verify rather than just accept.
    
    Read the prompt builders in `src/qikly/agent_api/prompts/`, read the function that strips `acceptance_criteria` out of the specification before the prompt is built, and watch a run do it.
    
    That verification is the point of publishing the engine at all.
    
    *Try it in thirty seconds:*
    
    ```bash
    pip install qikly
    export GEMINI_API_KEY=...
    qikly --demo
    ```
    
    Then do the thing that actually tests the idea: write one specification of your own, split it into `requirements` and `acceptance_criteria`, and read the criteria the tool derives against the ones you would have written by hand. If it finds a gap you missed, that is the argument. If it does not, I want to know which specification broke it.
    
    Issues and results, welcome and wanted:
    [github.com/gal-a/qikly/issues](https://github.com/gal-a/qikly/issues)
    
    ---
    
    ## Appendix: reference
    
    ### Example tasks
    
    Thirteen tasks ship with the tool, in five domain families.
    
    | Family | Task | What it does |
    |---|---|---|
    | `CALC_*` arithmetic and precision | `CALC_TAX` | Per-line order tax; stresses rounding and currency precision |
    | | `CALC_DISCOUNT` | Discount amounts; validation versus computation edge cases |
    | | `CALC_CALENDAR` | Date-range charge from a monthly rate; calendar arithmetic, leap years |
    | `ETL_*` string and format validation | `ETL_EMAIL` | Validates and normalises contact records with an email field |
    | | `ETL_ADDRESS` | Postal addresses; deduplicates across two files |
    | | `ETL_NAME_SPLIT` | Splits a full-name field into first, last and suffix |
    | `MERGE_*` combining two sources | `MERGE_SALES` | Two sales exports with overlapping ranges; cross-file dedup |
    | | `MERGE_STOCK` | Inventory transactions into per-SKU stock; cross-file ordering |
    | | `MERGE_CONTACTS` | Contact records from two systems; conflict resolution by recency |
    | `AGG_*` event-stream aggregation | `AGG_RUNLOG` | Summarises append-only JSONL run logs |
    | `ADAS_*` vehicle sensor data | `ADAS_HEADWAY` | Following-distance warnings from forward-radar samples; inclusive limits and the two-second boundary |
    | | `ADAS_TTC` | Time to collision from radar tracks; a closing speed that can be zero or negative, and a three-second boundary |
    | | `ADAS_SPEED_LIMIT` | Speed-limit compliance; a unit conversion, an enforcement tolerance, and an enumerated set of valid limits |
    
    Every task ships with `requirements`, an `interface`, and a full set of
    `acceptance_criteria`, so each one is a worked example of the task format
    as well as something to run.
    
    ### Supplying your own code or tests
    
    The loop takes three inputs, and each can be yours or generated, independently
    and in any combination. `seed.implementation` points the tool at code you
    already have, so the run skips generating a first implementation and goes
    straight to testing and repairing yours. `seed.tests` keeps a suite you already
    trust, so the loop repairs the code against your tests rather than its own.
    
    That last combination is the one worth naming, because it is the strongest use
    of the tool: **your suite, its code.** The bar was written by a person and the
    implementation has to satisfy it without ever having read it.
    
    Setup, the exact YAML, and what `--scaffold` writes are in
    [README.md](https://github.com/gal-a/qikly#bringing-acceptance-criteria-you-have-already-written).
    
    ### Measuring rather than producing
    
    ```bash
    python -m qikly.orchestrator.run_all --repeat 10
    ```
    
    Repeats the whole sweep N times and writes one aggregate report to `outputs/reports/`. Concretely, that report contains:
    
    | Output | What it is |
    |---|---|
    | Convergence rate per task | Converged runs over total, with a 95% confidence interval, so a task at 6/10 is reported as a range rather than as "60%" |
    | Per-stage pass rates | How often integration, system and unit each cleared, which is what identifies the blocking stage |
    | Right-censored runs | Budget-exhausted runs counted separately from failures, because "did not finish in 10 attempts" is not the same fact as "cannot be done" |
    | Stall signatures | Non-converging runs grouped by the tests they persistently failed, which is what turns a pile of stalls into a short list of recurring subjects |
    | Cycle counts | FIX-to-PATCH cycles per run, so cost per converged task is visible |
    
    The point is that it emits **intervals and groupings, not a single headline number.** A rate without an interval invites exactly the mistake described below.
    
    Use `--repeat` when making a claim. Use a single run when you want the output. That distinction has already caught us out: a criteria change we were confident about looked like a fix, and ten repeated runs showed 5 out of 10 against 6 out of 10 before, no improvement at all and well inside the noise.
    
    ### Watching a run live
    
    ```bash
    python -m qikly.orchestrator.live_view
    ```
    
    Tails the current run's transaction log and renders it as it happens: which stage, which iteration, which tests failed, and what the model did about it.
    
  • EXISTING_CODE_AND_HELPERS.md 4.6 KB
    # Pointing qikly at code you already have
    
    What qikly can see, what it can change, and what happens when the defect turns
    out to be somewhere it may not write. Split out of [FAQ.md](FAQ.md) because the
    answer outgrew a FAQ entry.
    
    The YAML for everything below is in
    [TASK_FILE_REFERENCE.md](TASK_FILE_REFERENCE.md); this page is the reasoning.
    
    ## qikly does not go looking for your code
    
    There is no discovery step. qikly does not scan your repository and work out
    what is available, so it will not find your driver or parser classes on its
    own. What the test-writing agent targets is what the task file's `interface`
    block declares, plus the requirements and the acceptance criteria. You describe
    the surface.
    
    Two things make that less manual than it sounds:
    
    - **`qikly --scaffold path/to/module.py`** reads the real signatures out of a
      file you point it at and writes the `interface` block for you.
    - **`seed.implementation`** points the run at code you already have, and
      **`seed.tests`** keeps a suite you already trust, per stage, so the loop
      repairs code against your tests rather than against generated ones.
    
    ## Can the tests import my other classes once they exist?
    
    Yes, with one condition. Each stage runs as `python -m pytest` from your
    project root, which puts the project root on `sys.path`. So a package sitting
    at the project root, or one installed into the same virtualenv, imports
    normally from both the generated tests and the generated implementation.
    
    A package in a subdirectory that is not on the path does not. The run stops
    with `no test ran: the module could not be imported`, and no amount of
    iterating fixes it because the problem is the path rather than the code. Put
    the directory on `PYTHONPATH`, or `pip install -e .` your own package.
    
    ## What it can read, and what it can change, are decided by the seed
    
    `seed.implementation` takes a directory as well as a file. A directory is
    treated as a package: it is copied in under its own name and becomes
    importable by that name, so `from mypkg.money import to_cents` resolves as it
    does in your own tree.
    
    | | Inside the seed | Outside it |
    |---|---|---|
    | Imports at runtime | Yes | Yes, if on `sys.path` as above |
    | The coding agent **sees the source** | Yes | No. It reasons about them from their behaviour |
    | The coding agent **may write** to them | Yes | No |
    
    **The seed is the boundary you choose, and that is deliberate.** qikly does not
    follow your imports and decide for itself which of your files an agent may
    rewrite: the transitive closure of a real package has no natural edge, and "it
    rewrote a shared module I never named" is a worse outcome than naming a folder.
    So put inside the seed what you want worked on, and leave a vendor library, or
    a large shared module you do not want touched, outside it.
    
    ## When the defect is outside the seed
    
    A test that fails because your `utils.py` is wrong does fail, correctly, which
    is the point. But the agent cannot patch a file outside the seed, so left to
    itself it would either stop without converging or change a module it does own
    to work around a defect that is somewhere else.
    
    From 0.5.5 a run says which of those happened, rather than leaving you to infer
    it from a patch history:
    
    - **At the start**, if the task's module imports anything local it may not
      write, one line names those files. It is information, not a warning: they
      import and run normally, and most of the time they are fine.
    - **When a run stops without converging**, the message names the file instead
      of ending at "exceeded 10 attempts".
    - **When a run passes but a failure along the way was traced into one of those
      files**, it says so, because a green suite reached that way may be green
      because the code was bent around a defect that is still there. **This is the
      one worth reading twice.**
    
    **Nothing here stops or slows a run.** A project with shared helpers is normal,
    and runs in one converge exactly as before. The guard only speaks up.
    
    If it was not what you wanted, the fix is usually one line: seed the package
    that holds `utils.py` instead of the single module.
    
    ## `--score-code` has no such limit
    
    It neither writes code nor calls a model. Point it at a package and it plants
    faults throughout it, helpers included, and tells you which ones your suite
    noticed. No task file, no seed, nothing of yours modified.
    
    ## One thing worth knowing before you start
    
    The loop reruns pytest on every iteration, so it suits the deterministic layer
    best. Protocol parsing is a good fit: bytes in, structured records out, driven
    from recorded captures. Code talking to live hardware is better left behind the
    test doubles you already have, because a suite that needs a rig attached is a
    suite the loop cannot rerun freely.
    
  • FAQ.md 7 KB
    # FAQ
    
    Questions people have actually asked, with the answer checked against the code
    rather than remembered. Where an answer has a boundary, the boundary is stated:
    a qualified yes is more use than an unqualified one.
    
    ## 1. Does it run locally or in the cloud?
    
    Locally. It is a `pip install` and a command line tool on your machine, or in
    CI through the bundled GitHub Action. There is no qikly service in the middle
    and nothing to sign up for.
    
    The only thing that leaves your machine is the prompt sent to whichever model
    provider you configure, using your own API key. Gemini, OpenAI and Anthropic
    are supported; see [PROVIDER_KEY_SETUP.md](PROVIDER_KEY_SETUP.md).
    
    ## 2. Does my code or my data go to you?
    
    No. Nothing is sent to the project, and the tool collects no telemetry. What
    reaches your model provider is what the prompts contain: the task file as each
    agent receives it, and the test failures the loop is working through. Your
    provider's own terms then govern that traffic.
    
    ## 3. Is there anything I can run before committing an API key?
    
    Yes, three things, all offline and free:
    
    ```bash
    qikly --explain MERGE_SALES          # what each agent is shown, and the difference
    qikly --explain MERGE_SALES --html   # the same as one page you can share
    qikly --validate                     # check your task files
    qikly --score-code src/yours.py --score-tests tests/test_yours.py
    ```
    
    `--explain` is the one worth running first: no model call, about a second.
    
    `--score-code` is the one that runs on **your** code rather than on a bundled
    task. It plants one fault at a time and reports which ones your existing tests
    did not notice. No task file, no run, no key, and nothing of yours is modified.
    Read the score as a floor: it says how much of the code that is there your
    tests would notice changing, and nothing about a rule nobody implemented.
    
    ## 4. What is the difference between `--score-code` and `--score-tests`?
    
    They are not alternatives. They are the two halves of one command, and it does
    nothing useful without both:
    
    ```bash
    qikly --score-code src/pricing.py --score-tests tests/test_pricing.py
    ```
    
    `--score-code` is **the code to plant faults in**, a module or a whole package.
    `--score-tests` is **the suite to run against each planted fault**, a file or a
    directory of them. The report names the faults your tests did not notice.
    
    Neither writes anything: the faults go into a temporary copy, and your files
    are not modified. No task file, no run, no API key, no model call.
    
    ## 5. What is the difference between `--demo` and `--example`?
    
    `--demo` is for watching, `--example` is for editing.
    
    `qikly --demo` runs a bundled task end to end in a throwaway folder you can
    delete, so you can see a real run before deciding anything. It needs an API
    key, takes about thirty seconds and costs under a cent. It changes nothing in
    your project.
    
    `qikly --example` writes a finished worked example **into your project**: a
    module, its two task files and sample data, at the paths the task files name.
    Nothing is left for you to fill in, so `qikly --tasks MY_METRICS_VERIFY` runs
    immediately, and it is the thing to copy when writing your own. It never
    overwrites an existing file.
    
    If you are deciding which to type first: `--demo` to see whether this is for
    you, `--example` once you have decided it is.
    
    ## 6. Does the coding agent really never see the acceptance criteria?
    
    That is what `--explain` exists to show, on your own task rather than on a
    claim in a README: it prints the task file as test generation receives it, then
    as the coding agent receives it, then the difference. It builds those strings
    through the same function a real run uses, so it demonstrates the mechanism
    instead of describing it.
    
    The property is also held by the test suite: no call site in the codebase can
    pass a criterion to the coding agent, so the removal cannot be undone by a
    later change without a test failing.
    
    ## 7. What does one passing run prove?
    
    That this task converged this once. A run is a loop with a variable trip count,
    so one run is an artifact and never a rate. If you want a number you can quote,
    repeat the run and report the spread. The project's own performance figures are
    in [design_2_performance.md](design_2_performance.md), with the sample sizes
    they rest on.
    
    ## 8. What will a run cost?
    
    Every run prints a projection before it starts: the expected number of model
    calls, tokens in and out, and a price from a static table rather than from your
    bill. Treat it as a projection, because the trip count varies.
    
    ## 9. Can I use it commercially?
    
    Yes. qikly is Apache 2.0, which permits commercial use, modification and
    redistribution. Note that the licence grants no trademark rights, so building a
    service on it is fine and naming that service after the project is a separate
    conversation.
    
    ## 10. Can I drive it from an editor or an agent?
    
    Yes, it ships an MCP server, so Claude Code, VS Code and other MCP hosts can
    call it. See [mcp.md](mcp.md). The server deliberately never returns acceptance
    criteria, for the same reason the coding agent never receives them.
    
    
    ## 11. Does qikly know about the classes I already have, such as hardware drivers or protocol parsers?
    
    Not by discovery: it targets what your task file's `interface` block declares,
    and `qikly --scaffold your_module.py` writes that block from the real
    signatures. Your other classes import normally at run time as long as they are
    importable from your project root.
    
    What it may **change** is a separate question from what it can import, and the
    answer is the seed: `seed.implementation` can name a single module or a whole
    package, and everything inside it is visible to the coding agent and
    repairable. A defect in a helper you left outside the seed is caught by the
    tests, cannot be repaired, and the run now says so rather than working around
    it.
    
    The full picture, including what a run prints in each case, is in
    [EXISTING_CODE_AND_HELPERS.md](EXISTING_CODE_AND_HELPERS.md).
    
    ## 12. My module is not one file. Can qikly work on a package with a nested `utils/`?
    
    Yes, from 0.5.5. Point `seed.implementation` at the directory rather than the
    file, and name the module by the import path your own code already uses:
    
    ```yaml
    interface:
      module: "mypkg.pricing"
    
    seed:
      implementation: "mypkg"
    ```
    
    The package is copied in under its own name, so `from mypkg.utils.money import
    to_cents` resolves exactly as it does in your tree. Relative imports work too,
    at any depth: `from .utils.money import to_cents`. No `__init__.py` is
    required, and one that is there is kept.
    
    **Every module in the package can be repaired**, nested ones included, and a
    single patch may change more than one of them.
    
    Two things that will not work, and both fail loudly rather than quietly: a
    package whose name is also a standard library module's, such as `json`; and
    modules that import each other circularly at the top level, which plain Python
    rejects too. The full matrix, including the import shapes that do not work,
    is in [TASK_FILE_REFERENCE.md](TASK_FILE_REFERENCE.md#seeding-a-package-when-the-implementation-is-more-than-one-module).
    
  • favicon.ico 4.9 KB · in bundle
  • favicon.svg 1.5 KB · in bundle
  • index.html 52.9 KB · in bundle
  • mcp.md 11.4 KB
    # qikly over MCP
    
    `qikly` runs as an [MCP](https://modelcontextprotocol.io) server, so an agent in
    any MCP host can start a run and read its result without you leaving the
    conversation to type a command. It is tested in VS Code and Claude Code so far;
    the setup for other hosts below follows the MCP standard but has not been
    verified end to end.
    
    With [uv](https://docs.astral.sh/uv/) installed, nothing else needs installing.
    In Claude Code:
    
    ```bash
    claude mcp add qikly -- uvx --from "qikly[mcp]" qikly-mcp
    ```
    
    Other hosts take the same command in their own config file, and it speaks
    stdio:
    
    ```json
    {
      "mcpServers": {
        "qikly": { "command": "uvx", "args": ["--from", "qikly[mcp]", "qikly-mcp"] }
      }
    }
    ```
    
    Without uv, install with pip and register the `qikly-mcp` command instead:
    
    ```bash
    pip install "qikly[mcp]"
    claude mcp add qikly -- qikly-mcp
    ```
    
    `uvx` keeps the version it first installed. To pick up a new release, run
    `uv cache clean qikly` and restart the host.
    
    Run it from the project directory, the one holding `inputs_private/`. That is
    how it finds your tasks, and it is the same rule the CLI follows.
    
    ## One-click install, and which editors have it
    
    The README badge installs into **VS Code**. It writes the `uvx` command above,
    so it needs uv and installs nothing itself. The same redirect serves Insiders
    from its own host:
    
    ```markdown
    [![VS Code Insiders](https://img.shields.io/badge/VS_Code_Insiders-Install_qikly_MCP-24bfa5?logo=visualstudiocode&logoColor=white)](https://insiders.vscode.dev/redirect/mcp/install?name=qikly&config=%7B%22name%22%3A%22qikly%22%2C%22command%22%3A%22uvx%22%2C%22args%22%3A%5B%22--from%22%2C%22qikly%5Bmcp%5D%22%2C%22qikly-mcp%22%5D%7D)
    ```
    
    A common mistake in other projects' READMEs is pointing both buttons at
    `insiders.vscode.dev`. Stable is `vscode.dev`, Insiders is
    `insiders.vscode.dev`, and the payload is identical.
    
    **Other editors have their own deeplink schemes**, and qikly does not yet ship
    buttons for them because none has been tested against a real install here:
    
    | Editor | Scheme |
    | --- | --- |
    | Visual Studio | `vs-open.link/mcp-install` |
    | Cursor | `cursor://anysphere.cursor-deeplink/mcp/install` |
    | Goose | `goose://install-mcp` |
    | LM Studio | `lmstudio://add_mcp` |
    
    They differ in how the config is encoded, and at least one uses base64 where
    VS Code uses URL-encoded JSON. A button that writes a malformed config is worse
    than no button, because the reader blames the tool rather than the link. Until
    each is tried, use the plain configuration above: **every one of these editors
    accepts a hand-written config**, and the button only ever saves you a paste.
    
    ## Let it register itself
    
    ```bash
    qikly --install-mcp --dry-run    # what it would write, and where
    qikly --install-mcp              # write it
    ```
    
    It writes **project-local files only**: `.mcp.json` for Claude Code and
    `.vscode/mcp.json` for VS Code. Never `~/.claude.json`, never the VS Code user
    profile. Those hold every other server you have, and the project file is also
    the right place on its own terms, because the setting that reliably goes wrong
    is which folder the server treats as the project.
    
    It merges rather than overwrites, so your other servers and every other key in
    the file survive, and it copies the file to a timestamped backup first. Running
    it twice changes nothing and says so.
    
    It refuses in three cases, and prints the block for you to paste instead:
    
    | It stops when | Because |
    |---|---|
    | the file holds comments | VS Code's `mcp.json` is JSONC, and writing it back as JSON would delete every comment you wrote |
    | a `qikly` entry exists and differs | you changed it on purpose. `--force` says otherwise |
    | the file is not valid JSON | guessing what you meant is how a config gets lost |
    
    **If `claude` is not a command**, it is bundled inside the VS Code extension
    rather than installed on `PATH`, at
    `%USERPROFILE%\.vscode\extensionsnthropic.claude-code-<version>-<platform>
    esources
    ative-binary\claude.exe`.
    The version is in the path, so it moves on every extension update. You do not
    need the CLI for any of this; it is only how you check the registration from a
    terminal.
    
    **Claude Code needs one approval after this.** A project-scoped `.mcp.json`
    is not trusted on sight, and it should not be: cloning a repository must not
    silently run whatever it ships. `claude mcp list` will show qikly as pending
    until you run `claude` once in that directory and approve it.
    
    Add `--install-mcp claude` or `--install-mcp vscode` to target one host.
    
    Cursor and Codex CLI are not written yet. Cursor takes the `mcpServers` shape
    below; Codex needs TOML, which the standard library cannot write on any Python
    this supports.
    
    ## Two things that will bite you on Windows
    
    Both were hit on a real machine before anyone else saw them, and neither
    announces itself clearly, so they are worth reading before you debug.
    
    **The command has to be on `PATH`.** The uv route and the one-click button run
    `uvx`, which the uv installer puts on `PATH`, though an editor opened before
    you installed uv will not see it until you restart the editor. The pip route
    names a bare `qikly-mcp`, and on Windows `pip` often
    installs console scripts into a `Scripts` directory that is not on `PATH`, and
    the host then reports only that the command was not found. Check it:
    
    ```powershell
    Get-Command qikly-mcp
    ```
    
    If that finds nothing, let qikly print a config that does not depend on
    `PATH`. In your project folder:
    
    ```powershell
    python -m qikly --mcp-config
    ```
    
    It prints a VS Code config naming the full path of the Python that has qikly,
    run as `python -m qikly.mcp_server`, and the folder you ran it from as
    `QIKLY_PROJECT_ROOT`, which also settles the folder problem below.
    `--mcp-config claude` prints the `mcpServers` shape for Claude Code, Cursor
    and most other hosts. It only prints; it never edits a config file, because
    those hold your other servers too. Use `python -m qikly` rather than `qikly`
    here: when `qikly-mcp` is not on `PATH`, neither is `qikly`.
    
    **The host starts the server in the folder your editor has open**, which is
    often the parent of your qikly project rather than the project itself. The
    server starts, and every task looks absent. Name the project explicitly:
    
    ```json
    {
      "servers": {
        "qikly": {
          "type": "stdio",
          "command": "qikly-mcp",
          "env": { "QIKLY_PROJECT_ROOT": "C:\\path\\to\\your\\project" }
        }
      }
    }
    ```
    
    The tools detect this case rather than reporting an empty project: if the
    resolved root has no `inputs_private/`, they say so and name the directory they
    resolved, because that one fact is the difference between a misconfiguration
    and an apparently broken tool.
    
    ## The four tools
    
    | Tool | What it achieves | Input | Returns | Safe to call unattended? | CLI equivalent |
    | --- | --- | --- | --- | --- | --- |
    | `qikly_run` | Starts a full run for one task: generates the suite from the acceptance criteria, writes an implementation from the requirements alone, and repairs it against the suite until every stage passes or the retry budget is spent. Returns **immediately**, because a run takes minutes to hours. | `task_id`, optional `provider` and `model` | A `run_id` to poll with | **No.** It writes code and tests, calls a model provider over the network, and spends real money. A second call is a second run, not a repeat: the agents are not deterministic even at a fixed seed. | `qikly --run <task>` |
    | `qikly_status` | Reports where a run got to, reading what the run itself wrote rather than watching the process. That is what lets it tell a crash from a failing suite: `stalled` means the process is gone without writing a summary. While a run is in flight it also gives the stage and iteration. | `run_id` | One of `running`, `passed`, `failed`, `stalled`, `unknown`, plus stage and iteration | **Yes.** Reads local run records only. Free, no network. | `qikly --status <run_id>` |
    | `qikly_validate` | Checks a task file offline before you spend anything on it: valid YAML, the required sections present, `acceptance_criteria` a list rather than one long string, and fixture paths that actually resolve. It does **not** look for contradictions between requirements and criteria: that is `qikly --check-criteria`, a paid model call this server deliberately does not expose. | `task_id` | Counts and a verdict | **Yes.** No model call, so no network and no cost. Returns no criterion text. | `qikly --validate --tasks <task>` |
    | `qikly_scaffold` | Turns a Python file you already have into a task that tests it: the module path, the real signatures of its public functions, a guessed entrypoint, and a seed pointing back at the file. `requirements` and `acceptance_criteria` are left as `TODO` on purpose, because criteria read out of an implementation can only describe what it already does. | `file_path` | The task YAML as text, for you to save under `inputs_private/config/tasks/` | **Yes.** Returns the YAML rather than writing it, so saving stays your decision. | `qikly --scaffold <file>` |
    
    Every tool carries the four MCP behaviour annotations, so a host can act on the
    column above rather than guess: `readOnlyHint`, `destructiveHint`,
    `idempotentHint` and `openWorldHint`. Three of the four are read-only, free and
    local. Only `qikly_run` is none of those things.
    
    ## Why `qikly_run` does not wait
    
    A run takes minutes to hours. No MCP host will hold a tool call open that long,
    so `qikly_run` starts the run in a detached process and hands back an id:
    
    ```
    qikly_run(task_id="CALC_TAX")
      -> {"run_id": "CALC_TAX_20260910_113412", "state": "running"}
    ```
    
    The agent then polls:
    
    ```
    qikly_status(run_id="CALC_TAX_20260910_113412")
      -> {"state": "running", "progress": {"stage": "unit", "iteration": 2}}
    ```
    
    Because the run is detached, you can close the terminal, close the editor, and
    ask for the status tomorrow. The run outlives the conversation that started it.
    
    `stalled` is worth knowing about: it means the run died without writing a
    summary, which is a crash rather than a failing suite. The two need different
    reactions, so they get different words.
    
    ## What these tools will not tell you
    
    No qikly tool returns your acceptance criteria. Not on success, and especially
    not on failure, which is exactly when a helpful tool wants to explain why and
    "why" is the criterion.
    
    That is the point of the whole library, so it is enforced rather than intended:
    `tests/test_mcp_withholding.py` asserts on the serialised JSON that crosses the
    wire, for every tool, including the error paths.
    
    Your agent will see that four tests failed, and the pytest output for each. It
    will not see the rule it broke. That is the same information a human developer
    gets from a failing CI run, and it is what stops the agent writing code aimed at
    the test instead of at the requirement.
    
    ## One thing you have to do yourself
    
    These tools control what a *response* contains. They cannot control what your
    agent reads off your disk, and **the generated tests under `outputs/tests/` are
    written from your criteria**. An agent that opens those files has the answer
    key, without any tool call being involved.
    
    This matters more than it first sounds, because the natural way to use this
    server is to let the same agent both write your code and call `qikly_run`.
    
    So tell your agent to leave the generated tests alone. In Claude Code, add to
    `.claude/settings.json`:
    
    ```json
    {
      "permissions": {
        "deny": ["Read(./outputs/tests/**)"]
      }
    }
    ```
    
    Or add `outputs/tests/` to whatever your host uses to keep files out of an
    agent's reach. Reading the *summary* is fine and useful. Reading the tests is
    the one habit that quietly undoes the reason you installed this.
    
  • PROVIDER_KEY_SETUP.md 8 KB
    # Setting up a provider key
    
    Every qikly run calls one model provider, and that provider needs a key. This
    page covers all three, on PowerShell, bash and CI, plus how to check a key took
    and how to stop paying for one you forgot about.
    
    If you only want the shortest path: get a
    [Gemini key](https://aistudio.google.com/apikey), which has a free tier and
    needs no card, then `qikly --demo`.
    
    ---
    
    ## Which provider
    
    | Provider | Key from | Install | Notes |
    |---|---|---|---|
    | **Gemini** (default) | [aistudio.google.com/apikey](https://aistudio.google.com/apikey) | included | Free tier, no card. Cheapest by roughly 20x. Every published qikly figure was measured on it |
    | **OpenAI** | [platform.openai.com/api-keys](https://platform.openai.com/api-keys) | `pip install "qikly[openai]"` | Requires billing set up |
    | **Anthropic** | [console.anthropic.com/settings/keys](https://console.anthropic.com/settings/keys) | `pip install "qikly[anthropic]"` | Requires billing. No seed and no temperature, so runs vary more than the others. A reasoning model, so it produces thinking tokens you are billed for on top of the answer |
    
    You do not have to choose one forever. All three keys can sit in your
    environment at once; `LLM_PROVIDER` decides which is used, and if you set no
    provider and hold exactly one key, qikly uses that one and says so. The recipes
    below set it anyway, because each one is about a named provider. With a single
    key you can leave it out, which is why the quick start does not set it.
    
    **Restricting the key.** qikly calls exactly one endpoint per provider, so a
    minimal key is enough. On OpenAI, choose **Restricted** and grant only
    **Chat completions (`/v1/chat/completions`)**; every other row stays `None`.
    Read-only will not work, because creating a completion is a write.
    
    ---
    
    ## Which model
    
    Pick a small fast one. A run makes one model call per stage per iteration, so
    the length of a run is mostly the length of a single call, and a model that
    reasons before it answers turns a two minute run into twenty.
    
    | Provider | Start with | Costs you a long wait |
    |---|---|---|
    | **Gemini** | `gemini-3.5-flash-lite`, the default, or any flash model | `gemini-3.5-pro` |
    | **OpenAI** | a mini model | the reasoning models |
    | **Anthropic** | `claude-haiku-4-5` | `claude-sonnet-5`, which thinks before every answer and bills the thinking |
    
    qikly says this in the run banner when the model you chose is one that thinks,
    and again the first time a call runs long, so a slow run explains itself rather
    than looking like a hang.
    
    **Why the smallest model is the one every figure was measured on.** Every
    published qikly number comes from `gemini-3.5-flash-lite`, over hundreds of
    runs. That is the cheapest and least capable of the three defaults, and the
    choice was deliberate: a number measured there is a floor rather than a best
    case. A later flash model should clear it, not fall short of it. If yours does
    not, that is a result worth reporting.
    
    None of which makes a thinking model wrong. It writes a stricter suite, and if
    that is what you are after, the wait is what it costs. It does mean not to
    reach for one first, and that a run which seems stuck is usually this.
    
    ---
    
    ## Windows PowerShell
    
    ### Gemini
    
    ```powershell
    $env:GEMINI_API_KEY = "AIza..."
    $env:LLM_PROVIDER   = "gemini"
    
    Write-Output "provider: $env:LLM_PROVIDER"
    Write-Output "key:      $($env:GEMINI_API_KEY.Length) chars, ends ...$($env:GEMINI_API_KEY.Substring($env:GEMINI_API_KEY.Length-4))"
    
    qikly --demo
    ```
    
    ### OpenAI
    
    ```powershell
    pip install "qikly[openai]"
    
    $env:OPENAI_API_KEY = "sk-proj-..."
    $env:LLM_PROVIDER   = "openai"
    
    Write-Output "provider: $env:LLM_PROVIDER"
    Write-Output "key:      $($env:OPENAI_API_KEY.Length) chars, ends ...$($env:OPENAI_API_KEY.Substring($env:OPENAI_API_KEY.Length-4))"
    
    qikly --demo
    ```
    
    ### Anthropic
    
    ```powershell
    pip install "qikly[anthropic]"
    
    $env:ANTHROPIC_API_KEY = "sk-ant-..."
    $env:LLM_PROVIDER      = "anthropic"
    
    Write-Output "provider: $env:LLM_PROVIDER"
    Write-Output "key:      $($env:ANTHROPIC_API_KEY.Length) chars, ends ...$($env:ANTHROPIC_API_KEY.Substring($env:ANTHROPIC_API_KEY.Length-4))"
    
    qikly --demo
    ```
    
    The check prints the length and last four characters rather than the key. A
    plain `$env:OPENAI_API_KEY` writes the whole secret into your terminal
    scrollback and your PSReadLine history file, which is a bad habit for something
    that bills you.
    
    ### Making it stick
    
    `$env:` lasts for that window only. Close the terminal and the key is gone.
    
    ```powershell
    [Environment]::SetEnvironmentVariable("OPENAI_API_KEY", "sk-proj-...", "User")
    ```
    
    Then **open a new terminal**: the current one will not see it.
    
    ```powershell
    # what is persisted, as opposed to only set here
    [Environment]::GetEnvironmentVariable("OPENAI_API_KEY", "User")
    
    # everything qikly might pick up, in this session
    Get-ChildItem Env: | Where-Object Name -match 'API_KEY|LLM_'
    
    # remove one
    Remove-Item Env:\OPENAI_API_KEY
    ```
    
    That last listing is the fastest way to find a stale `LLM_PROVIDER` from an
    earlier experiment, which is the usual reason a run goes somewhere unexpected.
    
    ---
    
    ## macOS and Linux
    
    ```bash
    # Gemini
    export GEMINI_API_KEY=AIza...
    export LLM_PROVIDER=gemini
    
    # OpenAI
    pip install "qikly[openai]"
    export OPENAI_API_KEY=sk-proj-...
    export LLM_PROVIDER=openai
    
    # Anthropic
    pip install "qikly[anthropic]"
    export ANTHROPIC_API_KEY=sk-ant-...
    export LLM_PROVIDER=anthropic
    
    qikly --demo
    ```
    
    Check it took without printing it:
    
    ```bash
    echo "${#OPENAI_API_KEY} chars, ends ...${OPENAI_API_KEY: -4}"
    env | grep -E 'API_KEY|LLM_' | sed 's/=.*/=<set>/'
    ```
    
    To persist, add the `export` lines to `~/.zshrc` or `~/.bashrc`, or keep them
    in a `.env` you source. Do not commit either.
    
    ---
    
    ## CI
    
    **GitHub Actions.** Store the key as a repository secret, never in the
    workflow file:
    
    ```yaml
    - uses: gal-a/qikly@v0.5.7
      with:
        api-key: ${{ secrets.OPENAI_API_KEY }}
        provider: openai
        qikly-version: "qikly==0.5.7"
    ```
    
    A secret is masked in logs. A literal is not, and a key pushed to a public repo
    is compromised within minutes, whether or not the commit is later removed.
    
    ---
    
    ## Checking it worked
    
    ```bash
    qikly --validate     # free, and will NOT catch a bad key: it opens no socket
    qikly --demo         # one task, real calls, a few cents
    ```
    
    `--demo` is the real test. It reports the provider, the model, the estimated
    cost, and whether the task converged.
    
    **What a wrong key looks like:**
    
    ```
    OpenAI API rejected the request: Error code: 401 ...
    If this looks like an auth error, check API_KEY for typos/whitespace
    ```
    
    Whitespace is the usual culprit: a trailing space or newline copied along with
    the key. Compare the length you printed above against the length on the
    provider's console.
    
    **What a wrong model looks like:**
    
    ```
    The model `gemini-3.5-flash-lite` does not exist or you do not have access
    ```
    
    That is a provider and model that disagree. Check `LLM_MODEL` is unset or
    correct for the provider you chose.
    
    ---
    
    ## Cost, and how not to be surprised
    
    Every run prints an estimate before it starts and the real usage afterwards.
    The estimate comes from your own run history once you have some.
    
    A single `--demo` on the default provider is a fraction of a cent. The same
    demo on gpt-4o is roughly twenty times that, and it will be slower.
    
    Three ways to bound it:
    
    - `QIKLY_MAX_CALLS=50` stops a run dead after N model calls.
    - `QIKLY_MAX_OUTPUT_TOKENS` caps Anthropic's per-reply budget, default 16384.
      It has to cover thinking as well as the answer: at 4096 claude-sonnet-5 spent
      the entire budget reasoning and returned no answer at all.
    - Set a spend limit on the provider's own console. **This is the one that
      actually protects you**, because it applies whatever calls the API.
    - On OpenAI, put the key in its own project and cap that project. Then a
      mistake here cannot spend what you budgeted for something else.
    
    ---
    
    ## If you rotate or revoke a key
    
    Nothing in qikly stores a key. It reads the environment on every call, so
    revoking on the provider's console and exporting a new value is the whole
    procedure. There is no cache to clear and no config file to edit.
    
  • QUICK_START_ON_YOUR_OWN_DATA.md 18 KB
    # Quick start on your own data
    
    From nothing to a first run on your own module. Try the bundled demo, follow
    the five steps, then read the one question that decides where each line of your
    specification goes. That is the whole path, and it is the whole of this page.
    
    Everything else lives in [the task file reference](https://github.com/gal-a/qikly/blob/main/docs/TASK_FILE_REFERENCE.md): which command fits what
    you already have, the file field by field, criteria you wrote elsewhere,
    seeding your own code or tests, and the fixture rows a criterion needs before
    it can be checked at all. Come back for those when you want them; you do not
    need any of it for a first run.
    
    ## Try it first
    
    Requires **Python 3.10+** and **GNU `patch`** on `PATH`. On Windows it ships
    with Git under `usr\bin\patch.exe`, which the tool finds on its own. **On macOS
    you have to install it:** the system `patch` is Apple's BSD one, which rejects
    the options qikly sends, so no generated diff will apply.
    
    ```bash
    brew install gpatch     # macOS only
    ```
    
    qikly looks for `gpatch` before `patch`, so nothing else is needed afterwards
    and your system `patch` is left alone.
    
    **Install into a virtual environment**, and not only out of habit. Two reasons
    specific to this tool. qikly pulls in a provider SDK, so a bare install can
    upgrade a package your own project pinned. And qikly resolves where to read and
    write from the environment it is running in, so "which interpreter am I in" is
    a question you will want a clean answer to the first time something behaves
    oddly. `qikly --version` is that answer: it prints the version, the package
    directory and the interpreter together.
    
    ```bash
    python -m venv .venv && source .venv/bin/activate    # Windows: .venv\Scripts\activate
    pip install qikly
    qikly --version                # which build, from where, on which Python
    export GEMINI_API_KEY=...      # or OPENAI_API_KEY / ANTHROPIC_API_KEY
    qikly --demo
    ```
    
    On Windows, in PowerShell, where `export` is not a command:
    
    ```powershell
    python -m venv .venv
    .venv\Scripts\activate
    pip install qikly
    qikly --version
    $env:GEMINI_API_KEY = "..."    # or OPENAI_API_KEY / ANTHROPIC_API_KEY
    qikly --demo
    ```
    
    Let the install finish before running anything. The provider SDK is a long
    dependency chain and pip installs qikly itself last, so there is a window where
    its dependencies are present and qikly is not, which looks exactly like a
    broken install.
    
    Only `--demo` needs a key. `--version`, `--explain`, `--validate`,
    `--init`, `--example` and `--scaffold` make no model call and cost
    nothing.
    
    Gemini is the default and its key is `GEMINI_API_KEY`. For OpenAI or
    Anthropic, which need an extra install, see [the provider table in
    docs/CONFIGURATION.md](https://github.com/gal-a/qikly/blob/main/docs/CONFIGURATION.md),
    which also says where to get each key.
    
    **The default model differs by provider**, and they are not the same size:
    Gemini gets `gemini-3.5-flash-lite`, OpenAI `gpt-4o`, Anthropic
    `claude-sonnet-5`, which is a reasoning model and takes minutes rather than
    seconds per run. [Which default you get, and how to change
    it](https://github.com/gal-a/qikly/blob/main/docs/TROUBLESHOOTING.md#provider-defaults).
    
    One key is enough and `LLM_PROVIDER` is optional: qikly uses the one key it
    finds and prints which. Set `LLM_PROVIDER` to `gemini`, `openai` or `anthropic`
    only when you hold more than one and want to choose. The key is a shell
    variable rather than a venv one, so set it once in the terminal and the venv
    sees it too. Full recipes, including making a key survive a new terminal, are
    in [docs/PROVIDER_KEY_SETUP.md](https://github.com/gal-a/qikly/blob/main/docs/PROVIDER_KEY_SETUP.md).
    
    `--demo` runs one task end to end in a throwaway `demo/throwaway_<timestamp>/` directory
    and prints what it built and where. It writes nothing outside that directory,
    so a first run leaves everything else untouched. About 30 seconds on the Gemini
    default, and [several minutes on the Anthropic
    one](https://github.com/gal-a/qikly/blob/main/docs/TROUBLESHOOTING.md#provider-defaults).
    After ten seconds of silence a run starts printing one line every fifteen
    saying how long it has been waiting, so a slow call is distinguishable from a
    hang.
    
    For the same thing on vehicle sensor data rather than an order pipeline:
    
    ```bash
    qikly --demo --tasks ADAS_HEADWAY
    ```
    
    `ADAS_HEADWAY` checks following distance from forward-radar samples. Its
    requirements give the limits and the two-second rule; its withheld criteria
    pin what happens exactly at each limit, including that a gap of zero metres is
    not a measurement.
    
    From a clone instead:
    
    ```bash
    pip install -r requirements.txt
    python run.py --demo
    ```
    
    `run.py` is a shim around `src/qikly/cli.py`, the same entry point the
    installed `qikly` command calls, so a clone and an install run identical
    code.
    
    ## Your own module, start to first run
    
    Rather see it work before you point it at your own code?
    
    ```bash
    qikly --example
    ```
    
    **That is the five steps below, already done, on a module you do not have
    to write.** It lays down `my_metrics.py` at your project root, both task
    files a scaffold of it produces, and sample data at the paths those task
    files name. Its `requirements` and `acceptance_criteria` are written in,
    which a scaffold of your own code cannot do for you, so it runs as it
    stands:
    
    ```bash
    qikly --validate --tasks MY_METRICS_VERIFY    # free, no model call
    qikly --tasks MY_METRICS_VERIFY
    ```
    
    | The five steps below, on your module | What `--example` hands you instead |
    |---|---|
    | 1. Scaffold a task from the module | `my_metrics.py`, and the two task files a scaffold of it writes |
    | 2. Put your input data where the task says | two sample CSVs, at the paths those task files name |
    | 3. Write the two sections only you can write | written already, so you can read a finished pair before writing your own |
    | 4. Check it, for free | the same command |
    | 5. Run it | the same command |
    
    So the difference is step 3, and step 3 is the part that matters: those
    two sections are the whole mechanism, and reading a worked pair is the
    fastest way to see the split. [What it lays down, and what to look at in
    it](https://github.com/gal-a/qikly/blob/main/docs/SCAFFOLDED_TASK_EXAMPLE.md).
    
    It leaves the rest of your project alone, and never overwrites, so a
    second run keeps anything you edited.
    
    **Which starter is which.** `--init` writes `MY_FIRST_TASK`, an empty form with
    `TODO` where your rules go and a stub CSV, and you bring the code. `--example`
    writes `MY_METRICS`, the same form filled in, with a module and real sample data
    behind it. Neither command writes the other's task, so whichever you ran is the
    only task in your project.
    
    > **No module to start from?** You do not need one. `qikly --init` writes a
    > starter task with `requirements` and `acceptance_criteria` and no `seed:`
    > block, so a run writes the first implementation from your requirements
    > instead of testing code you already have. Skip step 1, fill in those two
    > sections, and the rest of the steps are unchanged. Scaffolding exists to read
    > signatures out of code that already exists, which is the only part you are
    > missing.
    
    1. **Scaffold a task from the module.**
    
       ```bash
       qikly --scaffold my_metrics.py
       ```
    
       It reads the real function signatures and writes a task file, then prints
       what to do next.
    
       The name comes from the module's own filename, upper-cased, and which of two
       task files you get depends on the job:
    
       | Command | Writes | What that task does |
       |---|---|---|
       | `qikly --scaffold my_metrics.py` | [`MY_METRICS_VERIFY.yaml`](https://github.com/gal-a/qikly/blob/main/src/qikly/inputs_public/examples/config/tasks/MY_METRICS_VERIFY.yaml) | Tests the code you already have |
       | `qikly --scaffold my_metrics.py --fresh` | [`MY_METRICS.yaml`](https://github.com/gal-a/qikly/blob/main/src/qikly/inputs_public/examples/config/tasks/MY_METRICS.yaml) | Writes a fresh implementation of the same interface, and tests that |
    
       Both land in `inputs_private/config/tasks/`. The `_VERIFY` suffix is what
       keeps them apart, so scaffolding the same module both ways never overwrites
       the first file with the second. The rest of this section uses
       `MY_METRICS_VERIFY` as the example; substitute your own.
    
    2. **Put your input data where the task says.** Its `inputs:` list names the
       files a run reads, here `inputs_private/data/MY_METRICS/input_01.csv`. Both
       task files point at that one folder, named for the module rather than the
       task, so scaffolding both ways does not split your data in two. Scaffold
       does not create them, so copy a real sample of your data there. Until you
       do, `qikly --validate` reports `input file not found`. What a good sample
       looks like: [`input_01.csv`](https://github.com/gal-a/qikly/blob/main/src/qikly/inputs_public/examples/data/MY_METRICS/input_01.csv) and
       [`input_02.csv`](https://github.com/gal-a/qikly/blob/main/src/qikly/inputs_public/examples/data/MY_METRICS/input_02.csv).
    
    3. **Write the two sections only you can write.** `requirements` holds the
       decisions and `acceptance_criteria` the consequences; the rule for telling
       them apart is under [Getting the two halves right](#getting-the-two-halves-right).
       Already written them in a page or a ticket? This takes the criteria from it:
       `qikly --scaffold my_metrics.py --from-doc feature.md`. Replace the
       remaining `TODO` lines too, and check the entrypoint scaffold marks as
       guessed.
    
    4. **Check it, for free.**
    
       ```bash
       qikly --validate --tasks MY_METRICS_VERIFY
       ```
    
       No model call and no cost. Without `--tasks` it also checks every bundled
       example. It checks that the file parses, that every input
       path exists, that no `TODO` placeholder is left, that criteria name values
       rather than adjectives, and that no requirement restates a criterion.
    
    5. **Run it.**
    
       ```bash
       qikly --tasks MY_METRICS_VERIFY
       ```
    
       Start reading at `outputs/reports/iterations/<task>_<timestamp>_report.html`.
    
       A `_VERIFY` task tests code you already have, which qikly did not write, so
       it cannot enforce that the code's author never saw your acceptance criteria.
       It evidences what it can: the run opens by asking git whether this task file
       was last changed before that code's first commit, and prints the answer with
       what it does not show. A commit date is not a writing date, and criteria
       settled first is the precondition of the opposite problem, someone coding to
       the bar. So read it as corroboration, and read the bad answer, criteria
       revised after the code landed, as the question it is.
    
    **Tried it?** [Tell us what happened](https://github.com/gal-a/qikly/discussions/6), whether it worked, stalled
    or never got past install.
    
    **Keeping the suite?** Add the badge to your project's README, so the people
    reading it know the tests were written by an agent kept apart from the code:
    
    ```markdown
    [![tested with qikly](https://img.shields.io/badge/tested_with-qikly-2b8f95)](https://test.qikly.com/?ref=badge)
    ```
    
    ### Step 1 from inside VS Code
    
    With the qikly MCP server connected (setup in
    [docs/mcp.md](https://github.com/gal-a/qikly/blob/main/docs/mcp.md)), ask Copilot
    in agent mode:
    
    > Use the qikly_scaffold MCP tool on `src/my_metrics.py`, and save the task it
    > returns under `inputs_private/config/tasks/`.
    
    It returns the same task the command writes, one that tests the code you
    already have, and says which filename to save it as. Steps 2 to 5 are the same,
    and `qikly_validate` runs step 4 from the chat at no cost.
    
    ### What a real run on your own code looks like
    
    Not the bundled demo. This is an ordinary module that totals invoice lines,
    scaffolded with `qikly --scaffold`, with the two human sections filled in by
    hand. The whole run cost **$0.003** and seven model calls.
    
    ```
    [stage 1/3] [iteration 1] 3/4 passed | FAILED: test_quantity_boundary_zero_and_negative
    [stage 1/3] [iteration 2] 4/4 passed
    [stage 2/3] [iteration 1] 0/1 passed | FAILED: test_system_entrypoint_output_structure_and_types
    [stage 2/3] [iteration 2] 1/1 passed
    [stage 1/3] [iteration 2-regcheck] 4/4 passed
    [stage 3/3] [iteration 1] 5/5 passed
    ```
    
    **Line 1 is the whole point.** The existing code validated that `qty` parsed as
    a number and stopped there, so a quantity of zero or minus one went through as
    a real order. One acceptance criterion said otherwise:
    
    ```yaml
    acceptance_criteria:
      - "A qty of 0 is rejected and a qty of 1 is accepted; a negative qty is rejected"
    ```
    
    The suite was written from that criterion before the run touched the code, and
    the coding agent never saw the criterion itself. What it received was pytest's
    output for the failing test: the name, the test's own source and docstring, and
    the assertion error. From that it produced this patch:
    
    ```diff
             try:
                 qty = int(row["qty"])
                 unit = float(row["unit_price"])
    +            if qty <= 0:
    +                rejected.append(dict(row, reason="qty must be positive"))
    +                continue
             except (ValueError, TypeError, KeyError):
    ```
    
    That is a real bug in code that already existed, found by a test written from a
    rule the agent repairing it could not read.
    
    **The generated tests say where they came from**, so the suite is reviewable
    rather than a black box:
    
    ```python
    def test_quantity_boundary_zero_and_negative():
        """Verify that a quantity of 0 is rejected, 1 is accepted, and negative
        quantities are rejected.
        # Requirements: 2, 4
        # Criteria: 2
        """
    ```
    
    **One honest note about the other criterion.** The same task asked for
    round-half-away-from-zero on currency, and the test for it passed on the first
    attempt against code using plain `round()`. Not because the code was right in
    general, but because on this particular input the binary representation of
    10.005 lands just above the halfway point and `round()` returns 30.02 anyway.
    A criterion is only as good as the input that exercises it, which is the same
    point as [fixture coverage](https://github.com/gal-a/qikly/blob/main/docs/TASK_FILE_REFERENCE.md#proposing-fixture-rows) further down.
    
    ### Getting the two halves right
    
    This matters more than anything else in the file, and one question settles
    most of it.
    
    **`requirements` holds the decisions.** Anything a person chose that could have
    gone another way: a threshold, a unit, a measurement convention, an exemption.
    Nobody can guess a decision, so the coding agent has to be told.
    
    **`acceptance_criteria` holds the consequences.** What must be true if those
    decisions were implemented correctly: the exact boundary, the identity that has
    to hold, the case a careless reading gets wrong. They are withheld because they
    are the exam.
    
    > **Given only the requirements, could two competent developers legitimately
    > disagree about this line?** If yes, it is a decision and belongs in
    > `requirements`. If no, it is a consequence and belongs in
    > `acceptance_criteria`.
    
    | In `requirements`, because it is a decision | In `acceptance_criteria`, because it follows |
    |---|---|
    | Keep at least 2.5 m from the vehicle ahead, measured centre to centre | At exactly 2.5 m, no violation is raised |
    | Amounts are currency, rounded to the nearest cent | For every accepted row, total equals subtotal plus tax, exactly |
    | Dates are written YYYY-MM-DD | 2026-02-30 is rejected, because it is not a real date |
    
    **The gap between them is the entire mechanism**: with nothing withheld, both
    sides read the spec identically and every test passes first try, which proves
    nothing.
    
    **Both mix-ups have a signature. Learn to read them.**
    
    **A decision in `acceptance_criteria`** is withheld from the one agent that
    needed it, so the coding agent has to guess a choice nobody told it. It does
    not produce a harder test, it produces repetition, in one of two shapes. Either
    the same test fails while the FIX and PATCH come back near identical each time,
    because nothing the agent can see would lead it anywhere else, or two tests
    disagree and each patch makes one pass and the other fail. Restrict street
    suffixes to three valid values and the model keeps widening them, since
    everything it knows says "Boulevard" is a suffix.
    
    Sometimes, though, the agent simply guesses right and the run goes green. That
    is the worse outcome, because nothing then tells you a decision was in the wrong
    half. On a bundled task whose criteria alone settled whether exactly 2.00
    seconds of headway raises a warning, three runs in ten converged anyway. Do not
    rely on the loop to find these for you; apply the question above when you write
    the spec. [TROUBLESHOOTING.md](https://github.com/gal-a/qikly/blob/main/docs/TROUBLESHOOTING.md#5-the-same-patch-appearing-over-and-over)
    has the diagnosis for the repeating case.
    
    **A consequence in `requirements`** is the quieter mistake. Both agents read the
    same boundary value, so the test that checks it passes on the first attempt and
    proves nothing. The rest of the suite is unaffected and still bites, which is
    what makes it easy to miss: the run looks entirely normal. Nothing fails, and
    nothing was learned about that boundary.
    
    **The loop cannot fix either one.** It can tighten a bar the code already
    attempts, and it cannot tell you a line is in the wrong place.
    
    An assistant reading your spec can help with half of it. Given the question
    above, it can say which line looks misplaced and why, and qikly's
    [agent Skill](https://github.com/gal-a/qikly/blob/main/docs/skill.md) exists
    partly to make it good at that. What it cannot do is settle the decision:
    whether exactly 2.00 s warns, or whether an empty string counts as missing, is
    not in the specification, which is what makes it a decision. Someone has to
    choose, and that someone is you.
    
    **To see a task that follows the rule,** run `qikly --explain CALC_TAX`. Its
    requirements say amounts are "currency amounts rounded to the nearest cent",
    and the withheld criteria pin what that already means, down to a float result
    of 434.99999999999994 reporting as 435.00. More in
    [docs/design_3_mechanism.md](https://github.com/gal-a/qikly/blob/main/docs/design_3_mechanism.md#using-it).
    
    
    ## Everything else
    
    The task file field by field, the four routes in, criteria you have
    already written, seeding your own implementation or suite, where each
    artefact lands, and fixture coverage: [task file reference](https://github.com/gal-a/qikly/blob/main/docs/TASK_FILE_REFERENCE.md).
    
  • SCAFFOLDED_TASK_EXAMPLE.md 6.4 KB
    # A scaffolded task, start to finish
    
    ```bash
    qikly --example
    ```
    
    That puts the whole example in your project, ready to run:
    
    | Written to | What it is |
    |---|---|
    | `my_metrics.py` | The module you already have. Four functions, no implementation: scaffold reads signatures, never bodies. Each body carries a `TODO` saying what is expected there |
    | `inputs_private/config/tasks/MY_METRICS_VERIFY.yaml` | What `qikly --scaffold my_metrics.py` writes, with the two human sections filled in. Tests the code you already have, so it carries a `seed:` block pointing back at the module |
    | `inputs_private/config/tasks/MY_METRICS.yaml` | The same, as `qikly --scaffold my_metrics.py --fresh` writes it. A fresh implementation of the same interface, so no `seed:` block |
    | `inputs_private/data/MY_METRICS/input_01.csv` | Sample input data, first file |
    | `inputs_private/data/MY_METRICS/input_02.csv` | Sample input data, second file |
    
    It makes the `inputs_private/` layout first if you have not run `--init`,
    and leaves out that command's blank `MY_FIRST_TASK` starter, so the
    example is the only task in your project.
    
    It never overwrites, so running it twice is safe and your edits survive. Then:
    
    ```bash
    qikly --validate --tasks MY_METRICS_VERIFY    # free, no model call
    qikly --tasks MY_METRICS_VERIFY
    ```
    
    **What to expect from that run: it converged 9 times in 10.** Measured on
    2026-09-23, ten runs from a fresh project each time, on the default model. The
    nine took 4 to 6 iterations and 26 to 52 seconds; the tenth exhausted its
    budget after 17 repair attempts and exited non-zero, which is the loop working
    rather than the example being broken. A first run is a draw, not a promise, and
    this one is a better draw than the bundled tasks at roughly 6 in 10 because the
    implementation is supplied and the specification is short.
    
    Every one of those runs did real work before it converged: the module ships as
    signatures with no bodies, so the first integration attempt fails every test
    and the coding agent writes the implementation from those failures.
    
    `--validate` reports the task clean, with nothing left to fill in. That is
    the one way this differs from scaffolding your own module: scaffold leaves
    `requirements` and `acceptance_criteria` as `TODO`, because neither can be
    derived from code, and here they are written so you can read a finished
    pair. [The quick
    start](QUICK_START_ON_YOUR_OWN_DATA.md#getting-the-two-halves-right) has
    the rule for telling the two apart.
    
    ## Why these files exist
    
    Nothing here is illustrative. Everything in the two YAML files that scaffold
    decides is the output of `qikly --scaffold` run on `my_metrics.py`: the task
    id, the module path, the signatures, the output path and the `seed:` block.
    `tests/test_scaffold_examples.py` regenerates both files and fails if any of
    that drifts, so what you open is what scaffold writes today rather than what
    it wrote once. What it cannot regenerate is `requirements` and
    `acceptance_criteria`, which are written here rather than left as `TODO`,
    and a second test holds them to having nothing left to fill in.
    
    The copies that ship live under [`src/qikly/inputs_public/examples/`](../src/qikly/inputs_public/examples),
    laid out the way a real project is rather than as a flat folder of samples, so
    the paths you read there are the paths `--example` writes them to. They
    sit under `examples/` rather than beside the bundled tasks deliberately: a
    scaffolded pair breaks two rules the bundled tasks keep. `MY_METRICS_VERIFY`
    ends in the suffix reserved for scaffolding, and the pair shares one data
    folder where bundled ids map one to one. Keeping it out of the bundled task
    list also keeps `qikly --validate` quiet for everyone who never asked for the
    example. Once `--example` copies it in, it is an ordinary task in your
    project like any other, and a bare `qikly` with no `--tasks` will run it.
    
    ## What to look at in the two task files
    
    They differ in one thing that matters. `diff` them and almost everything is the
    task id and the paths derived from it. The real difference is the last block:
    `MY_METRICS_VERIFY.yaml` ends with
    
    ```yaml
    seed:
      implementation: "my_metrics.py"
    ```
    
    and `MY_METRICS.yaml` ends with a comment saying there is no `seed:` block, so
    a run writes the implementation itself. That one block is the whole difference
    between "check this code" and "write this code and check it".
    
    Both arrive with `requirements` and `acceptance_criteria` written, which is
    the one thing scaffold cannot do for your own module. Scaffold could read
    plausible criteria out of an implementation, and deliberately does not:
    criteria derived from code can only describe what that code already does,
    which is a bar it passes by construction. So when you scaffold your own
    module those two sections arrive as `TODO`, each carrying an `e.g.` showing
    the shape of a useful answer, and the pair here is what a filled in version
    looks like.
    
    Read the criteria against the requirements above them. Every one names a
    boundary value, and none of them restates a decision: that a humidity of 100
    is the last usable one follows from the requirement that the range is 0 to
    100, and could not be guessed from it by someone who had not been told where
    the range ends.
    
    ## About the sample data
    
    The `inputs:` list in a scaffolded task names one file,
    `inputs_private/data/MY_METRICS/input_01.csv`, because scaffold cannot know
    how many you have. Add the rest yourself; the example names both of its
    files, which is what that looks like once you do.
    
    Look at what is in them, because it is the part people get wrong. Between the
    two files there are readings that are clean, a missing temperature, a
    non-numeric temperature, a humidity of exactly 100 and one above it, a
    temperature of exactly 0.0, a negative temperature, and a malformed
    timestamp. Every one of those is the boundary named by one of the six
    acceptance criteria, and a test in `tests/test_scaffold_examples.py` fails if
    a criterion loses the row that reaches it. That is deliberate: **a criterion
    no row can trigger is a criterion nothing checks.**
    A suite written against data containing only clean readings will pass whatever
    the code does about bad ones, and report nothing about the rule you cared most
    about. Two thirds of the faults this project planted in its own research went
    unnoticed for exactly that reason.
    
    So when you copy your own data in, check that something in it reaches every
    criterion you wrote. These two files are sample rows, not a template: replace
    them wholesale with a real sample of your own.
    
  • skill.md 11.5 KB
    # qikly as an agent Skill
    
    **What this gets you:** your coding agent stops needing to be told about
    qikly. Ask it for tests you can trust and it reaches for the tool on its own,
    and it knows the part that is hard to guess, which line of your specification
    is a decision the coder needs and which is a consequence to withhold.
    
    A Skill is a folder of Markdown your agent reads when what you are asking
    matches what the Skill says it is for. No server, no configuration, no process
    to keep running.
    
    ## Install it
    
    **First, make sure the qikly you are about to run is the current one.** The
    Skill ships inside the package, so an old qikly writes an old Skill, and it
    then describes commands that do not exist yet.
    
    ```bash
    pip uninstall qikly -y
    pip install qikly
    qikly --version
    ```
    
    Uninstall first rather than `--upgrade`: an interrupted or repeated upgrade can
    leave more than one version's metadata behind, and pip then reports one version
    while the files on disk are another's. The clean pair takes seconds and removes
    the question. Compare what `--version` prints against
    [the latest release](https://pypi.org/project/qikly/).
    
    Then, from the directory of the project you want the Skill in:
    
    ```bash
    qikly --install-skill
    ```
    
    That writes `.claude/skills/qikly/`, which is where agents look. Then open your
    agent in that directory and ask for something ordinary, without mentioning
    qikly:
    
    > Write tests for `src/following_distance.py` that would actually catch a bug.
    
    `src/following_distance.py` stands in for a module you actually have; name a
    real one, because an agent asked about a file that does not exist will spend
    its answer asking you which file you meant. Every example on this page uses
    following distance from radar samples, which is the worked case in the Skill
    itself, so the two read together.
    
    It should reach for qikly by itself, and it does: first time in **Claude Code,
    Gemini CLI, Codex and Cursor**, in a project with a rival testing skill
    installed beside this one and a request that never mentioned qikly. That is the
    hard version of the test, because the agent had a competing option and no hint.
    Your own project has more skills in it than that one did.
    
    If it does not, see [if your agent does not pick it up](#if-your-agent-does-not-pick-it-up).
    
    **Using GitHub Copilot in VS Code?** Copilot reads none of the skill folders
    above. Run `qikly --install-skill copilot`, which writes the Skill into
    `.github/instructions/` along with the `qikly.instructions.md` file Copilot
    actually opens, then use Copilot Chat in Agent mode. That path has not been
    watched loading yet, so tell us if it works for you. The MCP server is the
    other way in and is tested there:
    [docs/mcp.md](https://github.com/gal-a/qikly/blob/main/docs/mcp.md).
    
    **A few options, none of them needed the first time.** `--dry-run` shows what
    it would write. `--force` replaces a Skill you have already edited, keeping a
    timestamped copy of the old one. And `--install-skill agents`, `cursor` or
    `gemini` write where those tools document their own skill folders, rather than
    Claude's.
    
    ## What is in it
    
    - **The one question**: given only the requirements, could two competent
      developers legitimately disagree about this line? Yes means it is a decision
      and belongs in `requirements`; no means it is a consequence and belongs in
      `acceptance_criteria`.
    - **Both mistakes and their signatures**, so an agent can recognise a stalled
      loop as a spec problem rather than a code problem.
    - **How to write a criterion that can be tested**: name values not adjectives,
      name both sides of a boundary, and make sure the data contains what you name.
    - **The five steps** on your own module, and which commands are free.
    - **What to do when a run does not converge**, and the one thing never to do,
      which is loosen a criterion to get green.
    - **The honest limit**, so an agent does not oversell it on your behalf.
    - **Two bundled references**: the task file reference and the troubleshooting
      guide, complete, so the Skill works with no network.
    
    ## Skill or MCP server, and why qikly has both
    
    **Skip this if you are not using the MCP server.** The Skill works on its own.
    This is here because the two look interchangeable and are not: they answer
    different halves of the same problem, and neither replaces the other.
    
    | | MCP server | Skill |
    |---|---|---|
    | What it ships | a running process exposing typed tools | a folder of Markdown |
    | What it gives the model | the ability to call `qikly_scaffold`, `qikly_validate`, `qikly_run`, `qikly_explain` | the judgement around those calls |
    | Setup | `qikly --install-mcp`, then host configuration | `qikly --install-skill` |
    | Can enforce a rule | **yes**, in server code | no, it is instructions |
    | Works when the agent has no shell | yes | no |
    
    **The MCP server is the one that can enforce things.** `qikly_run` is built so
    it cannot return the acceptance criteria, and a test in qikly's own suite fails
    the build if any code path lets a criterion through. That guarantee lives in
    code, and it is why the server exists.
    
    **The Skill is the one that can teach.** Which line of your spec is a decision
    and which is a consequence; that a repeating identical patch means a decision
    is in the wrong half; that a criterion naming a value your data never holds
    produces a test that passes whatever the code does. None of that is a function
    call, and an agent that does not know it will use qikly and get less out of it.
    
    **The guarantee does not rest on the Skill.** A Skill is text in a context
    window, so it can inform but not enforce. The withholding is enforced by the
    tool, in code, whether or not this Skill is ever installed. The Skill's own
    text says so, and a test asserts that it still does.
    
    ## If your agent does not pick it up
    
    Three causes, in the order to check them.
    
    **1. Your agent is not in that directory.** The Skill is per project. It is
    files on disk, so an agent started in another folder, or running in a browser
    with its own sandbox rather than on your machine, cannot see them. Start the
    agent in the directory you installed into. This is the commonest cause by some
    distance.
    
    **2. Just name it.** This always works, because it does not depend on the
    agent being told what the Skill is for:
    
    > Use the qikly skill to write tests for `src/following_distance.py`.
    
    If naming it works and the neutral request did not, the Skill is fine and the
    problem is discovery.
    
    **3. The host did not pass the description along.** Whether an agent reaches
    for a Skill unprompted depends on how much it explores before it starts typing,
    and on what the host told it. Claude Code reserves a fraction of the context
    window for the whole skill listing, 1% by default; when the listing does not
    fit, Anthropic's own bundled skills keep their descriptions and everything else
    is ranked by how often you have used it. So a skill you have never invoked can
    arrive as a bare name with nothing to match against. Raising
    `skillListingBudgetFraction` in your Claude Code settings gives the listing more
    room, and that single change turned four failed routing attempts into a clean
    one during testing.
    
    **Which is why naming it once is more than a workaround.** That ranking is by
    how often you have used each skill, decaying over about a week, so a skill you
    have never invoked sorts below every skill you have. Name it in one request and
    it moves above them, and the next neutral request has a much better chance of
    finding it by itself. The first invocation is the only hard one.
    
    ## Checking that it works
    
    A Skill either loads or it does not, and it never tells you which, so it is
    worth five minutes once. **If you only do one of these, do number four:** the
    others check that the Skill arrived, and that one checks that it is right.
    
    **1. The files are where your agent looks for them.** In the project you ran
    `qikly --install-skill` in:
    
    ```bash
    ls .claude/skills/qikly/SKILL.md        # macOS, Linux
    dir .claude\skills\qikly\SKILL.md       # Windows PowerShell
    ```
    
    **And start your agent in that same directory.** The Skill is per project, not
    per machine: an agent started somewhere else, or running in a browser with its
    own sandbox rather than on your computer, cannot see these files and will never
    load them. That is the commonest reason a correctly installed Skill appears to
    do nothing.
    
    **2. It loads on a request that should trigger it.** Start a fresh session and
    ask for something in its territory, **naming a real module of your own** and
    not mentioning qikly:
    
    > Write tests for `src/following_distance.py` that would actually catch a bug in it.
    
    The agent should mention qikly, or the decisions-and-consequences split,
    unprompted. If it does not, see [when an agent does not pick it
    up](#if-your-agent-does-not-pick-it-up) below before changing anything.
    
    **3. It does not load when it should not.** Ask something unrelated, such as
    "rename this variable everywhere", and it should stay quiet. A Skill that loads
    for everything costs context on every request.
    
    **4. It gives the right answer to the question that matters.** This is the one
    to do if you do only one. Ask:
    
    > My spec says "warn when following distance breaks the two-second rule", and
    > the acceptance criteria say a headway of exactly 2.00 s does not warn. Is
    > that the right split?
    
    The answer should be no, and the reason should be that "breaks" can be read as
    "below" or "at or below", so the boundary is a decision and belongs in the
    requirements. That is the single most valuable thing in the Skill, and if it
    comes back wrong the rest does not matter much.
    
    **5. It does not claim more than the tool does.** Ask what guarantees the
    coding agent never sees the criteria. The answer should point at qikly's code
    and its build-failing test, not at the Skill.
    
    ## Keeping it current
    
    **An installed Skill does not update itself, and until 0.5.4 nothing told you.**
    `pip install --upgrade qikly` replaces the package; it cannot touch a folder
    copied into your project, so after an upgrade you can be following instructions
    that name a different set of commands. From 0.5.4 any qikly command says so
    when it notices:
    
    ```
    note: the qikly Skill in .claude\skills\qikly is older than this qikly, so
    it describes a different set of commands. `qikly --install-skill --force`
    replaces it and keeps a copy of the old one.
    ```
    
    The path is printed with your platform's own separator, so it reads with forward slashes on macOS and Linux.
    
    `--force` keeps a timestamped copy of what it replaces, so a Skill you have
    edited is recoverable. That is the whole update mechanism: qikly tells you, and
    you run one command. There is no background process and nothing phones home;
    the check is two version strings read from two files on your disk.
    
    The Skill's version is its own and moves when its instructions move, not when
    qikly releases, so a release that does not touch it produces no notice.
    
    ---
    
    The Skill itself lives at
    [`src/qikly/skills/qikly/`](https://github.com/gal-a/qikly/tree/main/src/qikly/skills/qikly)
    and ships inside the installed package, so `--install-skill` works from
    `pip install qikly` as well as from a clone.
    
    **A stale Skill is worse than no Skill**, because an agent quotes it with
    confidence and the reader has no way to tell. So it is updated whenever qikly
    changes in a way it describes: a new or renamed flag, a change to which
    commands are free, a re-measured convergence figure, a new failure mode worth
    carrying.
    
    Half of it cannot go stale on its own. The two bundled references are asserted
    byte-identical to `docs/TASK_FILE_REFERENCE.md` and `docs/TROUBLESHOOTING.md`
    by a test, so a drifted copy fails the build. The prose in `SKILL.md` has no
    such test, which is why it is on the release checklist instead.
    
  • TASK_FILE_REFERENCE.md 21.9 KB
    # Task file reference
    
    The parts of working on your own data that you look up rather than read
    through. The path to a first run is [the quick start](https://github.com/gal-a/qikly/blob/main/docs/QUICK_START_ON_YOUR_OWN_DATA.md); this is what it
    deliberately leaves out.
    
    ## You probably do not have to write the task file by hand
    
    The criteria usually exist already, in a feature page or a ticket, and the
    interface exists in the code. qikly reads both.
    
    ```bash
    # a markdown page, a ticket export, or a .feature file
    qikly --criteria-from feature.md --task-id MY_TASK
    
    # straight from Jira: needs JIRA_BASE_URL, JIRA_EMAIL, JIRA_API_TOKEN
    qikly --criteria-from-jira PROJ-412 --task-id MY_TASK
    
    # both halves at once: criteria from the page, interface from the module
    qikly --scaffold src/metrics/band.py --from-doc feature.md
    ```
    
    Bullet lists, a headed `Acceptance Criteria` section and Gherkin `Scenario:`
    blocks are all understood. Your page stays the source of truth and nobody
    retypes anything.
    
    **One section is never filled for you: `requirements`.** The coding agent reads
    it, and a feature page usually restates its own acceptance criteria in the
    prose above them, so lifting requirements across would hand the criteria to the
    one agent that must never see them. `qikly --validate` warns if what you write
    there restates a criterion.
    
    ## Which command depends on which parts you already have
    
    A task file is one YAML file with three parts, and the split above is a split
    between them:
    
    1. **`requirements`** what the code must do, in the words a person would use.
       The coding agent reads this.
    2. **`interface`** the contract, and a description rather than code: the
       function signatures and the dotted path where the module will live. Both
       agents read it, and neither is handed an implementation to read from it.
       When the integration and system tests are written there is not one yet.
    3. **`acceptance_criteria`** what counts as correct, each one checkable and
       naming its boundary value. **Only test generation reads this.**
    
    "Spec" below means 1 and 2 together, which is what the coding agent is given.
    A tick means you already have that part.
    
    One thing the three parts do not say, and it matters: **test generation never
    reads the implementation either.** Integration and system tests are written
    before any code exists, from the specification alone. The unit stage is the
    single exception, written last from the code that just cleared the earlier
    stages, because unit tests have to name real functions.
    
    | Where you are starting | #1 | #2 | #3 | Run | What happens |
    |---|:-:|:-:|:-:|---|---|
    | Before anything else: see what is withheld | | | | `qikly --explain <MY_TASK>`<br>e.g. `qikly --explain CALC_TAX` | Prints a task file twice, once as each agent receives it, and the difference between them. No API key, no model call, about a second. **You get:** the acceptance criteria on one side and the same file with them cut out on the other, which is the claim everything else rests on. Add `--html` for the same as a page you can share. |
    | Just looking | | | | `qikly --demo` | A bundled task end to end in a throwaway folder. Thirty seconds, under a cent. **You get:** a working implementation, three test suites, and the full record of every FIX and PATCH, in a directory you can delete. |
    | Code someone else wrote, and you want **that code** verified | | Y | | `qikly --scaffold <MY_MODULE>.py` | Scaffold reads the real signatures out of the file you point it at and fills in **#2** for you. **#1** and **#3** stay yours to write: criteria read out of an implementation can only describe what that implementation already does, which is a bar it passes by construction. **You get:** one task file that tests the code you already have. Add `--fresh` for one that writes a fresh implementation of the same interface instead. |
    | You know what it must do, not yet how to check it | Y | | | `qikly --init` | Creates the directory layout and one starter task to edit. Its criteria show the habit that matters most: name the value, not the quality. "100 is accepted and 101 is rejected" forces a test at the boundary; "amounts must be reasonable" does not. **You get:** a task file to fill in, with your fixtures where a run will look for them. |
    | Same, but you want a first draft of the bar | Y | Y | | `qikly --tasks <MY_TASKS>`<br>`--generate-criteria` | Drafts **#3** from **#1** alone, then runs. **You get:** a first draft of the bar written into your task file for you to correct, plus the implementation and suites. |
    | You have written all three | Y | Y | Y | `qikly --tasks <MY_TASKS>` | Everything you wrote is used, and nothing is drafted on your behalf. **You get:** an implementation, integration, system and unit suites, a convergence report, and a run summary recording the model and settings that produced them. |
    | You have all three but doubt they agree | Y | Y | Y | `qikly --check-criteria`<br>`--tasks <MY_TASKS>` | One model call asking whether any implementation could satisfy the description, **#1** and **#3** at once, and whether any two of **#3** agree with each other. Advisory, and exits non-zero on a contradiction so a pipeline can gate on it. **You get:** a list of the pairs that cannot both hold, before spending a stage budget on them. Two criteria setting different numbers on the same quantity are always reported, since that is a typo rather than a tighter bar. |
    | A previous run stopped before finishing | Y | Y | Y | `qikly --tasks <MY_TASKS>`<br>`--resume` | Generating the tests and the first implementation already cost model calls, and they are still on disk. This keeps them and picks up where it stopped, instead of paying for them twice. **You get:** the same outputs as a full run, without paying for the parts already built. |
    
    `<MY_TASKS>` is one task_id or several separated by commas. A task_id is a
    filename under `inputs_private/config/tasks/` without the `.yaml`:
    `--tasks CALC_TAX`, `--tasks CALC_TAX,MERGE_SALES`, or omit it to run every
    task found. `<MY_TASK>`, singular, takes exactly one.
    
    `QIKLY_MAX_CALLS=200 qikly` stops at a call limit rather than a bill.
    
    `--scaffold` reads the module path and the real signatures of every public
    function straight out of the file, because they are already there. It leaves
    `requirements` and `acceptance_criteria` for you, and that is deliberate:
    criteria derived from an implementation can only describe what that
    implementation already does, and a bar that agrees with the code by
    construction is the exact failure this tool exists to prevent.
    
    ## Writing a task by hand
    
    Nothing is written into the package, and nothing is written into your source
    tree.
    
    ### 1. Make the two directories
    
    Anywhere you want to work. The presence of `inputs_private/` is what marks a
    directory as your project.
    
    ```bash
    mkdir -p inputs_private/config/tasks
    mkdir -p inputs_private/data/MY_TASK
    ```
    
    Or let `qikly --init` create both, plus a starter task to copy.
    
    Until one of those exists, there is nothing marking your directory, and the
    fallback in [Where things live](#where-things-live) applies. From a
    `pip install -e` checkout that fallback finds the checkout itself, so a run
    started in an empty directory writes its outputs there instead of where you
    are standing. Make the directory first, or set `QIKLY_PROJECT_ROOT` to say
    exactly where you mean.
    
    ### 2. Drop your fixture data in
    
    Plain input files, whatever your code should read. CSV, JSON, JSONL, anything.
    
    ```bash
    cp ~/somewhere/orders_jan.csv inputs_private/data/MY_TASK/input_01.csv
    cp ~/somewhere/orders_feb.csv inputs_private/data/MY_TASK/input_02.csv
    ```
    
    Names are up to you, but they must match what you write in the task's
    `inputs:` list below. The generated program opens these **by literal relative
    path from your project directory**, so the path in the task file is the path
    that gets executed. That is also why the bundled fixtures are copied into
    `inputs_private/data/` on first run rather than resolved from inside the
    package: the generated code has no way to ask where the package lives.
    
    Fixtures are never overwritten once present, so an edited file stays edited.
    
    ### 3. Write the task file
    
    `inputs_private/config/tasks/MY_TASK.yaml`. The filename must match `task_id`.
    
    ```yaml
    task_id: "MY_TASK"                    # letters/digits/underscore, not starting with a digit
    task_name: "Order line-item tax"
    
    description: "Read two CSV files of order line items, validate them, compute
      tax per line, and write the result to a single JSON output alongside a
      reason for every rejected line."
    
    inputs:                               # literal paths, opened by the generated code
      - "inputs_private/data/MY_TASK/input_01.csv"
      - "inputs_private/data/MY_TASK/input_02.csv"
    
    outputs:
      - "outputs/data/MY_TASK/output.json"
    
    interface:                            # what test generation targets
      module: "outputs.agent_src.code.MY_TASK.calc"
      integration_functions:
        - "extract(input_path) -> list[dict]  # reads one input file, returns raw rows"
        - "transform(rows) -> dict  # validates and computes; returns {\"accepted\": [...], \"rejected\": [...]}"
        - "load(data, output_path) -> None  # writes the result as JSON"
      system_entrypoint: "run_calc(input_paths, output_path) -> None  # extract each path, then transform -> load"
    
    requirements:                         # THE DECISIONS. The coding agent sees only this.
      - "Read both CSV files listed in inputs and combine their rows before validation"
      - "Validate each row: order_id, item_price, quantity, tax_rate"
      - "Apply strict, real-world data-quality validation; reject anything malformed or out of range"
      - "For each valid row compute subtotal, tax owed, and line total as currency amounts"
      - "A rejected row is not silently dropped: record it with a brief, specific reason"
      - "Write a single JSON object with two keys, \"accepted\" and \"rejected\""
    
    acceptance_criteria:                  # THE CONSEQUENCES. Withheld from the coding agent.
      - "All computed currency amounts are rounded to two decimal places using round-half-up, not banker's rounding and not truncation"
      - "For every accepted row, the reported total equals the reported subtotal plus the reported tax, exactly, to the cent"
      - "A tax_rate of exactly 0 is valid: the computed tax is 0.00 and the total equals the subtotal"
      - "Each rejected row names the specific field that caused rejection, not a generic message"
    ```
    
    `interface.module` is the dotted path the generated tests will import. With
    no seed, or with a single-file seed, that is
    `outputs.agent_src.code.<task_id>.<name>`: the implementation is written
    there, so pick the final component freely and the rest is fixed by where
    outputs live.
    
    **A package seed is the exception**, because the package keeps its own name
    and becomes importable by it: write `interface.module: "mypkg.pricing"`, the
    path your own code already uses. See "Seeding a package" below.
    
    ### 4. Run it
    
    ```bash
    qikly --tasks <MY_TASKS>        # or: python run.py --tasks <MY_TASKS>
    ```
    
    Discovery is automatic; there is no registry to update. Results land in
    `outputs/`, and `outputs/reports/iterations/MY_TASK_<timestamp>_report.html`
    is the place to start reading.
    
    ## Bringing acceptance criteria you have already written
    
    Most teams have not got a blank page here. If you work in Jira, Linear, Azure
    DevOps or a design doc, the rules are usually already written down, because the
    process asks for them before any code is cut. A ticket routinely looks like
    this:
    
    ```
    PROJ-412  Merge overlapping sales exports
    
    Description
      Combine two CSV exports into one file...
    
    Acceptance Criteria
      - A transaction in both files at the same amount appears once
      - A negative or missing amount is rejected, naming the field
      - Dates must be YYYY-MM-DD
    ```
    
    Those bullets are exactly what `acceptance_criteria` wants. Save the ticket to
    a file and read them out:
    
    ```bash
    qikly --criteria-from ticket.md                    # print as YAML
    qikly --criteria-from ticket.md --task-id MY_TASK  # write into that task
    qikly --criteria-from ticket.md >> inputs_private/config/tasks/MY_TASK.yaml
    ```
    
    It understands plain bullet lists, an "Acceptance Criteria" heading in a longer
    document, and Gherkin `Scenario:` blocks with Given/When/Then. Only the criteria
    section is read, so pasting a whole ticket does not turn its description into
    part of the bar. Only YAML goes to stdout, so the third form above appends a
    valid block.
    
    **It will not invent criteria from prose.** A file with no list and no scenarios
    returns nothing and says so. A rule that nobody wrote is precisely the invented
    standard this tool exists to argue against, and once it is in the file it looks
    like every other line.
    
    There is no API token and no vendor integration involved. Copying the ticket
    into a file is the whole of it.
    
    **Read what comes out before you run.** Criteria lifted from a ticket are a
    draft: tickets are written for people, who fill in gaps that a test cannot. The
    criteria are the standard everything else is judged against, so they are worth
    a minute of your attention.
    
    ## Supplying your own acceptance criteria, code or tests
    
    The loop takes three inputs. **Each one can be yours or generated,
    independently and in any combination.**
    
    | Input | Default | To supply your own |
    |---|---|---|
    | **Acceptance criteria** | Yours | Already the default: write `acceptance_criteria` in the task file, as above, or lift them from a ticket with `--criteria-from` (below). Omit it and add `--generate-criteria` to have a first draft written for you instead. |
    | **Implementation** | Generated | `seed.implementation` in the task file. |
    | **Test suites** | Generated | `seed.tests`, per stage. |
    
    ### Auto-generating acceptance criteria
    
    A task with no `acceptance_criteria` still runs, with a warning rather than an
    error, because running one deliberately is a legitimate thing to do. What you
    lose is the point of the exercise: test generation has only `requirements` to
    work from, the coding agent has nothing sharper to fail against, and the run
    usually converges on the first attempt without exercising the loop at all.
    
    `--generate-criteria` writes a first draft from the requirements alone into
    `inputs_private/config/tasks/<task_id>.yaml` before the run starts. It is
    opt-in, it never touches a task that already has criteria, and it says what it
    wrote rather than editing your files quietly. Pointed at a bundled example it
    writes your own overriding copy and leaves the packaged original alone.
    
    **A generated bar is a draft, not ground truth.** It was written from the same
    requirements the coding agent reads, so a case it did not think to demand is
    not being withheld from anyone: the two halves agree because they came from one
    source, which is the failure mode this whole tool argues against.
    `--compare-criteria` scores a generated bar against yours when you want that
    difference measured rather than assumed.
    
    The optional `seed:` block:
    
    ```yaml
    seed:
      # A file or a directory, copied into outputs/agent_src/code/<task_id>/.
      # A single file keeps its own name, which must match interface.module.
      # A directory is treated as a package: see "Seeding a package" below.
      implementation: "seeds/MY_TASK/calc.py"
    
      # Per stage. Seeding a stage suppresses generation for that stage only.
      tests:
        integration: "seeds/MY_TASK/test_integration.py"
        unit: "seeds/MY_TASK/unit/"
    ```
    
    Paths are relative to your project directory. Both keys are optional.
    
    **`seed.implementation` is how you point this at code you already have.** The
    run skips generating a first implementation and goes straight to testing and
    repairing yours. `--scaffold` writes this block for you by default, and `--fresh` writes a
    task without it, for a new implementation of the same interface. **`seed.tests` keeps a suite you already trust**, so the loop
    repairs the code against your tests rather than its own. Mixing works and is
    often what you want: seed the integration stage with your suite and let the
    tool generate unit tests against whatever code results.
    
    Three things to know:
    
    - **Seeded test suites are checked before the run starts.** Every file must
      parse, and at least one must be named `test_*.py` and contain a `def test_*`
      function. A problem raises immediately rather than retrying, since there is
      no second sample to draw from a file you wrote.
    - **Seeds are installed after the workspace reset, not instead of it.** Every
      run still begins from one declared state, so repeated runs stay comparable
      and no run inherits the previous one's residue.
    - **A seeded run measures something different from an unseeded one.** Do not
      pool them in a single rate. The orchestrator prints a NOTE on every seeded
      run to keep that visible.
    
    ### Seeding a package, when the implementation is more than one module
    
    **New in 0.5.5.** Point `seed.implementation` at a directory and it is treated
    as a package: it is copied in **under its own name**, and the task's code
    directory is put on the path for the test run, so the package resolves by the
    name your code already uses.
    
    ```yaml
    interface:
      module: "mypkg.pricing"   # the module under test, by its real import path
    
    seed:
      implementation: "mypkg"   # the package it lives in, copied in whole
    ```
    
    Inside `mypkg/`, write imports exactly as you already do. All three shapes
    work: `from mypkg.money import to_cents`, `from .money import to_cents`, and
    `from .utils.rounding import half_up`. No `__init__.py` is required, and one
    that is there is kept.
    
    **Every module in the package is visible to the coding agent and every one is
    repairable**, and a single patch may change more than one of them. That is the
    difference the package form makes: with a single-file seed, a defect in a
    helper is found by the tests and cannot be fixed, and the run tells you so
    rather than working around it.
    
    **The directory you name is the boundary.** qikly does not follow imports and
    decide for itself which of your files an agent may rewrite, because the
    transitive closure of a real package has no natural edge and "it rewrote a
    shared module I never named" is a worse outcome than naming a folder. So put
    inside the seed what you want worked on, and leave a vendor library or a
    module you do not want touched outside it.
    
    Data files inside the package are copied too, since your code may open them.
    `__pycache__` and `.pyc` files are not: they are stale copies of the very
    modules the run is about to rewrite.
    
    Four limits worth knowing before you start:
    
    - **Before 0.5.5 a seeded directory was flattened**, dropping the folder's
      name. If you wrote a task against that behaviour, `interface.module` needs
      the package name adding to it.
    - **Modules that import each other circularly at the top level fail**, the
      same way they do in plain Python. This is not something a run can repair.
    - **The package's name may not be a standard-library module's name.** A
      package called `json` would shadow the real one for everything the run
      imports, so a seed naming one is refused with a message rather than
      discovered halfway through a stage. Names that clash with an *installed
      third-party* package are not checked, because what is installed varies by
      environment: if your package is called `yaml` or `requests`, rename it or
      seed the single module instead.
    - **`seed.implementation` must name the folder itself**, not a path that
      resolves to `.` or `..`. Those are refused too, because the install would
      land outside the task's own directory.
    
    The full matrix of import shapes, including the ones that do not work, is
    pinned in `tests/test_multi_module_seed.py`.
    
    ## Where things live
    
    Task specs and shared defaults are read from `inputs_private/` in your project
    directory if present, otherwise from the copies bundled inside the package, so
    a fresh install runs immediately. Resolution is **per file**: dropping one task
    spec into `inputs_private/config/tasks/` overrides exactly that task and leaves
    everything else in place. Nothing is ever written back into the package.
    
    | Path | Contents |
    |---|---|
    | `config/tasks/<task_id>.yaml` | One task, as above. |
    | `data/<task_id>/` | That task's fixture data. |
    | `config/settings.yaml` | Retry budget, stage order, patch size limit. A private copy is overlaid section by section, so state only what you change. |
    | `agent_defs/*.md` | The prompts. `code_agent.md` and `test_agent.md` are the two system prompts; the rest are per-mode fragments. Not per-task: editing these changes every task's behaviour. |
    
    ## Proposing fixture rows
    
    A criterion no input row can trigger produces a test that passes whatever the
    code does. Across this project's own measurements roughly two thirds of
    deliberately planted faults were missed by every suite for that reason: the bar
    was unmeasurable rather than wrong.
    
    ```bash
    qikly --propose-fixtures --tasks <MY_TASKS>
    ```
    
    A separate agent reads your criteria and your fixture files and says, for each
    criterion, either `covered` or here is the smallest row that would reach it.
    The answer goes to `outputs/reports/fixture_proposals/`, laid out with each row
    printed under the criterion it exists to reach so you judge the two together.
    
    **It never edits a fixture.** To accept a row, paste it into the named file and
    append `  # proposed`. To reject one, do nothing. Two reasons for the gate,
    neither about the model being untrustworthy. A row is only right or wrong
    relative to its criterion, so it is harder to review than a sentence. And a
    fixture set that grows in whatever direction a model finds interesting stops
    resembling the data you actually process, at which point every rate measured on
    it describes a world that does not exist. The report is capped at eight
    proposals per round and prints what share of your rows a machine has written,
    so that drift is visible in aggregate rather than one plausible row at a time.
    
    You can of course add rows by hand at any time, and always could. This exists
    because noticing *which* criteria have no data behind them is the tedious part.
    
    The refinement loop does this for you on what it adds. When
    `refine_acceptance_criteria` finishes with new criteria, it asks for rows the
    same way, lists the new criteria first in the report, and logs how many have no
    data that reaches them, so a sharper bar does not arrive partly unmeasurable.
    It still applies nothing.
    
    
  • TROUBLESHOOTING.md 20.5 KB
    # When a run does not converge
    
    **If a run has not started yet, skip to [Before a run: where did my files
    go?](#before-a-run-where-did-my-files-go) at the end.** Everything above that
    section assumes a run has already failed, and the commonest reports this
    project receives are not about runs at all.
    
    A stall is a normal outcome, not a broken tool. The run exits non-zero, names
    the tests that blocked it, keeps the whole record, and ships nothing. Across
    every measurement this project has taken, roughly four runs in ten stop this
    way, and no run has ever reported success on code its own tests rejected.
    
    So the question is never "why is it broken". It is which of a short list of
    things is happening, and the list is short.
    
    ---
    
    ## Triage
    
    Match what you saw to where to look. The rows are in the order to work through
    them: the one change that moves convergence most, then the checks that settle
    what happened, then fixes to the task file, and only then more attempts.
    
    | What you saw | Section | What to do |
    |---|---|---|
    | Poor results on a provider you just set up, or on the default model | [1. Try a stronger model](#1-try-a-stronger-model) | Set `LLM_MODEL` to a mid-tier or larger model and run again |
    | Any stall, before changing the task file | [2. Read what actually blocked it](#2-read-what-actually-blocked-it) | Change nothing yet: this step decides what to change. In the run's `_report.html` timeline, byte-identical patches go to [5](#5-the-same-patch-appearing-over-and-over), an import or syntax error to [3](#3-a-collection-error-means-nothing-ran), steady progress to [10](#10-give-it-more-attempts), and different patches that never fix the same test to [1](#1-try-a-stronger-model) |
    | `0 passed, 0 failed, 1 error` | [3. A collection error means nothing ran](#3-a-collection-error-means-nothing-ran) | Make `interface.module` and the declared signatures match what the tests import |
    | Everything suddenly worse than last week | [4. Check nothing is set that you have forgotten](#4-check-nothing-is-set-that-you-have-forgotten) | Run `qikly --trends --by week` and look for a setting that changed, such as `criteria_per_batch` |
    | The same test failing every iteration, no progress | [5. The same patch appearing over and over](#5-the-same-patch-appearing-over-and-over) | Move the restriction into `requirements`, or widen the criterion to match reality |
    | Stopped with a message naming two tests, each fix for one breaking the other | [6. Two generated tests disagree](#6-two-generated-tests-disagree) | Compare the two tests with the acceptance criteria. If one contradicts a criterion, run again without `--resume` so the suites are written and checked again, and leave a correct spec alone |
    | A stage spends its whole budget and never gets closer | [7. Check the criteria and requirements do not contradict each other](#7-check-the-criteria-and-requirements-do-not-contradict-each-other) | Run `qikly --check-criteria --tasks <MY_TASKS>` and correct whichever statement is wrong |
    | Tests check arbitrary values rather than the boundary | [8. Check the criteria name values, not adjectives](#8-check-the-criteria-name-values-not-adjectives) | Run `qikly --validate` and rewrite each flagged criterion as a value: "100 is accepted and 101 is rejected" |
    | A test passes whatever the code does | [9. Check your fixtures can reach every criterion](#9-check-your-fixtures-can-reach-every-criterion) | Run `propose_fixtures` and add the input rows it suggests |
    | Steady progress, then the budget ran out | [10. Give it more attempts](#10-give-it-more-attempts) | Raise `orchestrator.max_retries_per_stage` (default 10), only when the report shows progress |
    | **Nothing has run yet, and files you were told about are missing** | [Before a run](#before-a-run-where-did-my-files-go) | Read the first line of the command's output: it names the directory it worked in, and says when the project root is somewhere else |
    | **You are in a `demo/throwaway_<timestamp>/` folder** | [Before a run](#before-a-run-where-did-my-files-go) | That is a throwaway copy. Start your own project somewhere else |
    | `--score-code` says the suite does not pass | [Before a run](#before-a-run-where-did-my-files-go) | Usually pytest collected no tests at the path given to `--score-tests` |
    | **On macOS, no patch ever applies, on any task** | [12. On macOS, no patch ever applies](#12-on-macos-no-patch-ever-applies) | `brew install gpatch`. The system `patch` is BSD and rejects the options qikly sends |
    | Integration and system pass, unit does not | [11. Expect the unit stage to be where it fails](#11-expect-the-unit-stage-to-be-where-it-fails) | Expected. Accept it, or leave the unit stage out with `orchestrator.test_order` |
    
    ---
    
    ## 1. Try a stronger model
    
    This moves convergence more than anything else here, and it is one environment
    variable.
    
    ```bash
    export LLM_MODEL=gpt-4o          # or a larger model on your provider
    qikly --tasks <MY_TASKS>         # e.g. --tasks CALC_TAX,MERGE_SALES
    ```
    
    Every convergence figure in this project was measured on
    `gemini-3.5-flash-lite`, a deliberately small and cheap model chosen so that
    sweeps of hundreds of runs were affordable. Treat those figures as a floor.
    
    An entry-level model on any provider may stall on a task a mid-tier one clears
    comfortably. If you are evaluating qikly, evaluate it on a model you would
    actually ship behind.
    
    **What a bigger model buys, and what it costs.** On one CALC_TAX run,
    `claude-sonnet-5` generated 44 tests against `gemini-3.5-flash-lite`'s 24, and
    its suite rejected code that Gemini's suite accepted, on five tests, while
    Gemini's suite accepted its code entirely. A stricter bar, in other words. It
    also took 403 seconds against 31, and cost \$0.81 against \$0.005.
    
    That trade is worth making deliberately rather than by default. A reasoning
    model produces thinking tokens you are billed for and wait on, which is why the
    cheapest model is the default here and why every published figure was measured
    on it: a 400-run sweep costs about \$3 on the default and roughly \$320 on a
    reasoning model.
    
    **Use the cheap model to measure and the expensive one to work.** If you need a
    convergence rate, take it on the default. If you need the strictest bar for one
    important specification, pay for it once. One run of each is an anecdote, not a
    comparison; the figures above are a single run per model.
    
    ### The default model is not the same size on every provider
    
    <a id="provider-defaults"></a>Set no `LLM_MODEL` and each provider gets its own
    default, and they are not the same class of model. This is the first thing to
    check when a run takes far longer on one provider than another:
    
    | Provider | Default model | What that means for a run |
    |---|---|---|
    | Gemini | `gemini-3.5-flash-lite` | Small, cheap, no reasoning step. Every published figure here was measured on it. A demo task runs in well under a minute |
    | OpenAI | `gpt-4o` | Mid-tier. Slower and dearer than the Gemini default |
    | Anthropic | `claude-sonnet-5` | A reasoning model. qikly sends no thinking configuration, and on this model that means adaptive thinking runs by default, so every call thinks before it answers |
    
    The Anthropic default is the one that surprises people. The same demo task that
    finishes in under a minute on the Gemini default has taken around sixteen
    minutes on it, for the same seven or so model calls. Nothing is wrong when that
    happens: you are watching a reasoning model think, and it produces a stricter
    suite for it.
    
    **For a like-for-like comparison with the Gemini default, name the model:**
    
    ```bash
    export LLM_MODEL=claude-haiku-4-5    # Windows PowerShell: $env:LLM_MODEL = "claude-haiku-4-5"
    ```
    
    Haiku is the closest Anthropic analogue to a flash-lite class model, and qikly
    sends no thinking budget, which that model needs before it will think at all.
    So a run on it spends no time or tokens on reasoning.
    
    Keep `claude-sonnet-5` when you want the stricter bar, and expect the run to
    take minutes rather than seconds. What qikly does not yet expose is the middle
    setting: the API takes a reasoning effort level, and a way to ask for less of
    it without changing model would make this a dial rather than a switch.
    
    **If it looks stuck**, it probably is not. A single call can legitimately run
    for minutes on a reasoning model. After ten seconds of silence a run starts
    saying so, one line every fifteen: `[patch] still waiting on the model, 45s`.
    Those lines are the difference between slow and stalled, and
    `QIKLY_NO_PROGRESS=1` turns them off. One
    call is abandoned after `QIKLY_REQUEST_TIMEOUT` seconds, 300 by default, which
    was chosen when the slowest observed call was well under a minute; on a
    reasoning model consider raising it, or a slow-but-working call is thrown away
    and retried from scratch.
    
    ## 2. Read what actually blocked it
    
    Every run writes a timeline:
    
    ```
    outputs/reports/iterations/<task>_<timestamp>_report.html
    ```
    
    Open it in a browser. It shows every iteration, the FIX reasoning and the PATCH
    diff for each failure, and, most usefully, **which patches applied cleanly and
    changed nothing.** A run full of those is not a run that needs more attempts.
    It is [3](#3-a-collection-error-means-nothing-ran) or
    [5](#5-the-same-patch-appearing-over-and-over).
    
    ## 3. A collection error means nothing ran
    
    ```
    [MY_TASK] [stage 1/3] [iteration 3] 0 passed, 0 failed, 1 error, 0 skipped
    ```
    
    No test failed, because no test ran. The module could not be imported. This is
    a different problem from a wrong answer, and until it is fixed nothing else can
    be assessed.
    
    The FIX prompt is told this explicitly, and the report carries the underlying
    `ImportError` or `SyntaxError`. The usual cause is a mismatch between what
    `interface` declares and what the agent wrote, so check that
    `interface.module` and the declared function signatures are exactly what the
    tests should be importing.
    
    ## 4. Check nothing is set that you have forgotten
    
    ```bash
    qikly --trends --by week
    ```
    
    Convergence per task over time, from the run summaries already on disk. Every
    period names the model and settings behind it, and a period where those changed
    is marked.
    
    This exists because of a specific, expensive mistake. `criteria_per_batch` in a
    settings file controls how many acceptance criteria a single test-generation
    call is shown. At `0` one call sees the whole bar. At `4` the bar is split into
    batches and each gets its own call, so a long bar produces roughly three times
    as many tests, and every run has three times as much to satisfy.
    
    Left set from an earlier experiment, it made convergence appear to collapse
    across nine tasks at once. Half a day went into diffing prompts, specs and
    provider parameters before anyone looked at the override.
    
    **A rate belongs to a tool, a model and a configuration together.** A rate that
    moved when the configuration moved is not a finding.
    
    ## 5. The same patch appearing over and over
    
    Identical diffs, not merely a repeated failure, is a specific signature: the
    model is fighting something it correctly knows about the world.
    
    Restrict a real-world field to an artificial subset, say three valid street
    suffixes, and the model will keep widening the restriction back. Not out of
    disobedience. Every piece of its training agrees that "Boulevard" is a street
    suffix, and your criterion is the outlier. At temperature zero this does not
    converge slowly; it does not converge at all.
    
    **Fix:** widen the criterion to match reality, or move the restriction into
    `requirements`, where the coding agent can read it and treat it as a given
    rather than as an error to correct.
    
    To confirm it, compare successive diffs under
    `outputs/logs/patches/<task>/<timestamp>/`. Byte-identical patches mean this.
    Different patches that never resolve the same test mean something else: a bug
    that needs more than the failure text to fix, which is [1](#1-try-a-stronger-model).
    
    ## 6. Two generated tests disagree
    
    ```
    Stopped on stage 'system' after 4 attempts: the last three patches alternated
    between the same two diffs. [...] The failing tests alternate between
    test_run_headway_two_second_rule_warning and test_integration_pipeline_flow in
    the 'system' and 'integration' suites: each fix for one breaks the other [...]
    ```
    
    Each patch makes one test pass and the other fail, because the two tests expect
    different results for the same input. No code can pass both, so more attempts
    cannot help, and the fault is in the tests rather than the specification. In
    the run behind this section, an integration test warned at exactly 2.00 seconds
    of headway and a system test did not, against a criterion saying exactly 2.00
    seconds raises no warning.
    
    The same mistake can also be made identically in both suites. Then they agree
    with each other and still contradict the criterion, the run fails without this
    message, and the place to look is the same: each test's comparison at every
    limit the criteria state.
    
    An opt-in check, `check_suites: true` under `test_generation` in settings,
    looks for these before any code is written and rewrites a suite once. Measured
    on one task it found every wrong suite but did not raise convergence, and a
    wrong finding once led a correct suite to be rewritten wrong, so it is off by
    default.
    
    **Fix:** compare the two named tests with the acceptance criteria. If one
    contradicts a criterion, run again without `--resume`, so the suites are written
    and checked again. Do not change a specification that is already right: this is
    the one stall where the spec is not the problem.
    
    To check suites already on disk without a run, one model call per task:
    
    ```bash
    python -m qikly.orchestrator.tuning.check_suites --tasks <MY_TASKS>
    ```
    
    ## 7. Check the criteria and requirements do not contradict each other
    
    ```bash
    qikly --check-criteria --tasks <MY_TASKS>
    ```
    
    One model call per task, and it changes nothing. A criterion that no
    implementation could satisfy alongside the requirements produces a stage that
    spends its entire budget discovering that the slow way. It exits non-zero on a
    contradiction, so a pipeline can gate on it.
    
    ## 8. Check the criteria name values, not adjectives
    
    ```bash
    qikly --validate
    ```
    
    Free, offline, and it flags criteria written as adjectives.
    
    > "Reject large amounts" invites a test at some arbitrary large number.
    > "100 is accepted and 101 is rejected" forces a test at the boundary.
    
    This is the highest-leverage habit in writing a bar. A suite that never tests a
    boundary cannot catch an error at that boundary, no matter how many other cases
    it covers, and off-by-one at a boundary is among the oldest defect classes in
    software.
    
    `--validate` also catches the quiet structural mistakes: `acceptance_criteria`
    written as one long string instead of a list, a `task_id` that disagrees with
    its filename, and fixture paths that do not resolve. Each of those otherwise
    surfaces twenty minutes and several dollars into a run.
    
    ## 9. Check your fixtures can reach every criterion
    
    ```bash
    qikly --propose-fixtures --tasks <MY_TASKS>
    ```
    
    A criterion that no input row can trigger produces a test that passes whatever
    the code does. The bar is not lower; part of it is absent.
    
    Eight of the ten tasks bundled with qikly had at least one before this was run
    on them, from one in `CALC_CALENDAR` to seven of thirteen in `MERGE_CONTACTS`.
    Assume yours do too.
    
    It writes proposals to a file and never edits your data.
    
    ## 10. Give it more attempts
    
    `orchestrator.max_retries_per_stage` in `config/settings.yaml`, default 10.
    
    Worth raising when the report shows steady progress that simply ran out of
    room. Not worth raising when it shows the same patch repeating: that run will
    fail identically with a hundred attempts, and cost ten times as much doing it.
    
    ## 11. Expect the unit stage to be where it fails
    
    About twenty points of the gap between "passes integration and system" and
    "passes everything" is the unit stage, consistently, across every sweep this
    project has run.
    
    The reason is structural rather than mysterious: the unit suite is the largest,
    runs last, and is the only one written with sight of the implementation.
    [design_2_performance.md](https://github.com/gal-a/qikly/blob/main/docs/design_2_performance.md#nearly-the-whole-gap-between-those-two-numbers-is-the-unit-stage)
    has the full explanation.
    
    If behavioural verification is what you need, `orchestrator.test_order` in
    settings can leave it out.
    
    ## 12. On macOS, no patch ever applies
    
    Every generated diff fails, on every task, from the first iteration, for a
    reason that reads like the model's fault and is not.
    
    The system `patch` on macOS is BSD, and it rejects the options qikly sends.
    Install GNU patch and the same run goes through:
    
    ```bash
    brew install gpatch
    ```
    
    Nothing else changes. If patches apply on one machine and fail on all of them
    on another, this is the first thing to check.
    
    ---
    
    ---
    
    ## One run is an artifact, not a rate
    
    The same task with the same seed converges on some runs and not others. Before
    concluding anything about a task, a model or a setting:
    
    ```bash
    python -m qikly.orchestrator.run_all --tasks <MY_TASKS> --repeat 10 \
        --skip-eval --skip-refine
    ```
    
    That writes an aggregate with a confidence interval instead of a pass count.
    Ten runs is usually enough to tell a real difference from noise, and it is
    worth knowing that at n=10 the intervals are wide: this project has measured
    the same unchanged task at 72% and then 90% on consecutive sweeps.
    
    If a change looks like an improvement after one run, it is not yet evidence of
    anything.
    
    ---
    
    ## Before a run: where did my files go?
    
    None of the sections above apply if nothing has run yet, and this is the
    question that arrives most often.
    
    ### It said it created files and they are not there
    
    They almost certainly are, somewhere you did not look. qikly moves to a
    resolved **project root** when it starts, which can be a different directory
    from the one you are standing in, and `--init`, `--example` and `--scaffold`
    write relative to one of those two.
    
    Every command now opens with a line that settles it:
    
    ```
    qikly 0.5.4  2026-09-27 11:27:18  run in C:\Users\you\my-project
      project root is elsewhere: C:\some\other\place
      (set QIKLY_PROJECT_ROOT to choose it, or cd there)
    ```
    
    The second and third lines appear only when the two differ. If you see them,
    that is your answer. If you are on an older version that does not print them,
    upgrade, or search for one of the files by name:
    
    ```powershell
    Get-ChildItem $HOME -Recurse -Filter "MY_METRICS*" -ErrorAction SilentlyContinue | Select-Object FullName
    ```
    
    To pin the project explicitly rather than let it be inferred, set
    `QIKLY_PROJECT_ROOT` to the directory you mean.
    
    ### You are standing in the demo's folder
    
    `qikly --demo` runs in a throwaway `demo/throwaway_<timestamp>/` directory so it cannot
    touch anything of yours, which also means **everything in it goes when you
    delete the folder, and nothing in it is yours**. Somebody who has just watched
    the demo work is standing in something that looks exactly like a working
    project, and the obvious next move is to start theirs there.
    
    qikly now refuses, names the folder, and says where to go instead. If you
    deliberately kept that directory and want to work in it, delete the
    `.qikly-demo` marker inside it and the refusal stops.
    
    ### The version you are running is not the version you installed
    
    Two things can disagree. `qikly --version` reports what actually runs;
    `pip show qikly` reports metadata that an interrupted or repeated upgrade can
    leave stale. Trust `--version`. If they disagree, clean it:
    
    ```powershell
    pip uninstall qikly -y
    pip install qikly
    ```
    
    `qikly --version` also prints the package directory and the interpreter, which
    is what to check when a flag the documentation describes does not exist.
    
    ### `--score-code` says your suite does not pass
    
    It refuses to score a suite that does not pass your untouched code, because
    every planted fault would then fail for the reason the original does and the
    number would mean nothing. Two causes, likeliest first:
    
    - **pytest collected nothing.** The path given to `--score-tests` holds no
      tests, or none that match its discovery rules. Run pytest on that path
      yourself and read what it says.
    - **Your suite genuinely fails.** Fix that first, then score it.
    
    There is no reachability warning in this mode, unlike `--score-suite`: there is
    no task file, so there are no criteria to be unreachable. A fault that survives
    may still sit on a line no test executes at all, which is a gap in what your
    tests reach rather than in what they assert.
    
    ---
    
    ## Still stuck
    
    The run kept everything. `outputs/logs/transactions_<task>_<timestamp>.jsonl`
    is an append-only record of every test run, every FIX, every PATCH and every
    apply outcome, and it is the source of truth that the reports are rendered
    from.
    
    Issues and results are welcome:
    [github.com/gal-a/qikly/issues](https://github.com/gal-a/qikly/issues).
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related