docs
Imported from gal-a/qikly/docs.
Install
npx skills add https://github.com/gal-a/qikly/tree/main/docs
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install gal-a-qikly@llmmart
git clone https://github.com/gal-a/qikly.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole gal-a/qikly collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
qikly as an agent Skill
What this gets you: your coding agent stops needing to be told about qikly. Ask it for tests you can trust and it reaches for the tool on its own, and it knows the part that is hard to guess, which line of your specification is a decision the coder needs and which is a consequence to withhold.
A Skill is a folder of Markdown your agent reads when what you are asking matches what the Skill says it is for. No server, no configuration, no process to keep running.
Install it
First, make sure the qikly you are about to run is the current one. The Skill ships inside the package, so an old qikly writes an old Skill, and it then describes commands that do not exist yet.
pip uninstall qikly -y
pip install qikly
qikly --version
Uninstall first rather than --upgrade: an interrupted or repeated upgrade can
leave more than one version's metadata behind, and pip then reports one version
while the files on disk are another's. The clean pair takes seconds and removes
the question. Compare what --version prints against
the latest release.
Then, from the directory of the project you want the Skill in:
qikly --install-skill
That writes .claude/skills/qikly/, which is where agents look. Then open your
agent in that directory and ask for something ordinary, without mentioning
qikly:
Write tests for
src/following_distance.pythat would actually catch a bug.
src/following_distance.py stands in for a module you actually have; name a
real one, because an agent asked about a file that does not exist will spend
its answer asking you which file you meant. Every example on this page uses
following distance from radar samples, which is the worked case in the Skill
itself, so the two read together.
It should reach for qikly by itself, and it does: first time in Claude Code, Gemini CLI, Codex and Cursor, in a project with a rival testing skill installed beside this one and a request that never mentioned qikly. That is the hard version of the test, because the agent had a competing option and no hint. Your own project has more skills in it than that one did.
If it does not, see if your agent does not pick it up.
Using GitHub Copilot in VS Code? Copilot reads none of the skill folders
above. Run qikly --install-skill copilot, which writes the Skill into
.github/instructions/ along with the qikly.instructions.md file Copilot
actually opens, then use Copilot Chat in Agent mode. That path has not been
watched loading yet, so tell us if it works for you. The MCP server is the
other way in and is tested there:
docs/mcp.md.
A few options, none of them needed the first time. --dry-run shows what
it would write. --force replaces a Skill you have already edited, keeping a
timestamped copy of the old one. And --install-skill agents, cursor or
gemini write where those tools document their own skill folders, rather than
Claude's.
What is in it
- The one question: given only the requirements, could two competent
developers legitimately disagree about this line? Yes means it is a decision
and belongs in
requirements; no means it is a consequence and belongs inacceptance_criteria. - Both mistakes and their signatures, so an agent can recognise a stalled loop as a spec problem rather than a code problem.
- How to write a criterion that can be tested: name values not adjectives, name both sides of a boundary, and make sure the data contains what you name.
- The five steps on your own module, and which commands are free.
- What to do when a run does not converge, and the one thing never to do, which is loosen a criterion to get green.
- The honest limit, so an agent does not oversell it on your behalf.
- Two bundled references: the task file reference and the troubleshooting guide, complete, so the Skill works with no network.
Skill or MCP server, and why qikly has both
Skip this if you are not using the MCP server. The Skill works on its own. This is here because the two look interchangeable and are not: they answer different halves of the same problem, and neither replaces the other.
| MCP server | Skill | |
|---|---|---|
| What it ships | a running process exposing typed tools | a folder of Markdown |
| What it gives the model | the ability to call qikly_scaffold, qikly_validate, qikly_run, qikly_explain |
the judgement around those calls |
| Setup | qikly --install-mcp, then host configuration |
qikly --install-skill |
| Can enforce a rule | yes, in server code | no, it is instructions |
| Works when the agent has no shell | yes | no |
The MCP server is the one that can enforce things. qikly_run is built so
it cannot return the acceptance criteria, and a test in qikly's own suite fails
the build if any code path lets a criterion through. That guarantee lives in
code, and it is why the server exists.
The Skill is the one that can teach. Which line of your spec is a decision and which is a consequence; that a repeating identical patch means a decision is in the wrong half; that a criterion naming a value your data never holds produces a test that passes whatever the code does. None of that is a function call, and an agent that does not know it will use qikly and get less out of it.
The guarantee does not rest on the Skill. A Skill is text in a context window, so it can inform but not enforce. The withholding is enforced by the tool, in code, whether or not this Skill is ever installed. The Skill's own text says so, and a test asserts that it still does.
If your agent does not pick it up
Three causes, in the order to check them.
1. Your agent is not in that directory. The Skill is per project. It is files on disk, so an agent started in another folder, or running in a browser with its own sandbox rather than on your machine, cannot see them. Start the agent in the directory you installed into. This is the commonest cause by some distance.
2. Just name it. This always works, because it does not depend on the agent being told what the Skill is for:
Use the qikly skill to write tests for
src/following_distance.py.
If naming it works and the neutral request did not, the Skill is fine and the problem is discovery.
3. The host did not pass the description along. Whether an agent reaches
for a Skill unprompted depends on how much it explores before it starts typing,
and on what the host told it. Claude Code reserves a fraction of the context
window for the whole skill listing, 1% by default; when the listing does not
fit, Anthropic's own bundled skills keep their descriptions and everything else
is ranked by how often you have used it. So a skill you have never invoked can
arrive as a bare name with nothing to match against. Raising
skillListingBudgetFraction in your Claude Code settings gives the listing more
room, and that single change turned four failed routing attempts into a clean
one during testing.
Which is why naming it once is more than a workaround. That ranking is by how often you have used each skill, decaying over about a week, so a skill you have never invoked sorts below every skill you have. Name it in one request and it moves above them, and the next neutral request has a much better chance of finding it by itself. The first invocation is the only hard one.
Checking that it works
A Skill either loads or it does not, and it never tells you which, so it is worth five minutes once. If you only do one of these, do number four: the others check that the Skill arrived, and that one checks that it is right.
1. The files are where your agent looks for them. In the project you ran
qikly --install-skill in:
ls .claude/skills/qikly/SKILL.md # macOS, Linux
dir .claude\skills\qikly\SKILL.md # Windows PowerShell
And start your agent in that same directory. The Skill is per project, not per machine: an agent started somewhere else, or running in a browser with its own sandbox rather than on your computer, cannot see these files and will never load them. That is the commonest reason a correctly installed Skill appears to do nothing.
2. It loads on a request that should trigger it. Start a fresh session and ask for something in its territory, naming a real module of your own and not mentioning qikly:
Write tests for
src/following_distance.pythat would actually catch a bug in it.
The agent should mention qikly, or the decisions-and-consequences split, unprompted. If it does not, see when an agent does not pick it up below before changing anything.
3. It does not load when it should not. Ask something unrelated, such as "rename this variable everywhere", and it should stay quiet. A Skill that loads for everything costs context on every request.
4. It gives the right answer to the question that matters. This is the one to do if you do only one. Ask:
My spec says "warn when following distance breaks the two-second rule", and the acceptance criteria say a headway of exactly 2.00 s does not warn. Is that the right split?
The answer should be no, and the reason should be that "breaks" can be read as "below" or "at or below", so the boundary is a decision and belongs in the requirements. That is the single most valuable thing in the Skill, and if it comes back wrong the rest does not matter much.
5. It does not claim more than the tool does. Ask what guarantees the coding agent never sees the criteria. The answer should point at qikly's code and its build-failing test, not at the Skill.
Keeping it current
An installed Skill does not update itself, and until 0.5.4 nothing told you.
pip install --upgrade qikly replaces the package; it cannot touch a folder
copied into your project, so after an upgrade you can be following instructions
that name a different set of commands. From 0.5.4 any qikly command says so
when it notices:
note: the qikly Skill in .claude\skills\qikly is older than this qikly, so
it describes a different set of commands. `qikly --install-skill --force`
replaces it and keeps a copy of the old one.
The path is printed with your platform's own separator, so it reads with forward slashes on macOS and Linux.
--force keeps a timestamped copy of what it replaces, so a Skill you have
edited is recoverable. That is the whole update mechanism: qikly tells you, and
you run one command. There is no background process and nothing phones home;
the check is two version strings read from two files on your disk.
The Skill's version is its own and moves when its instructions move, not when qikly releases, so a release that does not touch it produces no notice.
The Skill itself lives at
src/qikly/skills/qikly/
and ships inside the installed package, so --install-skill works from
pip install qikly as well as from a clone.
A stale Skill is worse than no Skill, because an agent quotes it with confidence and the reader has no way to tell. So it is updated whenever qikly changes in a way it describes: a new or renamed flag, a change to which commands are free, a re-measured convergence figure, a new failure mode worth carrying.
Half of it cannot go stale on its own. The two bundled references are asserted
byte-identical to docs/TASK_FILE_REFERENCE.md and docs/TROUBLESHOOTING.md
by a test, so a drifted copy fails the build. The prose in SKILL.md has no
such test, which is why it is on the release checklist instead.
Files (qikly)
-
images
-
qikly_demo.gif 972.5 KB · in bundle
-
qikly_flow.png 203.1 KB · in bundle
-
qikly_hero.png 1.1 MB · in bundle
-
qikly_hero.svg 8.1 KB · in bundle
-
qikly_icon.png 6.2 KB · in bundle
-
qikly_icon.svg 1.7 KB · in bundle
-
qikly_icon_light.png 5.9 KB · in bundle
-
qikly_icon_light.svg 2 KB · in bundle
-
qikly_social.png 359.8 KB · in bundle
-
qikly_wordmark.svg 3.2 KB · in bundle
-
-
.nojekyll 0 B · in bundle
-
apple-touch-icon.png 17.4 KB · in bundle
-
CNAME 15 B · in bundle
-
CONFIGURATION.md 6.8 KB
# Configuration and LLM provider Settings, environment variables and provider setup. For getting a key onto a machine, into CI, or diagnosing one that is wrong, see [PROVIDER_KEY_SETUP.md](https://github.com/gal-a/qikly/blob/main/docs/PROVIDER_KEY_SETUP.md). ## Configuration ### Timeouts Every model call has a deadline of **300 seconds**, set by `QIKLY_REQUEST_TIMEOUT` in seconds. `0` waits forever, which is what provider SDKs do by default and is why the setting exists: a stalled connection blocks a call that never raises, so nothing downstream can react to it. With a deadline the same stall becomes an ordinary transient error and is retried with backoff. A task process prints `still running, N minutes elapsed` every five minutes, so that "not answering" is visible rather than inferred. `config/settings.yaml`, shared across all tasks: | Key | Meaning | |---|---| | `orchestrator.max_retries_per_stage` | Attempt budget per stage before the run raises. If you see the exact same patch content repeating verbatim, that's usually a requirement fighting the model's real-world prior (see below) rather than a budget problem. If instead each attempt is a *different* patch that never resolves the same failing test, that's a different signal: a bug that needs more than the failure text to resolve, rather than an artificial requirement; see [Where it fits today](https://github.com/gal-a/qikly/blob/main/README.md#where-it-fits-today) on telling the two signatures apart. | | `orchestrator.test_order` | Stage order; `unit` is always forced last (it's generated from the implementation, which doesn't exist yet during integration/system). | | `agent.max_patch_size` | Rejects an oversized PATCH and asks the model to retry smaller. Tune per task if a bigger implementation needs more room. | | `logging.save_transactions` | Turns off `transactions_*.jsonl` logging entirely; also disables the HTML report, which reads that log. | ## LLM provider Every call in a run goes to one provider. One provider per run; mixing them per agent role is not supported. ### Getting a key | Provider | Where the key comes from | Install | |---|---|---| | **Gemini** (default) | [aistudio.google.com/apikey](https://aistudio.google.com/apikey) | included | | **OpenAI** | [platform.openai.com/api-keys](https://platform.openai.com/api-keys) | `pip install "qikly[openai]"` | | **Anthropic** | [console.anthropic.com/settings/keys](https://console.anthropic.com/settings/keys) | `pip install "qikly[anthropic]"` | Then export the key and pick the provider: ```bash # Gemini, the default. Nothing else needed. export GEMINI_API_KEY=... qikly --demo # OpenAI pip install "qikly[openai]" export OPENAI_API_KEY=... export LLM_PROVIDER=openai qikly --demo # Anthropic pip install "qikly[anthropic]" export ANTHROPIC_API_KEY=... export LLM_PROVIDER=anthropic qikly --demo ``` On Windows PowerShell, `$env:OPENAI_API_KEY = "..."` instead of `export`. `API_KEY` works for any of them, and each provider's own conventional variable is accepted too, so a machine already configured for one needs nothing extra. **A note on which provider to start with.** Every convergence figure in this README was measured on `gemini-3.5-flash-lite`, over hundreds of runs, and the bundled demo is tuned to that path. The other providers work and are far less travelled here, and an entry-level model on any of them may stall on tasks that the measured path clears. If a provider you have chosen converges poorly, reach for a stronger model on it before concluding anything about the tool: model choice moves convergence more than any setting in this file. **And which model.** Prefer a small fast one: a run makes one call per stage per iteration, so a model that reasons before it answers turns minutes into tens of minutes. `gemini-3.5-flash-lite` is the floor the published figures were measured on, not a best case, and a later flash model should clear it. The table of what to start with per provider, and when a thinking model is worth the wait, is in [docs/PROVIDER_KEY_SETUP.md](https://github.com/gal-a/qikly/blob/main/docs/PROVIDER_KEY_SETUP.md#which-model). **PowerShell, CI, persisting a key, restricting one, and what a wrong key or a wrong model looks like:** [docs/PROVIDER_KEY_SETUP.md](https://github.com/gal-a/qikly/blob/main/docs/PROVIDER_KEY_SETUP.md). ### Checking a key works, for about a cent ```bash qikly --validate # free: does not touch the network LLM_PROVIDER=openai qikly --demo # one task, about a cent ``` `--demo` is the real test. It makes actual calls, writes to a throwaway folder, and reports the model, the estimated cost and whether it converged. A wrong key fails on the first call with a message naming what to check. `--validate` will not catch a bad key, because it never opens a socket. That is the point of it. ### The variables | Variable | Meaning | |---|---| | `LLM_PROVIDER` | `gemini` (default), `openai` or `anthropic` | | `LLM_MODEL` | Overrides the provider default: `gemini-3.5-flash-lite`, `gpt-4o`, `claude-sonnet-5` | | `API_KEY` | The key. Provider-specific names above are accepted too | | `QIKLY_REQUEST_TIMEOUT` | Seconds per call, default 300. `0` waits forever | | `QIKLY_MAX_CALLS` | Hard stop after N model calls, for an unattended run | One provider per run. Mixing them per agent role is not supported. ### Determinism is best-effort, and uneven With a seed set, Gemini and OpenAI are called at `temperature=0` and are given the seed itself. **Anthropic gets neither.** Its Messages API has never had a seed, and SDK 1.x removed `temperature` from `messages.create()` entirely, so there is no sampling lever left to pull. That matters if you compare rates across providers: an Anthropic figure carries more run-to-run variance than the others by construction. No provider promises identical output either way, so treat all of this as reduced drift rather than reproducibility. ### Keeping providers working The SDKs are other people's code on other people's release schedules, and this is the part of qikly most likely to break without you touching it. Two habits cover it: ```bash pip install "qikly[all-providers]" python -m pytest tests/test_provider_signatures.py -v ``` That reads the signature of every SDK you have installed and compares it against what qikly sends, so a removed or renamed parameter fails a test rather than a user's first run. It is how the Anthropic `temperature` break was found. What it cannot catch is a parameter that still exists and now means something different, or a model name retired server-side. **One `--demo` per provider before each release** covers that, costs a few cents, and is the only check that exercises the real API. Dependencies carry upper bounds for the same reason. Raising one after testing is a two-line change; not having one lets a major version arrive unannounced. -
design_1_case_study.md 16.2 KB
<!-- One of three. See MAINTAIN.md: these three files, README.md, docs/index.html and the investor deck move together. A number changed here has to change in all of them. --> # Engineering a Better System for Testing AI-Generated Code **Your AI writes both the code and its tests. How do you know the tests are really valid?** **The solution: two agents.** One turns the acceptance criteria into tests. The other writes the code and **never sees the acceptance criteria.**  Imagine a student who writes the exam paper, writes the answer key, and then sits the exam. They pass. Obviously they pass. Nobody would accept that as evidence the student knows the material, and nobody should. That is what happens when one model is handed a specification containing the acceptance criteria and asked to produce both the implementation and the suite that checks it. It reads the spec, resolves every ambiguity in it, writes code according to those resolutions, and then writes tests according to *the same resolutions*. Where the spec said "reject malformed rows" and left "malformed" undefined, the agent picked a definition, implemented it, and then tested that definition. Everything goes green. It was always going to. Run one model at temperature zero on both jobs and the two resolutions are almost identical *by construction*: the verification step burns compute and returns a tick that carries no information about whether the code is correct. That effect should get worse as models improve, since every gain in determinism tightens the agreement between the code and the tests that judge it. That last sentence is a working assumption rather than a measurement, and nothing here tests it. The fix is to take the answer key away from the student. Test generation gets the acceptance criteria in full. The coding agent gets the same specification with that section cut out, and when a test fails it sees pytest's report of that failure, never the specification section it was cut out of. Now a green suite means something happened: code written by someone who could not read the standard nevertheless satisfies it. This post is about a tool I built to do exactly that, and part 1 shows what it is and one real repair followed end to end, so you can judge the idea on something concrete rather than on a claim. --- **This is part 1 of three.** | | | | |---|---|---| | **1. The case** (you are here) | [2. How well it works](design_2_performance.md) | [3. How it is built](design_3_mechanism.md) | ## How it works, in one diagram ```mermaid flowchart TD SPEC["<b>Full specification</b><br/>task.yaml<br/>requirements + interface<br/>acceptance_criteria"] REQ["requirements<br/>+ interface"] AC["acceptance_criteria"] CODE["<b>Coding agent</b><br/>writes the implementation<br/>FIX then PATCH on failure"] TEST["<b>Test-writing agent</b><br/>writes the suite"] IMPL["Implementation"] SUITE["<b>pytest suite</b><br/>Tests for:<br/>1 integration, 2 system,<br/>then 3 unit"] RUN{"Run the suite"} FAIL["<b>pytest output</b> only<br/>no acceptance criteria"] OUT["Converged<br/><b>outputs:</b> code + suite<br/>+ audit trail"] STALL["Did not converge<br/><b>failure errors and audit trail</b><br/>exits non-zero, ships nothing"] SPEC --> REQ SPEC --> AC AC -. "never reaches" .-x CODE REQ --> CODE REQ --> TEST AC --> TEST CODE --> IMPL TEST --> SUITE IMPL --> RUN SUITE --> RUN RUN -- pass --> OUT RUN -- fail --> FAIL FAIL -- "repair loop:<br/>FIX, then PATCH" --> CODE FAIL -- "retry budget spent" --> STALL IMPL -. "unit stage only:<br/>written last, from the code" .-> TEST classDef codeView fill:#f3e8ff,stroke:#7e22ce,color:#4c1d95 classDef standardView fill:#d9ebea,stroke:#0e6a70,color:#0b3d40 classDef converged fill:#dcfce7,stroke:#15803d,color:#14532d classDef stalled fill:#fdf0d5,stroke:#b45309,color:#78350f class CODE,IMPL,FAIL codeView class AC,TEST,SUITE standardView class OUT converged class STALL stalled linkStyle 10 stroke:#15803d,stroke-width:2px linkStyle 11,12 stroke:#7e22ce,stroke-width:2px linkStyle 13 stroke:#b45309,stroke-width:2px ``` **Purple is what the coding agent can see. Teal is what the standard is written from.** They never touch. The purple arrows are the repair loop, and that is where almost all of a run happens. A run that never converges is still worth having: it exits non-zero, names the blocking tests, and keeps the same complete record. A failing suite sends the coding agent the failure text and nothing else, it produces a fix and a patch, and the suite runs again. That cycle repeats until the stage passes or the retry budget runs out, and each stage clears before the next is generated. | | Receives | |---|---| | **The test generation agent** | The requirements, the interface contract, and every acceptance criterion in full. It writes integration, system and unit tests against the standard. | | **The coding agent** | The same file with the criteria section removed, plus the text of whatever test just failed. The same vague brief a developer usually works from. | Two dotted lines in the diagram, pointing opposite ways. The first is the whole idea: the coding agent is never told the acceptance criteria it is judged against, only the requirements and the interface, so when a test fails it has to reason from behaviour rather than recall an answer it was given. The second is the one exception, and it runs the other way: unit tests are written last, from the code, because they have to name real functions. --- ## One simple end to end repair loop, for a code implementation that did not originally enforce a maximum value for the tax rate (100%) Everything below is lifted from a single real run of the bundled `CALC_TAX` task, on `gemini-3.5-flash-lite`. It is abbreviated, not invented: the prompt fragments, the failure text, the reasoning and the diff are what actually passed through the loop. ### 1. What the coding agent is given The task file has three parts. The coding agent receives two of them. First the **requirements**, in the words a person would use: ```yaml requirements: - "Read both CSV files listed in this task's inputs (input_01.csv and input_02.csv), and combine their rows into a single list before validation" - "Validate each row: order_id, item_price, quantity, tax_rate (a percentage, e.g. 8.25 means 8.25%)" - "Apply reasonable, strict real-world data-quality validation ... reject anything malformed, implausible, or out of range" ``` The task reads **two** CSV files rather than one because the orders arrive as two separate batches, and combining them before validation is itself part of what the tests check. Nothing about the tax rule depends on there being two; it is there so the pipeline has a real extract step to get wrong. Note what "reasonable" and "out of range" do not say. They do not say where the range ends. A developer handed this would have to decide, and so does the agent. Then the **interface**, the contract as a description rather than as code: ```yaml interface: module: "outputs.agent_src.code.CALC_TAX.calc" integration_functions: - "extract(input_path) -> list[dict]" - "transform(rows) -> dict # {\"accepted\": [...], \"rejected\": [...]}" - "load(data, output_path) -> None" ``` ### 2. What it is not given The same file carries eleven acceptance criteria. This is one of them: ```yaml - "tax_rate is only valid if it is a non-negative number -- a negative tax_rate is invalid. A tax_rate representing more than 100% is also invalid" ``` That criterion goes to the test-writing agent in full. It is cut out of the file before the coding agent is handed it, and `qikly --explain CALC_TAX` prints both halves so you can see the cut. ### 3. The first code implementation, and the test it fails The agent writes `calc.py` from the spec above. Its validation reads: ```python tax_rate = _parse_decimal(rate_str) if tax_rate is None or tax_rate < 0: # reject ``` The coding agent only wrote half the rule. It has the half that rejects a negative rate, and it is missing the half that asserts the value does not exceed 100%. Therefore the test fails. The suite, written from the full criteria, contains a test the agent has never seen: ``` outputs/tests/CALC_TAX/integration/test_integration.py::test_tax_rate_validation_rules FAILED for row in result["accepted"]: rate = float(row["tax_rate"]) > assert 0 <= rate <= 100 E assert 150.0 <= 100 ``` 8 of 9 integration tests pass. This one does not. Here is the whole test, because how little there is of it is the part worth reading: ```python def test_tax_rate_validation_rules(): rows_1 = extract(INPUT_01) rows_2 = extract(INPUT_02) result = transform(rows_1 + rows_2) for row in result["accepted"]: rate = float(row["tax_rate"]) assert 0 <= rate <= 100 ``` **It checks the rule in one direction only**, and that is worth noticing rather than glossing. Every accepted row must be in range, which is the assertion that failed. Nothing here requires that a row rejected *for* `tax_rate` was actually out of range, so an implementation that rejected every row would satisfy this test. The criterion rules that out; this test does not enforce that half of it. That is the honest state of a generated suite, and it is the reason the tool reports which criteria a suite traces to rather than asking you to trust that the coverage is complete. **It constructs no input.** It reads whatever the fixture files hold and asserts a property of the output. That is forced by when it was written, which is the next point, and it is why the failing value is 150.0 rather than something the test chose: the bad row was already in the data. **This run predates criterion traceability**, which is why the test above carries no marker. Current runs put the criteria a test came from on the last line inside its docstring, as `# Criteria: 7`, so a reader can follow any test back to the line of the specification it enforces. Test generation writes that; `qikly --explain CALC_TAX` prints the criteria in the same order. The quick start shows the current shape. **Why an integration test and not a unit test?** Because of when it was written. At that point no implementation existed, so the only names the test-writing agent could use were the ones the `interface` section promised: `extract`, `transform`, `load`. This test calls **`extract` then `transform`**, feeding the first one's output into the second, and stops there: `load` only writes the result to disk, and the tax rule is already visible in what `transform` returns. Integration tests are not every pair of functions. They follow **the chain the requirements describe**, `extract` to `transform` to `load`, and each test walks as much of that chain as its property needs. There is one pipeline, and the number of tests is set by how many properties the acceptance criteria assert about it. Nine here, for eleven criteria. The single end-to-end entry point, `run_calc`, is tested separately at the system stage. That constraint shapes the assertion. It cannot feed the code a tax rate of 150 directly, because it does not choose the input; it reads whatever the fixture files contain. So it asserts an **invariant over the output**: whatever ends up accepted must have a rate between 0 and 100. The bad row was already sitting in the fixture data, and the invariant caught it. The unit stage, generated later from the finished code, writes the same rule the other way round, because by then it can import internal helpers and construct inputs: ```python invalid_rates = ["-0.01", "-5", "100.01", "150", "abc", ""] for tr in invalid_rates: res = transform([{"order_id": "F1", "item_price": "10.00", "quantity": "1", "tax_rate": tr}]) assert len(res["accepted"]) == 0 ``` Same criterion, two stages, two styles: an invariant over real data first, then the constructed edge cases once there is code to point at. ### 4. The FIX, what goes back to the coding agent Pytest's own output for the tests that are failing right now, and nothing else. Be precise about what that contains, because it is more than one line: the test name, its source from `def` down to the failing statement, its docstring and comments if it has any, the assertion that failed, and the intermediate values pytest prints underneath it. Source after the failing line is not shown. Tests that are currently passing give up their names, because pytest lists everything it collected, but not their bodies, docstrings or assertions. What it never contains is the specification. The acceptance criteria are stripped from the task before the coding agent sees it, and no part of the criteria text is ever placed in a FIX or PATCH prompt. So the agent can read the one case it just failed, in the test author's words, and cannot read the rule that case came from, the other criteria, or the tests it has not failed yet. That distinction is the whole design, and it is narrower than "the agent is blind". A developer handed a failing test sees the same thing. What neither of them can do is change a test that was written before the code existed. From that input it produces a **FIX**, which is **reasoning rather than code**: ``` failure_summary: test_tax_rate_validation_rules failed because a tax rate greater than 100 was incorrectly accepted. root_cause: The tax rate validation logic lacks an upper bound check, allowing rates above 100% (e.g. 150). plan: - Add a validation rule to ensure tax_rate is less than or equal to 100 (and non-negative). target_files: - outputs/agent_src/code/CALC_TAX/calc.py ``` It has reconstructed the withheld rule from one failing assertion. ### 5. The PATCH The FIX names the files; a second call produces a unified diff of only those files, applied all or nothing: ```diff --- a/outputs/agent_src/code/CALC_TAX/calc.py +++ b/outputs/agent_src/code/CALC_TAX/calc.py @@ -52,3 +52,3 @@ tax_rate = _parse_decimal(rate_str) - if tax_rate is None or tax_rate < 0: + if tax_rate is None or tax_rate < 0 or tax_rate > 100: rej = dict(row) ``` ### 6. The stage clears, and the earlier ones are re-checked A run works through three stages, and the console names each one by number. **Stage 1 is the integration tests** and **stage 2 the system tests**, both written from the specification before any code exists; stage 2 drives the single end-to-end entry point, `run_calc`, rather than the individual functions. **Stage 3 is the unit tests**, generated last, once there is code whose internal helpers they can import and call. Every stage has to pass before the next one is generated. ``` [stage 1/3] [iteration 3] 9/9 passed [stage 2/3] [iteration 1] 5/5 passed [stage 1/3] [iteration 1-regcheck] 9/9 passed ``` Then the unit stage is generated, from the code that now exists, and the same loop runs again. That run converged in 38 seconds and 11 model calls, for about \$0.006. ### What this example shows The coding agent eventually wrote `tax_rate > 100`, and only because a test told it 150 was wrong, not because a criterion told it 100 was the limit. Had it been handed the criteria, it would have written that bound on the first attempt, the test would have passed immediately, and the green would have meant only that one model agreed with itself twice. Note that in real scenarios it is often impractical to hand over all the criteria in advance, for example when many edge cases exist. The failure is the evidence. It is what a test is capable of producing, unlike a shared context, which cannot do that. --- ## Where to go next That is the whole idea, demonstrated once end to end. Two questions follow naturally, and each has its own part. **Does it actually work, and how often?** Three sweeps, 967 runs, what reproduced and what did not, including the results this project measured and then withdrew. [Part 2: How well it works](design_2_performance.md). **How is it built, and how do I use it on my own specification?** The five agents, the FIX and PATCH separation, the stage ordering, and the command to run for whichever parts of a task file you already have. [Part 3: How it is built](design_3_mechanism.md). -
design_2_performance.md 18.6 KB
<!-- One of three. See MAINTAIN.md: these three files, README.md, docs/index.html and the investor deck move together. A number changed here has to change in all of them. --> # How well it works: three sweeps, and the results we withdrew **Part 2 of three.** [Part 1](design_1_case_study.md) makes the case that the agent writing the code should never see the acceptance criteria, and shows one repair end to end. This part asks the harder question: does it work, how often, and how much of that is measurable? Everything below is from runs on `gemini-3.5-flash-lite`, a small cheap model chosen so that repeated sweeps were affordable. Treat the figures as a floor rather than a ceiling. Where a measurement did not survive scrutiny, it is reported as withdrawn rather than removed. --- **This is part 2 of three.** | | | | |---|---|---| | [1. The case](design_1_case_study.md) | **2. How well it works** (you are here) | [3. How it is built](design_3_mechanism.md) | ## How well it works Three claims, in descending order of how much weight they can carry. ### One you can check yourself, with no statistics at all **The coding agent never receives the acceptance criteria.** Not "is instructed not to look at them", and not a convention someone has to remember: the criteria are removed before the FIX and PATCH prompts are assembled, and a test in the repository fails the build if any call site lets one through. ```bash qikly --explain CALC_TAX ``` It prints the task file twice, once as each agent receives it, and the difference between them: all eleven acceptance criteria go to one side and are cut before the other side is handed the file. No API key, no model call, about a second, and it works from a plain `pip install`. That shows the criteria being removed once, for one task. The test suite checks something stronger: that no call site anywhere in the codebase can pass a criterion to the coding agent, so the removal cannot be undone by a future change. If you have cloned the repository rather than installed the package, you can run it: ```bash python -m pytest tests/test_withholding.py -v ``` That takes a few seconds and needs no API key, no sample size and no confidence interval. It is the strongest claim here precisely because it is not a measurement: it is a property of the code, and it cannot rot. ### It converges, and the rate reproduces Measured on `gemini-3.5-flash-lite`, a small cheap model chosen to make repeated sweeps affordable, so treat these as a floor rather than a ceiling. **Roughly 8 runs in 10 finish with the code passing every integration and system test. Roughly 6 in 10 pass everything including unit tests.** The reason those are round numbers is that they were measured three times, over 967 runs, and then re-measured on a changed system. | Sweep | Runs | Integration + system | Integration, system and unit | |---|---|---|---| | 14 August | 427 | 80% | 59% | | 30 August, same tasks and settings | 140 | 87% | 67% | | 31 August, after correcting the benchmark | 400 | 83% | 64% | | 23 and 24 September, on 0.5.1 | 390 | 80% | 60% | **The last row is not a further draw from the same urn, and is deliberately not pooled with the three above it.** 0.5.1 hardened the process itself. Six changes, every one of them either narrowing what the coding agent receives or removing a way a run could stall on something no implementation could satisfy: - The criteria strip stopped being a regular expression and started asking the YAML parser where the section ends, closing two leaks a regex could not see: a comment written between two criteria, and a criterion wrapped onto a line beginning at column 0. Either handed the coding agent real criteria. - `ADAS_HEADWAY`'s header comment stopped naming which way its own withheld boundary falls. - The `# Criteria: N` markers stopped reaching the coding agent. - `CALC_TAX` shed half its commentary, which the agents had been reading on every call. - Integration tests are told to name the seam they cross. - Test generation is told to call the function under test once per scenario, after a run stalled for eleven iterations against `transform(transform(rows))`, a test no implementation can satisfy. **None of it makes the task easier, and two items make `ADAS_HEADWAY` strictly harder**, since that task had been telling the coding agent the answer it was about to be tested on. So the right reading of these rows is not "the rate held while we changed things" but "the rate held while the bar was tightened". Pooling them with the 967 would produce one number describing two systems, which is the single reason every positive result this project has withdrawn was withdrawn. Reported separately they say something a larger sample could not: 80% and 60% land inside the spread of the three earlier sweeps, 80 to 87 and 59 to 67, so none of this moved convergence by anything a sample this size can detect. The 0.5.1 row is three sweeps of 130 runs pooled, taken over two days on the same tasks and settings. Pooled, it converges **80% [76-84] and 60% [55-65]**, which is the tightest estimate on this page and the same 8 in 10 and 6 in 10 the August sweeps found before any of the hardening above existed. The third sweep is the interesting one, and the reason is what happened between the second and the third. Asking a separate agent which criteria no input row could trigger turned up unreachable rules in eight of the ten tasks, from one in `CALC_CALENDAR` to seven of thirteen in `MERGE_CONTACTS`. A criterion nothing can reach produces a test that passes whatever the code does, so part of every bar was not lower, it was absent. Thirty-one rows were added, and the whole sweep was taken again. The prediction was that convergence would fall, because the bar had genuinely become enforceable. **It did not move.** 64% sits between the two earlier figures and inside both intervals. Either the added rows exercise behaviour the implementations were already getting right, or the difference is smaller than this measurement can see, and separating those needs fault injection rather than convergence. That experiment is not run. Three independent sweeps agreeing is worth more than any one number's decimal places, so the decimal places are not quoted. One thing about a rate like this is worth internalising before you run anything: **a single run is an artifact, not a rate.** The same task with the same seed converges on some runs and exhausts its budget on others. To make a claim about how often anything converges, repeat the sweep and read the interval. ### Nearly the whole gap between those two numbers is the unit stage This is the most stable finding here and the most useful one, because it tells you where runs actually fail. It held at about twenty points across both sweeps. The reason is structural rather than mysterious. The unit suite is the largest, so there is more to satisfy. It runs last, when the earlier stages already pass and regression re-checks force them to keep passing, so a fix has the least room to move. And it is the one stage whose tests are written with sight of the implementation, so it can assert on incidental internal structure rather than on required behaviour. That last point deserves emphasis: **the only stage where the blindness is broken is also the hardest stage.** Whether that is cause or coincidence is testable, and untested. ### What explains the variation The variation in the convergence rates is not due to size because every task has 6 to 7 spec requirements, 10 to 13 criteria, 3 interface functions and 2 input files. Three things drive the difference: 1. *Carried state of a task.* Whether row N's correctness depends on the rows before it. The three `CALC_*` tasks are row-independent and average 78%. The three `MERGE_*` tasks all carry state and average 46%. 2. *Cross-file coupling.* Whether two input files can be concatenated or must genuinely be interleaved. `MERGE_STOCK` must merge-sort by timestamp before computing anything; concatenation produces plausible, wrong answers. 3. *Conflict with the model's priors.* Where a criterion asks for a convention the LLM would otherwise resolve differently. ### When a run stalls A run that exhausts its budget exits non-zero, names the blocking tests, and keeps the complete record. **There is no path by which it reports success on code its own tests reject**, and in several hundred measured runs there has never been such a case. Stalls are usually near misses rather than wreckage: most of the blocking stage is already passing, and a large share fail exactly one test. They also cluster. The exact failing test names almost never repeat, but the *subjects* do, so what you get is a short list of nameable problems rather than a diffuse failure rate. **Three of the four causes are fixed by editing text, not by buying a bigger model.** They live in the specification or the bar rather than in the coding agent, which means most stalls are within your control and cheap to resolve. The signature of each, and what to do about it, is in [TROUBLESHOOTING.md](TROUBLESHOOTING.md). ## Is the bar any good? This is the question that matters, and the tool is built to answer it with evidence rather than assertion. A programme of work is running on exactly this: cross-testing implementations against other runs' suites, planting faults in code a suite has already accepted to see whether it notices, and asking whether automatically refined criteria catch more than a first draft. Those results will be published once substantial user data has accumulated and been carefully analysed, because a figure earned across many real specifications is worth far more than one earned across ten example tasks. One finding from that work is already solid enough to act on today, because it is a mechanism rather than a rate. **A suite that never tests the boundary cannot detect an error at the boundary**, no matter how many other cases it covers. Change a `>` to a `>=` in code a generated suite has already approved: ```python # the code the suite accepted if quantity > 100: reject(row, "quantity too large") # the same code with one character changed if quantity >= 100: reject(row, "quantity too large") ``` Those two versions disagree about exactly one input, `quantity == 100`, and agree about every other input in the universe. Suites generated for this task tested 5, 50 and 250: | Test input | Original | Mutated | Suite can tell? | |---|---|---|---| | 5 | accept | accept | no | | 50 | accept | accept | no | | 250 | reject | reject | no | | **100** | **accept** | **reject** | **yes, but no test used 100** | Adding more tests at 5, 50 and 250 would not help. Only a test at exactly 100 would. What this says is that the generated suites were testing that the logic works, not that it stops in the right place, which is the difference between a test written to demonstrate behaviour and a test written to catch a mistake. **And it tells you exactly what to fix, in your own file, today.** A criterion phrased "reject quantities above 100" invites a test at 250. A criterion phrased "100 is accepted and 101 is rejected" forces a test at the boundary. The blind spot is in how the criteria are worded, which you control, rather than in the model, which you do not. ## Where the bar comes from Everything above assumes a human wrote the criteria. The harder question is whether the tool can write them itself, because that is the expensive half of QA. The procedure goes as follows: draft initial criteria from spec requirements alone; converge a code implementation against that draft in an isolated workspace; then give a reviewing agent the specification, the current criteria, *and the finished implementation*, and ask what the code does that none of the criteria constrain? The reviewer is looking for a specific thing: the gap between "passes what is currently tested" and "actually correct". These are places where the validation is looser than it appears, where one code path applies a normalization and another does not, or where a boundary condition is handled correctly only by accident rather than by design. Each such finding becomes a new criterion, tagged with an appropriate category. Before trusting a drafted bar on a task where you have not written one, there is a way to find out what it would have cost you on a task where you have: ```bash qikly --compare-criteria <MY_TASK> ``` It drafts criteria from that task's requirements alone, then reports them against the ones you wrote, treating yours as the ground truth. Output is covered, partial and missed, with every gap quoted in full and the drafted criteria that match nothing of yours listed separately, since those are either noise or a rule you know and never wrote down. The matching is a model's reading rather than a measurement, and says so: two criteria can mean the same thing in different words. It also says nothing about test quality, because two bars can describe the same rule and produce suites that catch different faults, which is the whole reason this project measures a bar by fault detection rather than by its text. *Currently, the tool only makes additions.* It appends criteria and never modifies or removes one. **This constrains the machine, not you.** The criteria live in your YAML file. You can rewrite one, delete one, or throw the whole set out, at any time, exactly as you would edit any other file you own. Nothing is locked. What the *system* cannot do is remove a criterion by itself. The reason is narrow and worth stating plainly. The tool's success condition is "all criteria satisfied". If the tool could also edit the criteria, then the cheapest way to satisfy a failing criterion is to delete it, and every run would converge by definition. Success would stop meaning anything. Append-only is what keeps a passing run evidence of something rather than evidence of nothing. So the answer to "we let it add criteria that may be wrong and can never be fixed?" is no, on two counts. A proposed criterion is never adopted automatically: a human reads it and accepts it, because a drafted bar is a proposal, not ground truth. And once adopted, it is yours to correct like any other line in the file. The extension not yet built is letting the tool *propose* a removal, with the evidence for it, for a human to accept or reject through a normal gate such as a pull request, keeping the superseded criterion in the history rather than erasing it. The line that matters is between proposing and enacting, not between adding and removing. ### Does refining the criteria make the suite catch more? **An open question, and an active one.** Refinement reliably *grows* the bar: 54% more criteria, and `CALC_TAX` goes from 12 to 26 across three rounds. Growth and improvement are different things, so the question worth answering is whether the grown bar catches more real defects. Measuring that is a genuinely interesting experimental problem, and most of the work so far has gone into the apparatus rather than the answer. Comparing two bars means holding everything else constant: the code, the fault set, the suite size, and the substrate each suite is judged on. Each of those took a round to get right, and the current design controls all four. The remaining piece is fixture coverage. In a fault-injection comparison a planted fault can only be caught if some input row reaches the behaviour it changes, which is exactly what the [fixture proposal agent](design_3_mechanism.md#the-five-agents) was built to close. Results will follow once substantial user data has accumulated and been carefully analysed. Refinement is shipped and usable today. A quantified claim about how much it sharpens the bar is what the work above will produce. ## Should the specification iterate too? No, and that is a deliberate design decision worth explaining. The spec requirements are the fixed point everything is judged against. A system permitted to rewrite its own goal can always satisfy a failing test by weakening what was asked. That is the same failure mode append-only prevents on the criteria side, and automating both ends would remove the last thing making convergence meaningful. The legitimate version is different and worth building: when the review finds behaviour no criterion constrains, it is often because *the specification was ambiguous there*. Reporting "your spec does not say what happens when X, and the implementation chose Y" is a proposal to a human, not a self-edit. For a real team that may be the most valuable thing the tool produces. ## In summary - **An agent that writes its own tests is grading its own homework.** At temperature zero with identical inputs it is provably vacuous: the same model that wrote the bug writes the test that blesses it. - **The fix is structural, not procedural.** The coding agent never receives the acceptance criteria. Not "is told not to look", but never has them in its context. - **You can verify that in one command.** `qikly --explain <MY_TASK>` prints what each agent is given and the difference between them. No API key, no model call, no sample size. From a clone, `pytest tests/test_withholding.py` additionally proves no call site can leak one, including a call site added next year. Both are properties of the code rather than benchmark results, so neither can go stale. - **It converges, and the rate reproduces.** Roughly 8 runs in 10 pass every integration and system test, roughly 6 in 10 pass everything including unit tests, measured three times on a small cheap model over 967 runs, the third after correcting the benchmark itself, then re-measured on 0.5.1 over 390 runs after the process was hardened, and reproduced at 80% and 60%. Every run that does not converge exits non-zero and names its blocking tests, and none has ever reported success on code its own tests rejected. - **Nearly the whole gap between those two figures is the unit stage**, which is also the only stage whose tests are written with sight of the code. - **How you word a criterion decides how sharp the test is, and that is yours to control.** A suite tests the boundary when the criterion names the boundary: "100 is accepted and 101 is rejected" produces the test that "reject quantities above 100" leaves to chance. This is the highest-leverage thing you can do in your own file. - *Whether automatic criteria refinement produces a measurably sharper bar is the open question, and an active one.* It grows the bar reliably, and the experiment design to quantify the rest is built. The measurement programme continues alongside the tool, and its results will be published once substantial user data has accumulated and been carefully analysed. Figures earned across many real specifications, from many people, are worth considerably more than figures from ten example tasks. What is published here has already survived an independent re-measurement; the rest will meet the same bar before it joins it. -
design_3_mechanism.md 20.1 KB
<!-- One of three. See MAINTAIN.md: these three files, README.md, docs/index.html and the investor deck move together. A number changed here has to change in all of them. --> # How it is built, and how to use it **Part 3 of three.** [Part 1](design_1_case_study.md) makes the case and shows one repair end to end. [Part 2](design_2_performance.md) reports what was measured. This part is the architecture: what separates it from the alternatives, the five agents, and the command to run for whichever parts of a task file you already have. --- **This is part 3 of three.** | | | | |---|---|---| | [1. The case](design_1_case_study.md) | [2. How well it works](design_2_performance.md) | **3. How it is built** (you are here) | ## What makes this different **The tests come from the standard, not from the code.** This is the one that matters most. Every other AI test generator in this space derives its tests from an implementation that already exists: it reads the code to decide what to assert, which produces an excellent regression harness that locks in current behaviour. qikly writes the integration and system suites from the acceptance criteria **before any implementation exists**, so there is nothing for the standard to be shaped by. That is why a green suite here carries information: the tests describe what the code should do, not what it already does. **The withholding is a mechanism you can watch, not a promise.** The criteria are cut out of the file in code, before the FIX and PATCH prompts are assembled, so no representation of them exists in the coding agent's context. `qikly --explain <MY_TASK>` prints what test generation receives, what the coding agent receives, and the difference between them, from a plain install with no API key. A test fails the build if any call site lets a criterion through, including one added next year by someone who has never read this page. **What you get is an executable suite you keep.** Real pytest files, plus JUnit XML for whatever tracks tests where you work. Read them, run them, put them in continuous integration, and when one fails in six months it fails for a reason you can inspect and argue with. A suite is a durable asset in a way a model's verdict is not: a verdict cannot be re-run against tomorrow's commit. **It helps with writing the standard, not just checking against it.** `--scaffold` turns code you already have into a task file, `--criteria-from` reads criteria straight out of the ticket that already holds them, `--generate-criteria` drafts a first bar from requirements alone, `--check-criteria` looks for two statements in the specification that no implementation could satisfy at once, criterion against criterion included, and the refinement loop reviews converged code to propose criteria the first draft missed. Whether that last step produces a measurably *sharper* bar is an open question this project is still working on. **Every run is reproducible, and the whole trail is kept.** A run records the provider, the model, the settings and the version that produced it, next to every failing test, every FIX with its stated root cause, and every PATCH as a diff. You can read back exactly why a line of code exists: which assertion forced it, what the model concluded, and what it changed. That record is what makes a convergence rate a measurement rather than an anecdote, it is what let three of this project's own positive results be withdrawn on inspection, and it survives a run that never converges. Nothing here asks you to take a number on faith. ### "Why not just use two different models?" It is the first thing most people ask, and it does help a little. It does not reach the underlying issue, though, because both models still read the same criteria and so both still write to them. Where the criteria say "reject malformed rows" and never define malformed, two models resolve that ambiguity from the same sentence, and the implementation is still built around the resolution the tests will check for. Two models is also a habit rather than a mechanism: nothing checks they stayed different, and a settings change a year from now undoes it with no test to notice. Withholding removes the channel rather than the coincidence, so it holds whichever model is writing. The two compose nicely, incidentally, since qikly picks a provider and model per agent role: you can withhold *and* use two models. ### "Why not just add a reviewer agent?" The newer form of the same question, and the one to answer carefully, because independent verification steps are now shipping in mainstream coding agents: a second agent, usually from a different model family, reviews what the first produced. It helps, and it stops in the same place. A reviewer handed the same specification has read the same acceptance criteria and resolves the same ambiguity the same way. It catches what is visible from that context: an inconsistency, a requirement plainly skipped, a bug that looks like a bug. It cannot catch the case this design exists for, where the code and the standard agree because both came from one reading of a line that admitted two. Relative to the shared interpretation nobody in that loop is wrong, which is precisely why they all agree. The problem was never that nothing was checking. It is that everything checking had already read the answer key. Model diversity varies who is looking; withholding varies what they were shown, which is the only one of the two that changes what can be found. And it is enforced by a test rather than by an arrangement someone has to remember to keep. One more difference, and it is the one that outlasts the run: a reviewer emits a verdict, and this emits a pytest suite that is still there in six months, running against tomorrow's commit. ## The mechanism A task file has **three** parts, and the cut runs between the third and the first two: 1. **`requirements`** is the vague part, and it holds the decisions: what a real specification says before anyone sharpens it, including every choice that could have gone another way, such as a threshold, a unit or an exemption. 2. **`interface`** is the contract as a description rather than code, the function signatures and where the module will live. Both agents read it. 3. **`acceptance_criteria`** is the sharp part, and it holds the consequences: specific, objectively checkable statements of what must be true if those decisions were implemented correctly, including the boundary values and edge cases a careless reading gets wrong. A decision does not belong here: withheld from the one agent that needed it, it produces a stuck loop rather than a harder test. The test generation agent receives all three. The coding agent receives the first two, **with the acceptance criteria removed in code before the prompt is built**. It is not instructed to ignore them. It cannot be persuaded, prompted, or induced into seeing them, because no representation of them exists in its context. Everything else follows from that asymmetry. A run works through three stages, `integration` then `system` then `unit`: 1. **Integration and system tests are generated first**, from the spec alone, before any code exists. They cannot see an implementation because there is not one yet. 2. **The coding agent writes the implementation**, from the spec minus the criteria. 3. **pytest runs.** On failure the model produces a **FIX** (failure summary, root cause, plan, and the files it intends to touch) and then a **PATCH** (a unified diff of only those files), applied all or nothing. Repeat until the stage passes or the budget is spent. 4. **Unit tests are generated last**, once real code exists for them to name. This is the only stage allowed to see the implementation. 5. **Clearing a stage re-runs the earlier ones**, so a later fix cannot silently break something that already passed. Two things are worth noticing about step 3. The reasoning and the diff are separate LLM calls, which makes the reasoning independently auditable and restricts the diff's context to the files the reasoning identified. And the FIX is generated from pytest's output alone: the failing test's name, its source down to the failing line, its docstring and the assertion, with no acceptance criteria. The same feedback a developer sees in their terminal, and no more. There is no separate "write the initial implementation" step anywhere. The first run happens against an empty source tree, fails on a collection error, and that failure feeds the same FIX and PATCH cycle as every later repair. The first line of code and the hundredth are produced by one mechanism. ### The five agents Everything above is done by five agents, each with its own definition file in `inputs_public/agent_defs/` and its own entry under `agents:` in `settings.yaml`, so any of them can be pointed at a different model without touching the others. Two of them exist because a single agent doing both jobs is the failure this project is about. | Agent | Settings key | Reads | Never reads | Produces | |---|---|---|---|---| | **Coding agent** | `code` | requirements, interface, and pytest's output for whatever test just failed | the acceptance criteria | a FIX, the reasoning and the files it intends to touch, then a PATCH, a unified diff of only those files | | **Test-writing agent** | `test`, or `test_integration` / `test_system` / `test_unit` | requirements, interface, and every acceptance criterion in full | the implementation, except at the unit stage, which is written last and exists to name real functions | one pytest file per stage | | **Criteria drafter** | `criteria` | the requirements alone | any implementation, and any existing criteria | a first draft of the bar, only when `--generate-criteria` asks for one | | **Criteria reviewer** | `review` | the specification, the current criteria, and a finished implementation | nothing withheld: this is the one agent shown everything at once | proposed additional criteria, each tagged with a category, for a human to accept or reject | | **Fixture proposer** | `fixtures` | the requirements, the criteria, and the fixture files as they stand | the implementation, and the criteria it is not asked about | for each criterion nothing currently reaches, one row that would, written to a proposal file for a human to accept | The coding agent runs far more often than the others, once per repair attempt, which is why it is the one where a cheaper model pays for itself. The reviewer is the one asked to find what nobody wrote down, which makes it the likeliest to be worth a stronger model. Both are one line in `settings.yaml`. The fixture proposer is the newest and the only one that suggests changing the *inputs* rather than the code or the bar, which is why it is the only one whose output never lands anywhere automatically. A criterion nothing can trigger produces a test that passes whatever the code does, and in this project's own measurements roughly two thirds of deliberately planted faults were missed by every suite for that reason: the bar was unmeasurable rather than wrong. But a row is harder to review than a sentence, because it is only right or wrong relative to the criterion it was proposed for, and a fixture set that grows in whatever direction a model finds interesting stops resembling the data you actually process. So proposals go to a file, capped per round, each row printed under the criterion it exists to reach, and the file reports what share of your data a machine has written so the drift is visible in aggregate rather than one plausible row at a time. ## Using it A word first on how this was built, because it is not incidental to the subject. I architected the tool's objectives and its orchestration, and I guided the research: which experiments to run, which results to believe, and which of my own claims to discard when the numbers did not support them. **The code itself was largely written and tested by Claude Opus 5.0, across many iterations of review, correction and rework.** That feels worth stating plainly in a post about not letting one model mark its own homework. The separation this tool enforces is the same separation I relied on while building it: the measurements decided what was true, not the author of the code, and several of the conclusions below are ones I did not want. ```bash pip install qikly export GEMINI_API_KEY=... qikly --demo ``` The demo runs one task end to end in an output directory and prints what it built. That takes about thirty seconds. It writes nothing outside that directory. ### What you already have decides how you use it A task file has the three parts described under [The mechanism](#the-mechanism), and which of them you already have decides both what qikly does for you and which command you run. The short version: `qikly --explain <MY_TASK>` to see the split for yourself, `qikly --demo` to watch a whole run, `qikly --init` to start a task from nothing, and `qikly --scaffold <FILE>.py` to start one from code you already have. The full table, with the exact command for each starting point, is in [QUICK_START_ON_YOUR_OWN_DATA.md](https://github.com/gal-a/qikly/blob/main/docs/TASK_FILE_REFERENCE.md#which-command-depends-on-which-parts-you-already-have). The four that matter most, most valuable first: 1. **You have #2 and code somebody else wrote, and you want that code verified.** Something else produced the implementation. Point `--scaffold` at it and it reads the real signatures into the interface for you, leaving the requirements and the criteria yours to write. The suite is then written from criteria the implementation's author never saw. A suite generated from the same context as the code is a model agreeing with itself, and agreement is not evidence. This is the use no other tool in this space covers. One honest caveat comes with it, and the run states it rather than leaving you to find it. qikly did not write that implementation, so it cannot know what its author read: the separation is evidenced here rather than enforced. The run asks git whether the task file was last changed before the implementation's first commit and prints the answer next to what that answer does not show. A failing test is a real finding either way, since it was written from the criteria without reading the code. It is the passes the evidence qualifies. 2. **You have #1, #2 and #3, and no code.** A specification exists and an implementation does not. You get a first implementation plus the suite that justifies it, and nothing is drafted on your behalf. Everything lands in `outputs/` for review; nothing is written to your source tree. 3. **You have #1 and #2, and want a first draft of the bar.** `--generate-criteria` drafts #3 from the requirements alone, or `--criteria-from` lifts it out of the ticket where you already wrote it. Then the run proceeds as above, and the draft is yours to correct. 4. **You have everything and want it run unattended.** Non-zero exit on any non-convergence, so a scheduler or CI job can run many specifications and keep the reports as artifacts. There is a GitHub Action, and JUnit XML for whatever tracks tests where you work. ## Now it's your turn to test it on your specification The engine is open source under Apache 2.0, because the central claim is one nobody should take on faith. The whole argument rests on the coding agent genuinely never seeing the bar, and that is something you can test and verify rather than just accept. Read the prompt builders in `src/qikly/agent_api/prompts/`, read the function that strips `acceptance_criteria` out of the specification before the prompt is built, and watch a run do it. That verification is the point of publishing the engine at all. *Try it in thirty seconds:* ```bash pip install qikly export GEMINI_API_KEY=... qikly --demo ``` Then do the thing that actually tests the idea: write one specification of your own, split it into `requirements` and `acceptance_criteria`, and read the criteria the tool derives against the ones you would have written by hand. If it finds a gap you missed, that is the argument. If it does not, I want to know which specification broke it. Issues and results, welcome and wanted: [github.com/gal-a/qikly/issues](https://github.com/gal-a/qikly/issues) --- ## Appendix: reference ### Example tasks Thirteen tasks ship with the tool, in five domain families. | Family | Task | What it does | |---|---|---| | `CALC_*` arithmetic and precision | `CALC_TAX` | Per-line order tax; stresses rounding and currency precision | | | `CALC_DISCOUNT` | Discount amounts; validation versus computation edge cases | | | `CALC_CALENDAR` | Date-range charge from a monthly rate; calendar arithmetic, leap years | | `ETL_*` string and format validation | `ETL_EMAIL` | Validates and normalises contact records with an email field | | | `ETL_ADDRESS` | Postal addresses; deduplicates across two files | | | `ETL_NAME_SPLIT` | Splits a full-name field into first, last and suffix | | `MERGE_*` combining two sources | `MERGE_SALES` | Two sales exports with overlapping ranges; cross-file dedup | | | `MERGE_STOCK` | Inventory transactions into per-SKU stock; cross-file ordering | | | `MERGE_CONTACTS` | Contact records from two systems; conflict resolution by recency | | `AGG_*` event-stream aggregation | `AGG_RUNLOG` | Summarises append-only JSONL run logs | | `ADAS_*` vehicle sensor data | `ADAS_HEADWAY` | Following-distance warnings from forward-radar samples; inclusive limits and the two-second boundary | | | `ADAS_TTC` | Time to collision from radar tracks; a closing speed that can be zero or negative, and a three-second boundary | | | `ADAS_SPEED_LIMIT` | Speed-limit compliance; a unit conversion, an enforcement tolerance, and an enumerated set of valid limits | Every task ships with `requirements`, an `interface`, and a full set of `acceptance_criteria`, so each one is a worked example of the task format as well as something to run. ### Supplying your own code or tests The loop takes three inputs, and each can be yours or generated, independently and in any combination. `seed.implementation` points the tool at code you already have, so the run skips generating a first implementation and goes straight to testing and repairing yours. `seed.tests` keeps a suite you already trust, so the loop repairs the code against your tests rather than its own. That last combination is the one worth naming, because it is the strongest use of the tool: **your suite, its code.** The bar was written by a person and the implementation has to satisfy it without ever having read it. Setup, the exact YAML, and what `--scaffold` writes are in [README.md](https://github.com/gal-a/qikly#bringing-acceptance-criteria-you-have-already-written). ### Measuring rather than producing ```bash python -m qikly.orchestrator.run_all --repeat 10 ``` Repeats the whole sweep N times and writes one aggregate report to `outputs/reports/`. Concretely, that report contains: | Output | What it is | |---|---| | Convergence rate per task | Converged runs over total, with a 95% confidence interval, so a task at 6/10 is reported as a range rather than as "60%" | | Per-stage pass rates | How often integration, system and unit each cleared, which is what identifies the blocking stage | | Right-censored runs | Budget-exhausted runs counted separately from failures, because "did not finish in 10 attempts" is not the same fact as "cannot be done" | | Stall signatures | Non-converging runs grouped by the tests they persistently failed, which is what turns a pile of stalls into a short list of recurring subjects | | Cycle counts | FIX-to-PATCH cycles per run, so cost per converged task is visible | The point is that it emits **intervals and groupings, not a single headline number.** A rate without an interval invites exactly the mistake described below. Use `--repeat` when making a claim. Use a single run when you want the output. That distinction has already caught us out: a criteria change we were confident about looked like a fix, and ten repeated runs showed 5 out of 10 against 6 out of 10 before, no improvement at all and well inside the noise. ### Watching a run live ```bash python -m qikly.orchestrator.live_view ``` Tails the current run's transaction log and renders it as it happens: which stage, which iteration, which tests failed, and what the model did about it. -
EXISTING_CODE_AND_HELPERS.md 4.6 KB
# Pointing qikly at code you already have What qikly can see, what it can change, and what happens when the defect turns out to be somewhere it may not write. Split out of [FAQ.md](FAQ.md) because the answer outgrew a FAQ entry. The YAML for everything below is in [TASK_FILE_REFERENCE.md](TASK_FILE_REFERENCE.md); this page is the reasoning. ## qikly does not go looking for your code There is no discovery step. qikly does not scan your repository and work out what is available, so it will not find your driver or parser classes on its own. What the test-writing agent targets is what the task file's `interface` block declares, plus the requirements and the acceptance criteria. You describe the surface. Two things make that less manual than it sounds: - **`qikly --scaffold path/to/module.py`** reads the real signatures out of a file you point it at and writes the `interface` block for you. - **`seed.implementation`** points the run at code you already have, and **`seed.tests`** keeps a suite you already trust, per stage, so the loop repairs code against your tests rather than against generated ones. ## Can the tests import my other classes once they exist? Yes, with one condition. Each stage runs as `python -m pytest` from your project root, which puts the project root on `sys.path`. So a package sitting at the project root, or one installed into the same virtualenv, imports normally from both the generated tests and the generated implementation. A package in a subdirectory that is not on the path does not. The run stops with `no test ran: the module could not be imported`, and no amount of iterating fixes it because the problem is the path rather than the code. Put the directory on `PYTHONPATH`, or `pip install -e .` your own package. ## What it can read, and what it can change, are decided by the seed `seed.implementation` takes a directory as well as a file. A directory is treated as a package: it is copied in under its own name and becomes importable by that name, so `from mypkg.money import to_cents` resolves as it does in your own tree. | | Inside the seed | Outside it | |---|---|---| | Imports at runtime | Yes | Yes, if on `sys.path` as above | | The coding agent **sees the source** | Yes | No. It reasons about them from their behaviour | | The coding agent **may write** to them | Yes | No | **The seed is the boundary you choose, and that is deliberate.** qikly does not follow your imports and decide for itself which of your files an agent may rewrite: the transitive closure of a real package has no natural edge, and "it rewrote a shared module I never named" is a worse outcome than naming a folder. So put inside the seed what you want worked on, and leave a vendor library, or a large shared module you do not want touched, outside it. ## When the defect is outside the seed A test that fails because your `utils.py` is wrong does fail, correctly, which is the point. But the agent cannot patch a file outside the seed, so left to itself it would either stop without converging or change a module it does own to work around a defect that is somewhere else. From 0.5.5 a run says which of those happened, rather than leaving you to infer it from a patch history: - **At the start**, if the task's module imports anything local it may not write, one line names those files. It is information, not a warning: they import and run normally, and most of the time they are fine. - **When a run stops without converging**, the message names the file instead of ending at "exceeded 10 attempts". - **When a run passes but a failure along the way was traced into one of those files**, it says so, because a green suite reached that way may be green because the code was bent around a defect that is still there. **This is the one worth reading twice.** **Nothing here stops or slows a run.** A project with shared helpers is normal, and runs in one converge exactly as before. The guard only speaks up. If it was not what you wanted, the fix is usually one line: seed the package that holds `utils.py` instead of the single module. ## `--score-code` has no such limit It neither writes code nor calls a model. Point it at a package and it plants faults throughout it, helpers included, and tells you which ones your suite noticed. No task file, no seed, nothing of yours modified. ## One thing worth knowing before you start The loop reruns pytest on every iteration, so it suits the deterministic layer best. Protocol parsing is a good fit: bytes in, structured records out, driven from recorded captures. Code talking to live hardware is better left behind the test doubles you already have, because a suite that needs a rig attached is a suite the loop cannot rerun freely. -
FAQ.md 7 KB
# FAQ Questions people have actually asked, with the answer checked against the code rather than remembered. Where an answer has a boundary, the boundary is stated: a qualified yes is more use than an unqualified one. ## 1. Does it run locally or in the cloud? Locally. It is a `pip install` and a command line tool on your machine, or in CI through the bundled GitHub Action. There is no qikly service in the middle and nothing to sign up for. The only thing that leaves your machine is the prompt sent to whichever model provider you configure, using your own API key. Gemini, OpenAI and Anthropic are supported; see [PROVIDER_KEY_SETUP.md](PROVIDER_KEY_SETUP.md). ## 2. Does my code or my data go to you? No. Nothing is sent to the project, and the tool collects no telemetry. What reaches your model provider is what the prompts contain: the task file as each agent receives it, and the test failures the loop is working through. Your provider's own terms then govern that traffic. ## 3. Is there anything I can run before committing an API key? Yes, three things, all offline and free: ```bash qikly --explain MERGE_SALES # what each agent is shown, and the difference qikly --explain MERGE_SALES --html # the same as one page you can share qikly --validate # check your task files qikly --score-code src/yours.py --score-tests tests/test_yours.py ``` `--explain` is the one worth running first: no model call, about a second. `--score-code` is the one that runs on **your** code rather than on a bundled task. It plants one fault at a time and reports which ones your existing tests did not notice. No task file, no run, no key, and nothing of yours is modified. Read the score as a floor: it says how much of the code that is there your tests would notice changing, and nothing about a rule nobody implemented. ## 4. What is the difference between `--score-code` and `--score-tests`? They are not alternatives. They are the two halves of one command, and it does nothing useful without both: ```bash qikly --score-code src/pricing.py --score-tests tests/test_pricing.py ``` `--score-code` is **the code to plant faults in**, a module or a whole package. `--score-tests` is **the suite to run against each planted fault**, a file or a directory of them. The report names the faults your tests did not notice. Neither writes anything: the faults go into a temporary copy, and your files are not modified. No task file, no run, no API key, no model call. ## 5. What is the difference between `--demo` and `--example`? `--demo` is for watching, `--example` is for editing. `qikly --demo` runs a bundled task end to end in a throwaway folder you can delete, so you can see a real run before deciding anything. It needs an API key, takes about thirty seconds and costs under a cent. It changes nothing in your project. `qikly --example` writes a finished worked example **into your project**: a module, its two task files and sample data, at the paths the task files name. Nothing is left for you to fill in, so `qikly --tasks MY_METRICS_VERIFY` runs immediately, and it is the thing to copy when writing your own. It never overwrites an existing file. If you are deciding which to type first: `--demo` to see whether this is for you, `--example` once you have decided it is. ## 6. Does the coding agent really never see the acceptance criteria? That is what `--explain` exists to show, on your own task rather than on a claim in a README: it prints the task file as test generation receives it, then as the coding agent receives it, then the difference. It builds those strings through the same function a real run uses, so it demonstrates the mechanism instead of describing it. The property is also held by the test suite: no call site in the codebase can pass a criterion to the coding agent, so the removal cannot be undone by a later change without a test failing. ## 7. What does one passing run prove? That this task converged this once. A run is a loop with a variable trip count, so one run is an artifact and never a rate. If you want a number you can quote, repeat the run and report the spread. The project's own performance figures are in [design_2_performance.md](design_2_performance.md), with the sample sizes they rest on. ## 8. What will a run cost? Every run prints a projection before it starts: the expected number of model calls, tokens in and out, and a price from a static table rather than from your bill. Treat it as a projection, because the trip count varies. ## 9. Can I use it commercially? Yes. qikly is Apache 2.0, which permits commercial use, modification and redistribution. Note that the licence grants no trademark rights, so building a service on it is fine and naming that service after the project is a separate conversation. ## 10. Can I drive it from an editor or an agent? Yes, it ships an MCP server, so Claude Code, VS Code and other MCP hosts can call it. See [mcp.md](mcp.md). The server deliberately never returns acceptance criteria, for the same reason the coding agent never receives them. ## 11. Does qikly know about the classes I already have, such as hardware drivers or protocol parsers? Not by discovery: it targets what your task file's `interface` block declares, and `qikly --scaffold your_module.py` writes that block from the real signatures. Your other classes import normally at run time as long as they are importable from your project root. What it may **change** is a separate question from what it can import, and the answer is the seed: `seed.implementation` can name a single module or a whole package, and everything inside it is visible to the coding agent and repairable. A defect in a helper you left outside the seed is caught by the tests, cannot be repaired, and the run now says so rather than working around it. The full picture, including what a run prints in each case, is in [EXISTING_CODE_AND_HELPERS.md](EXISTING_CODE_AND_HELPERS.md). ## 12. My module is not one file. Can qikly work on a package with a nested `utils/`? Yes, from 0.5.5. Point `seed.implementation` at the directory rather than the file, and name the module by the import path your own code already uses: ```yaml interface: module: "mypkg.pricing" seed: implementation: "mypkg" ``` The package is copied in under its own name, so `from mypkg.utils.money import to_cents` resolves exactly as it does in your tree. Relative imports work too, at any depth: `from .utils.money import to_cents`. No `__init__.py` is required, and one that is there is kept. **Every module in the package can be repaired**, nested ones included, and a single patch may change more than one of them. Two things that will not work, and both fail loudly rather than quietly: a package whose name is also a standard library module's, such as `json`; and modules that import each other circularly at the top level, which plain Python rejects too. The full matrix, including the import shapes that do not work, is in [TASK_FILE_REFERENCE.md](TASK_FILE_REFERENCE.md#seeding-a-package-when-the-implementation-is-more-than-one-module). -
favicon.ico 4.9 KB · in bundle
-
favicon.svg 1.5 KB · in bundle
-
index.html 52.9 KB · in bundle
-
mcp.md 11.4 KB
# qikly over MCP `qikly` runs as an [MCP](https://modelcontextprotocol.io) server, so an agent in any MCP host can start a run and read its result without you leaving the conversation to type a command. It is tested in VS Code and Claude Code so far; the setup for other hosts below follows the MCP standard but has not been verified end to end. With [uv](https://docs.astral.sh/uv/) installed, nothing else needs installing. In Claude Code: ```bash claude mcp add qikly -- uvx --from "qikly[mcp]" qikly-mcp ``` Other hosts take the same command in their own config file, and it speaks stdio: ```json { "mcpServers": { "qikly": { "command": "uvx", "args": ["--from", "qikly[mcp]", "qikly-mcp"] } } } ``` Without uv, install with pip and register the `qikly-mcp` command instead: ```bash pip install "qikly[mcp]" claude mcp add qikly -- qikly-mcp ``` `uvx` keeps the version it first installed. To pick up a new release, run `uv cache clean qikly` and restart the host. Run it from the project directory, the one holding `inputs_private/`. That is how it finds your tasks, and it is the same rule the CLI follows. ## One-click install, and which editors have it The README badge installs into **VS Code**. It writes the `uvx` command above, so it needs uv and installs nothing itself. The same redirect serves Insiders from its own host: ```markdown [](https://insiders.vscode.dev/redirect/mcp/install?name=qikly&config=%7B%22name%22%3A%22qikly%22%2C%22command%22%3A%22uvx%22%2C%22args%22%3A%5B%22--from%22%2C%22qikly%5Bmcp%5D%22%2C%22qikly-mcp%22%5D%7D) ``` A common mistake in other projects' READMEs is pointing both buttons at `insiders.vscode.dev`. Stable is `vscode.dev`, Insiders is `insiders.vscode.dev`, and the payload is identical. **Other editors have their own deeplink schemes**, and qikly does not yet ship buttons for them because none has been tested against a real install here: | Editor | Scheme | | --- | --- | | Visual Studio | `vs-open.link/mcp-install` | | Cursor | `cursor://anysphere.cursor-deeplink/mcp/install` | | Goose | `goose://install-mcp` | | LM Studio | `lmstudio://add_mcp` | They differ in how the config is encoded, and at least one uses base64 where VS Code uses URL-encoded JSON. A button that writes a malformed config is worse than no button, because the reader blames the tool rather than the link. Until each is tried, use the plain configuration above: **every one of these editors accepts a hand-written config**, and the button only ever saves you a paste. ## Let it register itself ```bash qikly --install-mcp --dry-run # what it would write, and where qikly --install-mcp # write it ``` It writes **project-local files only**: `.mcp.json` for Claude Code and `.vscode/mcp.json` for VS Code. Never `~/.claude.json`, never the VS Code user profile. Those hold every other server you have, and the project file is also the right place on its own terms, because the setting that reliably goes wrong is which folder the server treats as the project. It merges rather than overwrites, so your other servers and every other key in the file survive, and it copies the file to a timestamped backup first. Running it twice changes nothing and says so. It refuses in three cases, and prints the block for you to paste instead: | It stops when | Because | |---|---| | the file holds comments | VS Code's `mcp.json` is JSONC, and writing it back as JSON would delete every comment you wrote | | a `qikly` entry exists and differs | you changed it on purpose. `--force` says otherwise | | the file is not valid JSON | guessing what you meant is how a config gets lost | **If `claude` is not a command**, it is bundled inside the VS Code extension rather than installed on `PATH`, at `%USERPROFILE%\.vscode\extensionsnthropic.claude-code-<version>-<platform> esources ative-binary\claude.exe`. The version is in the path, so it moves on every extension update. You do not need the CLI for any of this; it is only how you check the registration from a terminal. **Claude Code needs one approval after this.** A project-scoped `.mcp.json` is not trusted on sight, and it should not be: cloning a repository must not silently run whatever it ships. `claude mcp list` will show qikly as pending until you run `claude` once in that directory and approve it. Add `--install-mcp claude` or `--install-mcp vscode` to target one host. Cursor and Codex CLI are not written yet. Cursor takes the `mcpServers` shape below; Codex needs TOML, which the standard library cannot write on any Python this supports. ## Two things that will bite you on Windows Both were hit on a real machine before anyone else saw them, and neither announces itself clearly, so they are worth reading before you debug. **The command has to be on `PATH`.** The uv route and the one-click button run `uvx`, which the uv installer puts on `PATH`, though an editor opened before you installed uv will not see it until you restart the editor. The pip route names a bare `qikly-mcp`, and on Windows `pip` often installs console scripts into a `Scripts` directory that is not on `PATH`, and the host then reports only that the command was not found. Check it: ```powershell Get-Command qikly-mcp ``` If that finds nothing, let qikly print a config that does not depend on `PATH`. In your project folder: ```powershell python -m qikly --mcp-config ``` It prints a VS Code config naming the full path of the Python that has qikly, run as `python -m qikly.mcp_server`, and the folder you ran it from as `QIKLY_PROJECT_ROOT`, which also settles the folder problem below. `--mcp-config claude` prints the `mcpServers` shape for Claude Code, Cursor and most other hosts. It only prints; it never edits a config file, because those hold your other servers too. Use `python -m qikly` rather than `qikly` here: when `qikly-mcp` is not on `PATH`, neither is `qikly`. **The host starts the server in the folder your editor has open**, which is often the parent of your qikly project rather than the project itself. The server starts, and every task looks absent. Name the project explicitly: ```json { "servers": { "qikly": { "type": "stdio", "command": "qikly-mcp", "env": { "QIKLY_PROJECT_ROOT": "C:\\path\\to\\your\\project" } } } } ``` The tools detect this case rather than reporting an empty project: if the resolved root has no `inputs_private/`, they say so and name the directory they resolved, because that one fact is the difference between a misconfiguration and an apparently broken tool. ## The four tools | Tool | What it achieves | Input | Returns | Safe to call unattended? | CLI equivalent | | --- | --- | --- | --- | --- | --- | | `qikly_run` | Starts a full run for one task: generates the suite from the acceptance criteria, writes an implementation from the requirements alone, and repairs it against the suite until every stage passes or the retry budget is spent. Returns **immediately**, because a run takes minutes to hours. | `task_id`, optional `provider` and `model` | A `run_id` to poll with | **No.** It writes code and tests, calls a model provider over the network, and spends real money. A second call is a second run, not a repeat: the agents are not deterministic even at a fixed seed. | `qikly --run <task>` | | `qikly_status` | Reports where a run got to, reading what the run itself wrote rather than watching the process. That is what lets it tell a crash from a failing suite: `stalled` means the process is gone without writing a summary. While a run is in flight it also gives the stage and iteration. | `run_id` | One of `running`, `passed`, `failed`, `stalled`, `unknown`, plus stage and iteration | **Yes.** Reads local run records only. Free, no network. | `qikly --status <run_id>` | | `qikly_validate` | Checks a task file offline before you spend anything on it: valid YAML, the required sections present, `acceptance_criteria` a list rather than one long string, and fixture paths that actually resolve. It does **not** look for contradictions between requirements and criteria: that is `qikly --check-criteria`, a paid model call this server deliberately does not expose. | `task_id` | Counts and a verdict | **Yes.** No model call, so no network and no cost. Returns no criterion text. | `qikly --validate --tasks <task>` | | `qikly_scaffold` | Turns a Python file you already have into a task that tests it: the module path, the real signatures of its public functions, a guessed entrypoint, and a seed pointing back at the file. `requirements` and `acceptance_criteria` are left as `TODO` on purpose, because criteria read out of an implementation can only describe what it already does. | `file_path` | The task YAML as text, for you to save under `inputs_private/config/tasks/` | **Yes.** Returns the YAML rather than writing it, so saving stays your decision. | `qikly --scaffold <file>` | Every tool carries the four MCP behaviour annotations, so a host can act on the column above rather than guess: `readOnlyHint`, `destructiveHint`, `idempotentHint` and `openWorldHint`. Three of the four are read-only, free and local. Only `qikly_run` is none of those things. ## Why `qikly_run` does not wait A run takes minutes to hours. No MCP host will hold a tool call open that long, so `qikly_run` starts the run in a detached process and hands back an id: ``` qikly_run(task_id="CALC_TAX") -> {"run_id": "CALC_TAX_20260910_113412", "state": "running"} ``` The agent then polls: ``` qikly_status(run_id="CALC_TAX_20260910_113412") -> {"state": "running", "progress": {"stage": "unit", "iteration": 2}} ``` Because the run is detached, you can close the terminal, close the editor, and ask for the status tomorrow. The run outlives the conversation that started it. `stalled` is worth knowing about: it means the run died without writing a summary, which is a crash rather than a failing suite. The two need different reactions, so they get different words. ## What these tools will not tell you No qikly tool returns your acceptance criteria. Not on success, and especially not on failure, which is exactly when a helpful tool wants to explain why and "why" is the criterion. That is the point of the whole library, so it is enforced rather than intended: `tests/test_mcp_withholding.py` asserts on the serialised JSON that crosses the wire, for every tool, including the error paths. Your agent will see that four tests failed, and the pytest output for each. It will not see the rule it broke. That is the same information a human developer gets from a failing CI run, and it is what stops the agent writing code aimed at the test instead of at the requirement. ## One thing you have to do yourself These tools control what a *response* contains. They cannot control what your agent reads off your disk, and **the generated tests under `outputs/tests/` are written from your criteria**. An agent that opens those files has the answer key, without any tool call being involved. This matters more than it first sounds, because the natural way to use this server is to let the same agent both write your code and call `qikly_run`. So tell your agent to leave the generated tests alone. In Claude Code, add to `.claude/settings.json`: ```json { "permissions": { "deny": ["Read(./outputs/tests/**)"] } } ``` Or add `outputs/tests/` to whatever your host uses to keep files out of an agent's reach. Reading the *summary* is fine and useful. Reading the tests is the one habit that quietly undoes the reason you installed this. -
PROVIDER_KEY_SETUP.md 8 KB
# Setting up a provider key Every qikly run calls one model provider, and that provider needs a key. This page covers all three, on PowerShell, bash and CI, plus how to check a key took and how to stop paying for one you forgot about. If you only want the shortest path: get a [Gemini key](https://aistudio.google.com/apikey), which has a free tier and needs no card, then `qikly --demo`. --- ## Which provider | Provider | Key from | Install | Notes | |---|---|---|---| | **Gemini** (default) | [aistudio.google.com/apikey](https://aistudio.google.com/apikey) | included | Free tier, no card. Cheapest by roughly 20x. Every published qikly figure was measured on it | | **OpenAI** | [platform.openai.com/api-keys](https://platform.openai.com/api-keys) | `pip install "qikly[openai]"` | Requires billing set up | | **Anthropic** | [console.anthropic.com/settings/keys](https://console.anthropic.com/settings/keys) | `pip install "qikly[anthropic]"` | Requires billing. No seed and no temperature, so runs vary more than the others. A reasoning model, so it produces thinking tokens you are billed for on top of the answer | You do not have to choose one forever. All three keys can sit in your environment at once; `LLM_PROVIDER` decides which is used, and if you set no provider and hold exactly one key, qikly uses that one and says so. The recipes below set it anyway, because each one is about a named provider. With a single key you can leave it out, which is why the quick start does not set it. **Restricting the key.** qikly calls exactly one endpoint per provider, so a minimal key is enough. On OpenAI, choose **Restricted** and grant only **Chat completions (`/v1/chat/completions`)**; every other row stays `None`. Read-only will not work, because creating a completion is a write. --- ## Which model Pick a small fast one. A run makes one model call per stage per iteration, so the length of a run is mostly the length of a single call, and a model that reasons before it answers turns a two minute run into twenty. | Provider | Start with | Costs you a long wait | |---|---|---| | **Gemini** | `gemini-3.5-flash-lite`, the default, or any flash model | `gemini-3.5-pro` | | **OpenAI** | a mini model | the reasoning models | | **Anthropic** | `claude-haiku-4-5` | `claude-sonnet-5`, which thinks before every answer and bills the thinking | qikly says this in the run banner when the model you chose is one that thinks, and again the first time a call runs long, so a slow run explains itself rather than looking like a hang. **Why the smallest model is the one every figure was measured on.** Every published qikly number comes from `gemini-3.5-flash-lite`, over hundreds of runs. That is the cheapest and least capable of the three defaults, and the choice was deliberate: a number measured there is a floor rather than a best case. A later flash model should clear it, not fall short of it. If yours does not, that is a result worth reporting. None of which makes a thinking model wrong. It writes a stricter suite, and if that is what you are after, the wait is what it costs. It does mean not to reach for one first, and that a run which seems stuck is usually this. --- ## Windows PowerShell ### Gemini ```powershell $env:GEMINI_API_KEY = "AIza..." $env:LLM_PROVIDER = "gemini" Write-Output "provider: $env:LLM_PROVIDER" Write-Output "key: $($env:GEMINI_API_KEY.Length) chars, ends ...$($env:GEMINI_API_KEY.Substring($env:GEMINI_API_KEY.Length-4))" qikly --demo ``` ### OpenAI ```powershell pip install "qikly[openai]" $env:OPENAI_API_KEY = "sk-proj-..." $env:LLM_PROVIDER = "openai" Write-Output "provider: $env:LLM_PROVIDER" Write-Output "key: $($env:OPENAI_API_KEY.Length) chars, ends ...$($env:OPENAI_API_KEY.Substring($env:OPENAI_API_KEY.Length-4))" qikly --demo ``` ### Anthropic ```powershell pip install "qikly[anthropic]" $env:ANTHROPIC_API_KEY = "sk-ant-..." $env:LLM_PROVIDER = "anthropic" Write-Output "provider: $env:LLM_PROVIDER" Write-Output "key: $($env:ANTHROPIC_API_KEY.Length) chars, ends ...$($env:ANTHROPIC_API_KEY.Substring($env:ANTHROPIC_API_KEY.Length-4))" qikly --demo ``` The check prints the length and last four characters rather than the key. A plain `$env:OPENAI_API_KEY` writes the whole secret into your terminal scrollback and your PSReadLine history file, which is a bad habit for something that bills you. ### Making it stick `$env:` lasts for that window only. Close the terminal and the key is gone. ```powershell [Environment]::SetEnvironmentVariable("OPENAI_API_KEY", "sk-proj-...", "User") ``` Then **open a new terminal**: the current one will not see it. ```powershell # what is persisted, as opposed to only set here [Environment]::GetEnvironmentVariable("OPENAI_API_KEY", "User") # everything qikly might pick up, in this session Get-ChildItem Env: | Where-Object Name -match 'API_KEY|LLM_' # remove one Remove-Item Env:\OPENAI_API_KEY ``` That last listing is the fastest way to find a stale `LLM_PROVIDER` from an earlier experiment, which is the usual reason a run goes somewhere unexpected. --- ## macOS and Linux ```bash # Gemini export GEMINI_API_KEY=AIza... export LLM_PROVIDER=gemini # OpenAI pip install "qikly[openai]" export OPENAI_API_KEY=sk-proj-... export LLM_PROVIDER=openai # Anthropic pip install "qikly[anthropic]" export ANTHROPIC_API_KEY=sk-ant-... export LLM_PROVIDER=anthropic qikly --demo ``` Check it took without printing it: ```bash echo "${#OPENAI_API_KEY} chars, ends ...${OPENAI_API_KEY: -4}" env | grep -E 'API_KEY|LLM_' | sed 's/=.*/=<set>/' ``` To persist, add the `export` lines to `~/.zshrc` or `~/.bashrc`, or keep them in a `.env` you source. Do not commit either. --- ## CI **GitHub Actions.** Store the key as a repository secret, never in the workflow file: ```yaml - uses: gal-a/qikly@v0.5.7 with: api-key: ${{ secrets.OPENAI_API_KEY }} provider: openai qikly-version: "qikly==0.5.7" ``` A secret is masked in logs. A literal is not, and a key pushed to a public repo is compromised within minutes, whether or not the commit is later removed. --- ## Checking it worked ```bash qikly --validate # free, and will NOT catch a bad key: it opens no socket qikly --demo # one task, real calls, a few cents ``` `--demo` is the real test. It reports the provider, the model, the estimated cost, and whether the task converged. **What a wrong key looks like:** ``` OpenAI API rejected the request: Error code: 401 ... If this looks like an auth error, check API_KEY for typos/whitespace ``` Whitespace is the usual culprit: a trailing space or newline copied along with the key. Compare the length you printed above against the length on the provider's console. **What a wrong model looks like:** ``` The model `gemini-3.5-flash-lite` does not exist or you do not have access ``` That is a provider and model that disagree. Check `LLM_MODEL` is unset or correct for the provider you chose. --- ## Cost, and how not to be surprised Every run prints an estimate before it starts and the real usage afterwards. The estimate comes from your own run history once you have some. A single `--demo` on the default provider is a fraction of a cent. The same demo on gpt-4o is roughly twenty times that, and it will be slower. Three ways to bound it: - `QIKLY_MAX_CALLS=50` stops a run dead after N model calls. - `QIKLY_MAX_OUTPUT_TOKENS` caps Anthropic's per-reply budget, default 16384. It has to cover thinking as well as the answer: at 4096 claude-sonnet-5 spent the entire budget reasoning and returned no answer at all. - Set a spend limit on the provider's own console. **This is the one that actually protects you**, because it applies whatever calls the API. - On OpenAI, put the key in its own project and cap that project. Then a mistake here cannot spend what you budgeted for something else. --- ## If you rotate or revoke a key Nothing in qikly stores a key. It reads the environment on every call, so revoking on the provider's console and exporting a new value is the whole procedure. There is no cache to clear and no config file to edit. -
QUICK_START_ON_YOUR_OWN_DATA.md 18 KB
# Quick start on your own data From nothing to a first run on your own module. Try the bundled demo, follow the five steps, then read the one question that decides where each line of your specification goes. That is the whole path, and it is the whole of this page. Everything else lives in [the task file reference](https://github.com/gal-a/qikly/blob/main/docs/TASK_FILE_REFERENCE.md): which command fits what you already have, the file field by field, criteria you wrote elsewhere, seeding your own code or tests, and the fixture rows a criterion needs before it can be checked at all. Come back for those when you want them; you do not need any of it for a first run. ## Try it first Requires **Python 3.10+** and **GNU `patch`** on `PATH`. On Windows it ships with Git under `usr\bin\patch.exe`, which the tool finds on its own. **On macOS you have to install it:** the system `patch` is Apple's BSD one, which rejects the options qikly sends, so no generated diff will apply. ```bash brew install gpatch # macOS only ``` qikly looks for `gpatch` before `patch`, so nothing else is needed afterwards and your system `patch` is left alone. **Install into a virtual environment**, and not only out of habit. Two reasons specific to this tool. qikly pulls in a provider SDK, so a bare install can upgrade a package your own project pinned. And qikly resolves where to read and write from the environment it is running in, so "which interpreter am I in" is a question you will want a clean answer to the first time something behaves oddly. `qikly --version` is that answer: it prints the version, the package directory and the interpreter together. ```bash python -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate pip install qikly qikly --version # which build, from where, on which Python export GEMINI_API_KEY=... # or OPENAI_API_KEY / ANTHROPIC_API_KEY qikly --demo ``` On Windows, in PowerShell, where `export` is not a command: ```powershell python -m venv .venv .venv\Scripts\activate pip install qikly qikly --version $env:GEMINI_API_KEY = "..." # or OPENAI_API_KEY / ANTHROPIC_API_KEY qikly --demo ``` Let the install finish before running anything. The provider SDK is a long dependency chain and pip installs qikly itself last, so there is a window where its dependencies are present and qikly is not, which looks exactly like a broken install. Only `--demo` needs a key. `--version`, `--explain`, `--validate`, `--init`, `--example` and `--scaffold` make no model call and cost nothing. Gemini is the default and its key is `GEMINI_API_KEY`. For OpenAI or Anthropic, which need an extra install, see [the provider table in docs/CONFIGURATION.md](https://github.com/gal-a/qikly/blob/main/docs/CONFIGURATION.md), which also says where to get each key. **The default model differs by provider**, and they are not the same size: Gemini gets `gemini-3.5-flash-lite`, OpenAI `gpt-4o`, Anthropic `claude-sonnet-5`, which is a reasoning model and takes minutes rather than seconds per run. [Which default you get, and how to change it](https://github.com/gal-a/qikly/blob/main/docs/TROUBLESHOOTING.md#provider-defaults). One key is enough and `LLM_PROVIDER` is optional: qikly uses the one key it finds and prints which. Set `LLM_PROVIDER` to `gemini`, `openai` or `anthropic` only when you hold more than one and want to choose. The key is a shell variable rather than a venv one, so set it once in the terminal and the venv sees it too. Full recipes, including making a key survive a new terminal, are in [docs/PROVIDER_KEY_SETUP.md](https://github.com/gal-a/qikly/blob/main/docs/PROVIDER_KEY_SETUP.md). `--demo` runs one task end to end in a throwaway `demo/throwaway_<timestamp>/` directory and prints what it built and where. It writes nothing outside that directory, so a first run leaves everything else untouched. About 30 seconds on the Gemini default, and [several minutes on the Anthropic one](https://github.com/gal-a/qikly/blob/main/docs/TROUBLESHOOTING.md#provider-defaults). After ten seconds of silence a run starts printing one line every fifteen saying how long it has been waiting, so a slow call is distinguishable from a hang. For the same thing on vehicle sensor data rather than an order pipeline: ```bash qikly --demo --tasks ADAS_HEADWAY ``` `ADAS_HEADWAY` checks following distance from forward-radar samples. Its requirements give the limits and the two-second rule; its withheld criteria pin what happens exactly at each limit, including that a gap of zero metres is not a measurement. From a clone instead: ```bash pip install -r requirements.txt python run.py --demo ``` `run.py` is a shim around `src/qikly/cli.py`, the same entry point the installed `qikly` command calls, so a clone and an install run identical code. ## Your own module, start to first run Rather see it work before you point it at your own code? ```bash qikly --example ``` **That is the five steps below, already done, on a module you do not have to write.** It lays down `my_metrics.py` at your project root, both task files a scaffold of it produces, and sample data at the paths those task files name. Its `requirements` and `acceptance_criteria` are written in, which a scaffold of your own code cannot do for you, so it runs as it stands: ```bash qikly --validate --tasks MY_METRICS_VERIFY # free, no model call qikly --tasks MY_METRICS_VERIFY ``` | The five steps below, on your module | What `--example` hands you instead | |---|---| | 1. Scaffold a task from the module | `my_metrics.py`, and the two task files a scaffold of it writes | | 2. Put your input data where the task says | two sample CSVs, at the paths those task files name | | 3. Write the two sections only you can write | written already, so you can read a finished pair before writing your own | | 4. Check it, for free | the same command | | 5. Run it | the same command | So the difference is step 3, and step 3 is the part that matters: those two sections are the whole mechanism, and reading a worked pair is the fastest way to see the split. [What it lays down, and what to look at in it](https://github.com/gal-a/qikly/blob/main/docs/SCAFFOLDED_TASK_EXAMPLE.md). It leaves the rest of your project alone, and never overwrites, so a second run keeps anything you edited. **Which starter is which.** `--init` writes `MY_FIRST_TASK`, an empty form with `TODO` where your rules go and a stub CSV, and you bring the code. `--example` writes `MY_METRICS`, the same form filled in, with a module and real sample data behind it. Neither command writes the other's task, so whichever you ran is the only task in your project. > **No module to start from?** You do not need one. `qikly --init` writes a > starter task with `requirements` and `acceptance_criteria` and no `seed:` > block, so a run writes the first implementation from your requirements > instead of testing code you already have. Skip step 1, fill in those two > sections, and the rest of the steps are unchanged. Scaffolding exists to read > signatures out of code that already exists, which is the only part you are > missing. 1. **Scaffold a task from the module.** ```bash qikly --scaffold my_metrics.py ``` It reads the real function signatures and writes a task file, then prints what to do next. The name comes from the module's own filename, upper-cased, and which of two task files you get depends on the job: | Command | Writes | What that task does | |---|---|---| | `qikly --scaffold my_metrics.py` | [`MY_METRICS_VERIFY.yaml`](https://github.com/gal-a/qikly/blob/main/src/qikly/inputs_public/examples/config/tasks/MY_METRICS_VERIFY.yaml) | Tests the code you already have | | `qikly --scaffold my_metrics.py --fresh` | [`MY_METRICS.yaml`](https://github.com/gal-a/qikly/blob/main/src/qikly/inputs_public/examples/config/tasks/MY_METRICS.yaml) | Writes a fresh implementation of the same interface, and tests that | Both land in `inputs_private/config/tasks/`. The `_VERIFY` suffix is what keeps them apart, so scaffolding the same module both ways never overwrites the first file with the second. The rest of this section uses `MY_METRICS_VERIFY` as the example; substitute your own. 2. **Put your input data where the task says.** Its `inputs:` list names the files a run reads, here `inputs_private/data/MY_METRICS/input_01.csv`. Both task files point at that one folder, named for the module rather than the task, so scaffolding both ways does not split your data in two. Scaffold does not create them, so copy a real sample of your data there. Until you do, `qikly --validate` reports `input file not found`. What a good sample looks like: [`input_01.csv`](https://github.com/gal-a/qikly/blob/main/src/qikly/inputs_public/examples/data/MY_METRICS/input_01.csv) and [`input_02.csv`](https://github.com/gal-a/qikly/blob/main/src/qikly/inputs_public/examples/data/MY_METRICS/input_02.csv). 3. **Write the two sections only you can write.** `requirements` holds the decisions and `acceptance_criteria` the consequences; the rule for telling them apart is under [Getting the two halves right](#getting-the-two-halves-right). Already written them in a page or a ticket? This takes the criteria from it: `qikly --scaffold my_metrics.py --from-doc feature.md`. Replace the remaining `TODO` lines too, and check the entrypoint scaffold marks as guessed. 4. **Check it, for free.** ```bash qikly --validate --tasks MY_METRICS_VERIFY ``` No model call and no cost. Without `--tasks` it also checks every bundled example. It checks that the file parses, that every input path exists, that no `TODO` placeholder is left, that criteria name values rather than adjectives, and that no requirement restates a criterion. 5. **Run it.** ```bash qikly --tasks MY_METRICS_VERIFY ``` Start reading at `outputs/reports/iterations/<task>_<timestamp>_report.html`. A `_VERIFY` task tests code you already have, which qikly did not write, so it cannot enforce that the code's author never saw your acceptance criteria. It evidences what it can: the run opens by asking git whether this task file was last changed before that code's first commit, and prints the answer with what it does not show. A commit date is not a writing date, and criteria settled first is the precondition of the opposite problem, someone coding to the bar. So read it as corroboration, and read the bad answer, criteria revised after the code landed, as the question it is. **Tried it?** [Tell us what happened](https://github.com/gal-a/qikly/discussions/6), whether it worked, stalled or never got past install. **Keeping the suite?** Add the badge to your project's README, so the people reading it know the tests were written by an agent kept apart from the code: ```markdown [](https://test.qikly.com/?ref=badge) ``` ### Step 1 from inside VS Code With the qikly MCP server connected (setup in [docs/mcp.md](https://github.com/gal-a/qikly/blob/main/docs/mcp.md)), ask Copilot in agent mode: > Use the qikly_scaffold MCP tool on `src/my_metrics.py`, and save the task it > returns under `inputs_private/config/tasks/`. It returns the same task the command writes, one that tests the code you already have, and says which filename to save it as. Steps 2 to 5 are the same, and `qikly_validate` runs step 4 from the chat at no cost. ### What a real run on your own code looks like Not the bundled demo. This is an ordinary module that totals invoice lines, scaffolded with `qikly --scaffold`, with the two human sections filled in by hand. The whole run cost **$0.003** and seven model calls. ``` [stage 1/3] [iteration 1] 3/4 passed | FAILED: test_quantity_boundary_zero_and_negative [stage 1/3] [iteration 2] 4/4 passed [stage 2/3] [iteration 1] 0/1 passed | FAILED: test_system_entrypoint_output_structure_and_types [stage 2/3] [iteration 2] 1/1 passed [stage 1/3] [iteration 2-regcheck] 4/4 passed [stage 3/3] [iteration 1] 5/5 passed ``` **Line 1 is the whole point.** The existing code validated that `qty` parsed as a number and stopped there, so a quantity of zero or minus one went through as a real order. One acceptance criterion said otherwise: ```yaml acceptance_criteria: - "A qty of 0 is rejected and a qty of 1 is accepted; a negative qty is rejected" ``` The suite was written from that criterion before the run touched the code, and the coding agent never saw the criterion itself. What it received was pytest's output for the failing test: the name, the test's own source and docstring, and the assertion error. From that it produced this patch: ```diff try: qty = int(row["qty"]) unit = float(row["unit_price"]) + if qty <= 0: + rejected.append(dict(row, reason="qty must be positive")) + continue except (ValueError, TypeError, KeyError): ``` That is a real bug in code that already existed, found by a test written from a rule the agent repairing it could not read. **The generated tests say where they came from**, so the suite is reviewable rather than a black box: ```python def test_quantity_boundary_zero_and_negative(): """Verify that a quantity of 0 is rejected, 1 is accepted, and negative quantities are rejected. # Requirements: 2, 4 # Criteria: 2 """ ``` **One honest note about the other criterion.** The same task asked for round-half-away-from-zero on currency, and the test for it passed on the first attempt against code using plain `round()`. Not because the code was right in general, but because on this particular input the binary representation of 10.005 lands just above the halfway point and `round()` returns 30.02 anyway. A criterion is only as good as the input that exercises it, which is the same point as [fixture coverage](https://github.com/gal-a/qikly/blob/main/docs/TASK_FILE_REFERENCE.md#proposing-fixture-rows) further down. ### Getting the two halves right This matters more than anything else in the file, and one question settles most of it. **`requirements` holds the decisions.** Anything a person chose that could have gone another way: a threshold, a unit, a measurement convention, an exemption. Nobody can guess a decision, so the coding agent has to be told. **`acceptance_criteria` holds the consequences.** What must be true if those decisions were implemented correctly: the exact boundary, the identity that has to hold, the case a careless reading gets wrong. They are withheld because they are the exam. > **Given only the requirements, could two competent developers legitimately > disagree about this line?** If yes, it is a decision and belongs in > `requirements`. If no, it is a consequence and belongs in > `acceptance_criteria`. | In `requirements`, because it is a decision | In `acceptance_criteria`, because it follows | |---|---| | Keep at least 2.5 m from the vehicle ahead, measured centre to centre | At exactly 2.5 m, no violation is raised | | Amounts are currency, rounded to the nearest cent | For every accepted row, total equals subtotal plus tax, exactly | | Dates are written YYYY-MM-DD | 2026-02-30 is rejected, because it is not a real date | **The gap between them is the entire mechanism**: with nothing withheld, both sides read the spec identically and every test passes first try, which proves nothing. **Both mix-ups have a signature. Learn to read them.** **A decision in `acceptance_criteria`** is withheld from the one agent that needed it, so the coding agent has to guess a choice nobody told it. It does not produce a harder test, it produces repetition, in one of two shapes. Either the same test fails while the FIX and PATCH come back near identical each time, because nothing the agent can see would lead it anywhere else, or two tests disagree and each patch makes one pass and the other fail. Restrict street suffixes to three valid values and the model keeps widening them, since everything it knows says "Boulevard" is a suffix. Sometimes, though, the agent simply guesses right and the run goes green. That is the worse outcome, because nothing then tells you a decision was in the wrong half. On a bundled task whose criteria alone settled whether exactly 2.00 seconds of headway raises a warning, three runs in ten converged anyway. Do not rely on the loop to find these for you; apply the question above when you write the spec. [TROUBLESHOOTING.md](https://github.com/gal-a/qikly/blob/main/docs/TROUBLESHOOTING.md#5-the-same-patch-appearing-over-and-over) has the diagnosis for the repeating case. **A consequence in `requirements`** is the quieter mistake. Both agents read the same boundary value, so the test that checks it passes on the first attempt and proves nothing. The rest of the suite is unaffected and still bites, which is what makes it easy to miss: the run looks entirely normal. Nothing fails, and nothing was learned about that boundary. **The loop cannot fix either one.** It can tighten a bar the code already attempts, and it cannot tell you a line is in the wrong place. An assistant reading your spec can help with half of it. Given the question above, it can say which line looks misplaced and why, and qikly's [agent Skill](https://github.com/gal-a/qikly/blob/main/docs/skill.md) exists partly to make it good at that. What it cannot do is settle the decision: whether exactly 2.00 s warns, or whether an empty string counts as missing, is not in the specification, which is what makes it a decision. Someone has to choose, and that someone is you. **To see a task that follows the rule,** run `qikly --explain CALC_TAX`. Its requirements say amounts are "currency amounts rounded to the nearest cent", and the withheld criteria pin what that already means, down to a float result of 434.99999999999994 reporting as 435.00. More in [docs/design_3_mechanism.md](https://github.com/gal-a/qikly/blob/main/docs/design_3_mechanism.md#using-it). ## Everything else The task file field by field, the four routes in, criteria you have already written, seeding your own implementation or suite, where each artefact lands, and fixture coverage: [task file reference](https://github.com/gal-a/qikly/blob/main/docs/TASK_FILE_REFERENCE.md). -
SCAFFOLDED_TASK_EXAMPLE.md 6.4 KB
# A scaffolded task, start to finish ```bash qikly --example ``` That puts the whole example in your project, ready to run: | Written to | What it is | |---|---| | `my_metrics.py` | The module you already have. Four functions, no implementation: scaffold reads signatures, never bodies. Each body carries a `TODO` saying what is expected there | | `inputs_private/config/tasks/MY_METRICS_VERIFY.yaml` | What `qikly --scaffold my_metrics.py` writes, with the two human sections filled in. Tests the code you already have, so it carries a `seed:` block pointing back at the module | | `inputs_private/config/tasks/MY_METRICS.yaml` | The same, as `qikly --scaffold my_metrics.py --fresh` writes it. A fresh implementation of the same interface, so no `seed:` block | | `inputs_private/data/MY_METRICS/input_01.csv` | Sample input data, first file | | `inputs_private/data/MY_METRICS/input_02.csv` | Sample input data, second file | It makes the `inputs_private/` layout first if you have not run `--init`, and leaves out that command's blank `MY_FIRST_TASK` starter, so the example is the only task in your project. It never overwrites, so running it twice is safe and your edits survive. Then: ```bash qikly --validate --tasks MY_METRICS_VERIFY # free, no model call qikly --tasks MY_METRICS_VERIFY ``` **What to expect from that run: it converged 9 times in 10.** Measured on 2026-09-23, ten runs from a fresh project each time, on the default model. The nine took 4 to 6 iterations and 26 to 52 seconds; the tenth exhausted its budget after 17 repair attempts and exited non-zero, which is the loop working rather than the example being broken. A first run is a draw, not a promise, and this one is a better draw than the bundled tasks at roughly 6 in 10 because the implementation is supplied and the specification is short. Every one of those runs did real work before it converged: the module ships as signatures with no bodies, so the first integration attempt fails every test and the coding agent writes the implementation from those failures. `--validate` reports the task clean, with nothing left to fill in. That is the one way this differs from scaffolding your own module: scaffold leaves `requirements` and `acceptance_criteria` as `TODO`, because neither can be derived from code, and here they are written so you can read a finished pair. [The quick start](QUICK_START_ON_YOUR_OWN_DATA.md#getting-the-two-halves-right) has the rule for telling the two apart. ## Why these files exist Nothing here is illustrative. Everything in the two YAML files that scaffold decides is the output of `qikly --scaffold` run on `my_metrics.py`: the task id, the module path, the signatures, the output path and the `seed:` block. `tests/test_scaffold_examples.py` regenerates both files and fails if any of that drifts, so what you open is what scaffold writes today rather than what it wrote once. What it cannot regenerate is `requirements` and `acceptance_criteria`, which are written here rather than left as `TODO`, and a second test holds them to having nothing left to fill in. The copies that ship live under [`src/qikly/inputs_public/examples/`](../src/qikly/inputs_public/examples), laid out the way a real project is rather than as a flat folder of samples, so the paths you read there are the paths `--example` writes them to. They sit under `examples/` rather than beside the bundled tasks deliberately: a scaffolded pair breaks two rules the bundled tasks keep. `MY_METRICS_VERIFY` ends in the suffix reserved for scaffolding, and the pair shares one data folder where bundled ids map one to one. Keeping it out of the bundled task list also keeps `qikly --validate` quiet for everyone who never asked for the example. Once `--example` copies it in, it is an ordinary task in your project like any other, and a bare `qikly` with no `--tasks` will run it. ## What to look at in the two task files They differ in one thing that matters. `diff` them and almost everything is the task id and the paths derived from it. The real difference is the last block: `MY_METRICS_VERIFY.yaml` ends with ```yaml seed: implementation: "my_metrics.py" ``` and `MY_METRICS.yaml` ends with a comment saying there is no `seed:` block, so a run writes the implementation itself. That one block is the whole difference between "check this code" and "write this code and check it". Both arrive with `requirements` and `acceptance_criteria` written, which is the one thing scaffold cannot do for your own module. Scaffold could read plausible criteria out of an implementation, and deliberately does not: criteria derived from code can only describe what that code already does, which is a bar it passes by construction. So when you scaffold your own module those two sections arrive as `TODO`, each carrying an `e.g.` showing the shape of a useful answer, and the pair here is what a filled in version looks like. Read the criteria against the requirements above them. Every one names a boundary value, and none of them restates a decision: that a humidity of 100 is the last usable one follows from the requirement that the range is 0 to 100, and could not be guessed from it by someone who had not been told where the range ends. ## About the sample data The `inputs:` list in a scaffolded task names one file, `inputs_private/data/MY_METRICS/input_01.csv`, because scaffold cannot know how many you have. Add the rest yourself; the example names both of its files, which is what that looks like once you do. Look at what is in them, because it is the part people get wrong. Between the two files there are readings that are clean, a missing temperature, a non-numeric temperature, a humidity of exactly 100 and one above it, a temperature of exactly 0.0, a negative temperature, and a malformed timestamp. Every one of those is the boundary named by one of the six acceptance criteria, and a test in `tests/test_scaffold_examples.py` fails if a criterion loses the row that reaches it. That is deliberate: **a criterion no row can trigger is a criterion nothing checks.** A suite written against data containing only clean readings will pass whatever the code does about bad ones, and report nothing about the rule you cared most about. Two thirds of the faults this project planted in its own research went unnoticed for exactly that reason. So when you copy your own data in, check that something in it reaches every criterion you wrote. These two files are sample rows, not a template: replace them wholesale with a real sample of your own. -
skill.md 11.5 KB
# qikly as an agent Skill **What this gets you:** your coding agent stops needing to be told about qikly. Ask it for tests you can trust and it reaches for the tool on its own, and it knows the part that is hard to guess, which line of your specification is a decision the coder needs and which is a consequence to withhold. A Skill is a folder of Markdown your agent reads when what you are asking matches what the Skill says it is for. No server, no configuration, no process to keep running. ## Install it **First, make sure the qikly you are about to run is the current one.** The Skill ships inside the package, so an old qikly writes an old Skill, and it then describes commands that do not exist yet. ```bash pip uninstall qikly -y pip install qikly qikly --version ``` Uninstall first rather than `--upgrade`: an interrupted or repeated upgrade can leave more than one version's metadata behind, and pip then reports one version while the files on disk are another's. The clean pair takes seconds and removes the question. Compare what `--version` prints against [the latest release](https://pypi.org/project/qikly/). Then, from the directory of the project you want the Skill in: ```bash qikly --install-skill ``` That writes `.claude/skills/qikly/`, which is where agents look. Then open your agent in that directory and ask for something ordinary, without mentioning qikly: > Write tests for `src/following_distance.py` that would actually catch a bug. `src/following_distance.py` stands in for a module you actually have; name a real one, because an agent asked about a file that does not exist will spend its answer asking you which file you meant. Every example on this page uses following distance from radar samples, which is the worked case in the Skill itself, so the two read together. It should reach for qikly by itself, and it does: first time in **Claude Code, Gemini CLI, Codex and Cursor**, in a project with a rival testing skill installed beside this one and a request that never mentioned qikly. That is the hard version of the test, because the agent had a competing option and no hint. Your own project has more skills in it than that one did. If it does not, see [if your agent does not pick it up](#if-your-agent-does-not-pick-it-up). **Using GitHub Copilot in VS Code?** Copilot reads none of the skill folders above. Run `qikly --install-skill copilot`, which writes the Skill into `.github/instructions/` along with the `qikly.instructions.md` file Copilot actually opens, then use Copilot Chat in Agent mode. That path has not been watched loading yet, so tell us if it works for you. The MCP server is the other way in and is tested there: [docs/mcp.md](https://github.com/gal-a/qikly/blob/main/docs/mcp.md). **A few options, none of them needed the first time.** `--dry-run` shows what it would write. `--force` replaces a Skill you have already edited, keeping a timestamped copy of the old one. And `--install-skill agents`, `cursor` or `gemini` write where those tools document their own skill folders, rather than Claude's. ## What is in it - **The one question**: given only the requirements, could two competent developers legitimately disagree about this line? Yes means it is a decision and belongs in `requirements`; no means it is a consequence and belongs in `acceptance_criteria`. - **Both mistakes and their signatures**, so an agent can recognise a stalled loop as a spec problem rather than a code problem. - **How to write a criterion that can be tested**: name values not adjectives, name both sides of a boundary, and make sure the data contains what you name. - **The five steps** on your own module, and which commands are free. - **What to do when a run does not converge**, and the one thing never to do, which is loosen a criterion to get green. - **The honest limit**, so an agent does not oversell it on your behalf. - **Two bundled references**: the task file reference and the troubleshooting guide, complete, so the Skill works with no network. ## Skill or MCP server, and why qikly has both **Skip this if you are not using the MCP server.** The Skill works on its own. This is here because the two look interchangeable and are not: they answer different halves of the same problem, and neither replaces the other. | | MCP server | Skill | |---|---|---| | What it ships | a running process exposing typed tools | a folder of Markdown | | What it gives the model | the ability to call `qikly_scaffold`, `qikly_validate`, `qikly_run`, `qikly_explain` | the judgement around those calls | | Setup | `qikly --install-mcp`, then host configuration | `qikly --install-skill` | | Can enforce a rule | **yes**, in server code | no, it is instructions | | Works when the agent has no shell | yes | no | **The MCP server is the one that can enforce things.** `qikly_run` is built so it cannot return the acceptance criteria, and a test in qikly's own suite fails the build if any code path lets a criterion through. That guarantee lives in code, and it is why the server exists. **The Skill is the one that can teach.** Which line of your spec is a decision and which is a consequence; that a repeating identical patch means a decision is in the wrong half; that a criterion naming a value your data never holds produces a test that passes whatever the code does. None of that is a function call, and an agent that does not know it will use qikly and get less out of it. **The guarantee does not rest on the Skill.** A Skill is text in a context window, so it can inform but not enforce. The withholding is enforced by the tool, in code, whether or not this Skill is ever installed. The Skill's own text says so, and a test asserts that it still does. ## If your agent does not pick it up Three causes, in the order to check them. **1. Your agent is not in that directory.** The Skill is per project. It is files on disk, so an agent started in another folder, or running in a browser with its own sandbox rather than on your machine, cannot see them. Start the agent in the directory you installed into. This is the commonest cause by some distance. **2. Just name it.** This always works, because it does not depend on the agent being told what the Skill is for: > Use the qikly skill to write tests for `src/following_distance.py`. If naming it works and the neutral request did not, the Skill is fine and the problem is discovery. **3. The host did not pass the description along.** Whether an agent reaches for a Skill unprompted depends on how much it explores before it starts typing, and on what the host told it. Claude Code reserves a fraction of the context window for the whole skill listing, 1% by default; when the listing does not fit, Anthropic's own bundled skills keep their descriptions and everything else is ranked by how often you have used it. So a skill you have never invoked can arrive as a bare name with nothing to match against. Raising `skillListingBudgetFraction` in your Claude Code settings gives the listing more room, and that single change turned four failed routing attempts into a clean one during testing. **Which is why naming it once is more than a workaround.** That ranking is by how often you have used each skill, decaying over about a week, so a skill you have never invoked sorts below every skill you have. Name it in one request and it moves above them, and the next neutral request has a much better chance of finding it by itself. The first invocation is the only hard one. ## Checking that it works A Skill either loads or it does not, and it never tells you which, so it is worth five minutes once. **If you only do one of these, do number four:** the others check that the Skill arrived, and that one checks that it is right. **1. The files are where your agent looks for them.** In the project you ran `qikly --install-skill` in: ```bash ls .claude/skills/qikly/SKILL.md # macOS, Linux dir .claude\skills\qikly\SKILL.md # Windows PowerShell ``` **And start your agent in that same directory.** The Skill is per project, not per machine: an agent started somewhere else, or running in a browser with its own sandbox rather than on your computer, cannot see these files and will never load them. That is the commonest reason a correctly installed Skill appears to do nothing. **2. It loads on a request that should trigger it.** Start a fresh session and ask for something in its territory, **naming a real module of your own** and not mentioning qikly: > Write tests for `src/following_distance.py` that would actually catch a bug in it. The agent should mention qikly, or the decisions-and-consequences split, unprompted. If it does not, see [when an agent does not pick it up](#if-your-agent-does-not-pick-it-up) below before changing anything. **3. It does not load when it should not.** Ask something unrelated, such as "rename this variable everywhere", and it should stay quiet. A Skill that loads for everything costs context on every request. **4. It gives the right answer to the question that matters.** This is the one to do if you do only one. Ask: > My spec says "warn when following distance breaks the two-second rule", and > the acceptance criteria say a headway of exactly 2.00 s does not warn. Is > that the right split? The answer should be no, and the reason should be that "breaks" can be read as "below" or "at or below", so the boundary is a decision and belongs in the requirements. That is the single most valuable thing in the Skill, and if it comes back wrong the rest does not matter much. **5. It does not claim more than the tool does.** Ask what guarantees the coding agent never sees the criteria. The answer should point at qikly's code and its build-failing test, not at the Skill. ## Keeping it current **An installed Skill does not update itself, and until 0.5.4 nothing told you.** `pip install --upgrade qikly` replaces the package; it cannot touch a folder copied into your project, so after an upgrade you can be following instructions that name a different set of commands. From 0.5.4 any qikly command says so when it notices: ``` note: the qikly Skill in .claude\skills\qikly is older than this qikly, so it describes a different set of commands. `qikly --install-skill --force` replaces it and keeps a copy of the old one. ``` The path is printed with your platform's own separator, so it reads with forward slashes on macOS and Linux. `--force` keeps a timestamped copy of what it replaces, so a Skill you have edited is recoverable. That is the whole update mechanism: qikly tells you, and you run one command. There is no background process and nothing phones home; the check is two version strings read from two files on your disk. The Skill's version is its own and moves when its instructions move, not when qikly releases, so a release that does not touch it produces no notice. --- The Skill itself lives at [`src/qikly/skills/qikly/`](https://github.com/gal-a/qikly/tree/main/src/qikly/skills/qikly) and ships inside the installed package, so `--install-skill` works from `pip install qikly` as well as from a clone. **A stale Skill is worse than no Skill**, because an agent quotes it with confidence and the reader has no way to tell. So it is updated whenever qikly changes in a way it describes: a new or renamed flag, a change to which commands are free, a re-measured convergence figure, a new failure mode worth carrying. Half of it cannot go stale on its own. The two bundled references are asserted byte-identical to `docs/TASK_FILE_REFERENCE.md` and `docs/TROUBLESHOOTING.md` by a test, so a drifted copy fails the build. The prose in `SKILL.md` has no such test, which is why it is on the release checklist instead. -
TASK_FILE_REFERENCE.md 21.9 KB
# Task file reference The parts of working on your own data that you look up rather than read through. The path to a first run is [the quick start](https://github.com/gal-a/qikly/blob/main/docs/QUICK_START_ON_YOUR_OWN_DATA.md); this is what it deliberately leaves out. ## You probably do not have to write the task file by hand The criteria usually exist already, in a feature page or a ticket, and the interface exists in the code. qikly reads both. ```bash # a markdown page, a ticket export, or a .feature file qikly --criteria-from feature.md --task-id MY_TASK # straight from Jira: needs JIRA_BASE_URL, JIRA_EMAIL, JIRA_API_TOKEN qikly --criteria-from-jira PROJ-412 --task-id MY_TASK # both halves at once: criteria from the page, interface from the module qikly --scaffold src/metrics/band.py --from-doc feature.md ``` Bullet lists, a headed `Acceptance Criteria` section and Gherkin `Scenario:` blocks are all understood. Your page stays the source of truth and nobody retypes anything. **One section is never filled for you: `requirements`.** The coding agent reads it, and a feature page usually restates its own acceptance criteria in the prose above them, so lifting requirements across would hand the criteria to the one agent that must never see them. `qikly --validate` warns if what you write there restates a criterion. ## Which command depends on which parts you already have A task file is one YAML file with three parts, and the split above is a split between them: 1. **`requirements`** what the code must do, in the words a person would use. The coding agent reads this. 2. **`interface`** the contract, and a description rather than code: the function signatures and the dotted path where the module will live. Both agents read it, and neither is handed an implementation to read from it. When the integration and system tests are written there is not one yet. 3. **`acceptance_criteria`** what counts as correct, each one checkable and naming its boundary value. **Only test generation reads this.** "Spec" below means 1 and 2 together, which is what the coding agent is given. A tick means you already have that part. One thing the three parts do not say, and it matters: **test generation never reads the implementation either.** Integration and system tests are written before any code exists, from the specification alone. The unit stage is the single exception, written last from the code that just cleared the earlier stages, because unit tests have to name real functions. | Where you are starting | #1 | #2 | #3 | Run | What happens | |---|:-:|:-:|:-:|---|---| | Before anything else: see what is withheld | | | | `qikly --explain <MY_TASK>`<br>e.g. `qikly --explain CALC_TAX` | Prints a task file twice, once as each agent receives it, and the difference between them. No API key, no model call, about a second. **You get:** the acceptance criteria on one side and the same file with them cut out on the other, which is the claim everything else rests on. Add `--html` for the same as a page you can share. | | Just looking | | | | `qikly --demo` | A bundled task end to end in a throwaway folder. Thirty seconds, under a cent. **You get:** a working implementation, three test suites, and the full record of every FIX and PATCH, in a directory you can delete. | | Code someone else wrote, and you want **that code** verified | | Y | | `qikly --scaffold <MY_MODULE>.py` | Scaffold reads the real signatures out of the file you point it at and fills in **#2** for you. **#1** and **#3** stay yours to write: criteria read out of an implementation can only describe what that implementation already does, which is a bar it passes by construction. **You get:** one task file that tests the code you already have. Add `--fresh` for one that writes a fresh implementation of the same interface instead. | | You know what it must do, not yet how to check it | Y | | | `qikly --init` | Creates the directory layout and one starter task to edit. Its criteria show the habit that matters most: name the value, not the quality. "100 is accepted and 101 is rejected" forces a test at the boundary; "amounts must be reasonable" does not. **You get:** a task file to fill in, with your fixtures where a run will look for them. | | Same, but you want a first draft of the bar | Y | Y | | `qikly --tasks <MY_TASKS>`<br>`--generate-criteria` | Drafts **#3** from **#1** alone, then runs. **You get:** a first draft of the bar written into your task file for you to correct, plus the implementation and suites. | | You have written all three | Y | Y | Y | `qikly --tasks <MY_TASKS>` | Everything you wrote is used, and nothing is drafted on your behalf. **You get:** an implementation, integration, system and unit suites, a convergence report, and a run summary recording the model and settings that produced them. | | You have all three but doubt they agree | Y | Y | Y | `qikly --check-criteria`<br>`--tasks <MY_TASKS>` | One model call asking whether any implementation could satisfy the description, **#1** and **#3** at once, and whether any two of **#3** agree with each other. Advisory, and exits non-zero on a contradiction so a pipeline can gate on it. **You get:** a list of the pairs that cannot both hold, before spending a stage budget on them. Two criteria setting different numbers on the same quantity are always reported, since that is a typo rather than a tighter bar. | | A previous run stopped before finishing | Y | Y | Y | `qikly --tasks <MY_TASKS>`<br>`--resume` | Generating the tests and the first implementation already cost model calls, and they are still on disk. This keeps them and picks up where it stopped, instead of paying for them twice. **You get:** the same outputs as a full run, without paying for the parts already built. | `<MY_TASKS>` is one task_id or several separated by commas. A task_id is a filename under `inputs_private/config/tasks/` without the `.yaml`: `--tasks CALC_TAX`, `--tasks CALC_TAX,MERGE_SALES`, or omit it to run every task found. `<MY_TASK>`, singular, takes exactly one. `QIKLY_MAX_CALLS=200 qikly` stops at a call limit rather than a bill. `--scaffold` reads the module path and the real signatures of every public function straight out of the file, because they are already there. It leaves `requirements` and `acceptance_criteria` for you, and that is deliberate: criteria derived from an implementation can only describe what that implementation already does, and a bar that agrees with the code by construction is the exact failure this tool exists to prevent. ## Writing a task by hand Nothing is written into the package, and nothing is written into your source tree. ### 1. Make the two directories Anywhere you want to work. The presence of `inputs_private/` is what marks a directory as your project. ```bash mkdir -p inputs_private/config/tasks mkdir -p inputs_private/data/MY_TASK ``` Or let `qikly --init` create both, plus a starter task to copy. Until one of those exists, there is nothing marking your directory, and the fallback in [Where things live](#where-things-live) applies. From a `pip install -e` checkout that fallback finds the checkout itself, so a run started in an empty directory writes its outputs there instead of where you are standing. Make the directory first, or set `QIKLY_PROJECT_ROOT` to say exactly where you mean. ### 2. Drop your fixture data in Plain input files, whatever your code should read. CSV, JSON, JSONL, anything. ```bash cp ~/somewhere/orders_jan.csv inputs_private/data/MY_TASK/input_01.csv cp ~/somewhere/orders_feb.csv inputs_private/data/MY_TASK/input_02.csv ``` Names are up to you, but they must match what you write in the task's `inputs:` list below. The generated program opens these **by literal relative path from your project directory**, so the path in the task file is the path that gets executed. That is also why the bundled fixtures are copied into `inputs_private/data/` on first run rather than resolved from inside the package: the generated code has no way to ask where the package lives. Fixtures are never overwritten once present, so an edited file stays edited. ### 3. Write the task file `inputs_private/config/tasks/MY_TASK.yaml`. The filename must match `task_id`. ```yaml task_id: "MY_TASK" # letters/digits/underscore, not starting with a digit task_name: "Order line-item tax" description: "Read two CSV files of order line items, validate them, compute tax per line, and write the result to a single JSON output alongside a reason for every rejected line." inputs: # literal paths, opened by the generated code - "inputs_private/data/MY_TASK/input_01.csv" - "inputs_private/data/MY_TASK/input_02.csv" outputs: - "outputs/data/MY_TASK/output.json" interface: # what test generation targets module: "outputs.agent_src.code.MY_TASK.calc" integration_functions: - "extract(input_path) -> list[dict] # reads one input file, returns raw rows" - "transform(rows) -> dict # validates and computes; returns {\"accepted\": [...], \"rejected\": [...]}" - "load(data, output_path) -> None # writes the result as JSON" system_entrypoint: "run_calc(input_paths, output_path) -> None # extract each path, then transform -> load" requirements: # THE DECISIONS. The coding agent sees only this. - "Read both CSV files listed in inputs and combine their rows before validation" - "Validate each row: order_id, item_price, quantity, tax_rate" - "Apply strict, real-world data-quality validation; reject anything malformed or out of range" - "For each valid row compute subtotal, tax owed, and line total as currency amounts" - "A rejected row is not silently dropped: record it with a brief, specific reason" - "Write a single JSON object with two keys, \"accepted\" and \"rejected\"" acceptance_criteria: # THE CONSEQUENCES. Withheld from the coding agent. - "All computed currency amounts are rounded to two decimal places using round-half-up, not banker's rounding and not truncation" - "For every accepted row, the reported total equals the reported subtotal plus the reported tax, exactly, to the cent" - "A tax_rate of exactly 0 is valid: the computed tax is 0.00 and the total equals the subtotal" - "Each rejected row names the specific field that caused rejection, not a generic message" ``` `interface.module` is the dotted path the generated tests will import. With no seed, or with a single-file seed, that is `outputs.agent_src.code.<task_id>.<name>`: the implementation is written there, so pick the final component freely and the rest is fixed by where outputs live. **A package seed is the exception**, because the package keeps its own name and becomes importable by it: write `interface.module: "mypkg.pricing"`, the path your own code already uses. See "Seeding a package" below. ### 4. Run it ```bash qikly --tasks <MY_TASKS> # or: python run.py --tasks <MY_TASKS> ``` Discovery is automatic; there is no registry to update. Results land in `outputs/`, and `outputs/reports/iterations/MY_TASK_<timestamp>_report.html` is the place to start reading. ## Bringing acceptance criteria you have already written Most teams have not got a blank page here. If you work in Jira, Linear, Azure DevOps or a design doc, the rules are usually already written down, because the process asks for them before any code is cut. A ticket routinely looks like this: ``` PROJ-412 Merge overlapping sales exports Description Combine two CSV exports into one file... Acceptance Criteria - A transaction in both files at the same amount appears once - A negative or missing amount is rejected, naming the field - Dates must be YYYY-MM-DD ``` Those bullets are exactly what `acceptance_criteria` wants. Save the ticket to a file and read them out: ```bash qikly --criteria-from ticket.md # print as YAML qikly --criteria-from ticket.md --task-id MY_TASK # write into that task qikly --criteria-from ticket.md >> inputs_private/config/tasks/MY_TASK.yaml ``` It understands plain bullet lists, an "Acceptance Criteria" heading in a longer document, and Gherkin `Scenario:` blocks with Given/When/Then. Only the criteria section is read, so pasting a whole ticket does not turn its description into part of the bar. Only YAML goes to stdout, so the third form above appends a valid block. **It will not invent criteria from prose.** A file with no list and no scenarios returns nothing and says so. A rule that nobody wrote is precisely the invented standard this tool exists to argue against, and once it is in the file it looks like every other line. There is no API token and no vendor integration involved. Copying the ticket into a file is the whole of it. **Read what comes out before you run.** Criteria lifted from a ticket are a draft: tickets are written for people, who fill in gaps that a test cannot. The criteria are the standard everything else is judged against, so they are worth a minute of your attention. ## Supplying your own acceptance criteria, code or tests The loop takes three inputs. **Each one can be yours or generated, independently and in any combination.** | Input | Default | To supply your own | |---|---|---| | **Acceptance criteria** | Yours | Already the default: write `acceptance_criteria` in the task file, as above, or lift them from a ticket with `--criteria-from` (below). Omit it and add `--generate-criteria` to have a first draft written for you instead. | | **Implementation** | Generated | `seed.implementation` in the task file. | | **Test suites** | Generated | `seed.tests`, per stage. | ### Auto-generating acceptance criteria A task with no `acceptance_criteria` still runs, with a warning rather than an error, because running one deliberately is a legitimate thing to do. What you lose is the point of the exercise: test generation has only `requirements` to work from, the coding agent has nothing sharper to fail against, and the run usually converges on the first attempt without exercising the loop at all. `--generate-criteria` writes a first draft from the requirements alone into `inputs_private/config/tasks/<task_id>.yaml` before the run starts. It is opt-in, it never touches a task that already has criteria, and it says what it wrote rather than editing your files quietly. Pointed at a bundled example it writes your own overriding copy and leaves the packaged original alone. **A generated bar is a draft, not ground truth.** It was written from the same requirements the coding agent reads, so a case it did not think to demand is not being withheld from anyone: the two halves agree because they came from one source, which is the failure mode this whole tool argues against. `--compare-criteria` scores a generated bar against yours when you want that difference measured rather than assumed. The optional `seed:` block: ```yaml seed: # A file or a directory, copied into outputs/agent_src/code/<task_id>/. # A single file keeps its own name, which must match interface.module. # A directory is treated as a package: see "Seeding a package" below. implementation: "seeds/MY_TASK/calc.py" # Per stage. Seeding a stage suppresses generation for that stage only. tests: integration: "seeds/MY_TASK/test_integration.py" unit: "seeds/MY_TASK/unit/" ``` Paths are relative to your project directory. Both keys are optional. **`seed.implementation` is how you point this at code you already have.** The run skips generating a first implementation and goes straight to testing and repairing yours. `--scaffold` writes this block for you by default, and `--fresh` writes a task without it, for a new implementation of the same interface. **`seed.tests` keeps a suite you already trust**, so the loop repairs the code against your tests rather than its own. Mixing works and is often what you want: seed the integration stage with your suite and let the tool generate unit tests against whatever code results. Three things to know: - **Seeded test suites are checked before the run starts.** Every file must parse, and at least one must be named `test_*.py` and contain a `def test_*` function. A problem raises immediately rather than retrying, since there is no second sample to draw from a file you wrote. - **Seeds are installed after the workspace reset, not instead of it.** Every run still begins from one declared state, so repeated runs stay comparable and no run inherits the previous one's residue. - **A seeded run measures something different from an unseeded one.** Do not pool them in a single rate. The orchestrator prints a NOTE on every seeded run to keep that visible. ### Seeding a package, when the implementation is more than one module **New in 0.5.5.** Point `seed.implementation` at a directory and it is treated as a package: it is copied in **under its own name**, and the task's code directory is put on the path for the test run, so the package resolves by the name your code already uses. ```yaml interface: module: "mypkg.pricing" # the module under test, by its real import path seed: implementation: "mypkg" # the package it lives in, copied in whole ``` Inside `mypkg/`, write imports exactly as you already do. All three shapes work: `from mypkg.money import to_cents`, `from .money import to_cents`, and `from .utils.rounding import half_up`. No `__init__.py` is required, and one that is there is kept. **Every module in the package is visible to the coding agent and every one is repairable**, and a single patch may change more than one of them. That is the difference the package form makes: with a single-file seed, a defect in a helper is found by the tests and cannot be fixed, and the run tells you so rather than working around it. **The directory you name is the boundary.** qikly does not follow imports and decide for itself which of your files an agent may rewrite, because the transitive closure of a real package has no natural edge and "it rewrote a shared module I never named" is a worse outcome than naming a folder. So put inside the seed what you want worked on, and leave a vendor library or a module you do not want touched outside it. Data files inside the package are copied too, since your code may open them. `__pycache__` and `.pyc` files are not: they are stale copies of the very modules the run is about to rewrite. Four limits worth knowing before you start: - **Before 0.5.5 a seeded directory was flattened**, dropping the folder's name. If you wrote a task against that behaviour, `interface.module` needs the package name adding to it. - **Modules that import each other circularly at the top level fail**, the same way they do in plain Python. This is not something a run can repair. - **The package's name may not be a standard-library module's name.** A package called `json` would shadow the real one for everything the run imports, so a seed naming one is refused with a message rather than discovered halfway through a stage. Names that clash with an *installed third-party* package are not checked, because what is installed varies by environment: if your package is called `yaml` or `requests`, rename it or seed the single module instead. - **`seed.implementation` must name the folder itself**, not a path that resolves to `.` or `..`. Those are refused too, because the install would land outside the task's own directory. The full matrix of import shapes, including the ones that do not work, is pinned in `tests/test_multi_module_seed.py`. ## Where things live Task specs and shared defaults are read from `inputs_private/` in your project directory if present, otherwise from the copies bundled inside the package, so a fresh install runs immediately. Resolution is **per file**: dropping one task spec into `inputs_private/config/tasks/` overrides exactly that task and leaves everything else in place. Nothing is ever written back into the package. | Path | Contents | |---|---| | `config/tasks/<task_id>.yaml` | One task, as above. | | `data/<task_id>/` | That task's fixture data. | | `config/settings.yaml` | Retry budget, stage order, patch size limit. A private copy is overlaid section by section, so state only what you change. | | `agent_defs/*.md` | The prompts. `code_agent.md` and `test_agent.md` are the two system prompts; the rest are per-mode fragments. Not per-task: editing these changes every task's behaviour. | ## Proposing fixture rows A criterion no input row can trigger produces a test that passes whatever the code does. Across this project's own measurements roughly two thirds of deliberately planted faults were missed by every suite for that reason: the bar was unmeasurable rather than wrong. ```bash qikly --propose-fixtures --tasks <MY_TASKS> ``` A separate agent reads your criteria and your fixture files and says, for each criterion, either `covered` or here is the smallest row that would reach it. The answer goes to `outputs/reports/fixture_proposals/`, laid out with each row printed under the criterion it exists to reach so you judge the two together. **It never edits a fixture.** To accept a row, paste it into the named file and append ` # proposed`. To reject one, do nothing. Two reasons for the gate, neither about the model being untrustworthy. A row is only right or wrong relative to its criterion, so it is harder to review than a sentence. And a fixture set that grows in whatever direction a model finds interesting stops resembling the data you actually process, at which point every rate measured on it describes a world that does not exist. The report is capped at eight proposals per round and prints what share of your rows a machine has written, so that drift is visible in aggregate rather than one plausible row at a time. You can of course add rows by hand at any time, and always could. This exists because noticing *which* criteria have no data behind them is the tedious part. The refinement loop does this for you on what it adds. When `refine_acceptance_criteria` finishes with new criteria, it asks for rows the same way, lists the new criteria first in the report, and logs how many have no data that reaches them, so a sharper bar does not arrive partly unmeasurable. It still applies nothing. -
TROUBLESHOOTING.md 20.5 KB
# When a run does not converge **If a run has not started yet, skip to [Before a run: where did my files go?](#before-a-run-where-did-my-files-go) at the end.** Everything above that section assumes a run has already failed, and the commonest reports this project receives are not about runs at all. A stall is a normal outcome, not a broken tool. The run exits non-zero, names the tests that blocked it, keeps the whole record, and ships nothing. Across every measurement this project has taken, roughly four runs in ten stop this way, and no run has ever reported success on code its own tests rejected. So the question is never "why is it broken". It is which of a short list of things is happening, and the list is short. --- ## Triage Match what you saw to where to look. The rows are in the order to work through them: the one change that moves convergence most, then the checks that settle what happened, then fixes to the task file, and only then more attempts. | What you saw | Section | What to do | |---|---|---| | Poor results on a provider you just set up, or on the default model | [1. Try a stronger model](#1-try-a-stronger-model) | Set `LLM_MODEL` to a mid-tier or larger model and run again | | Any stall, before changing the task file | [2. Read what actually blocked it](#2-read-what-actually-blocked-it) | Change nothing yet: this step decides what to change. In the run's `_report.html` timeline, byte-identical patches go to [5](#5-the-same-patch-appearing-over-and-over), an import or syntax error to [3](#3-a-collection-error-means-nothing-ran), steady progress to [10](#10-give-it-more-attempts), and different patches that never fix the same test to [1](#1-try-a-stronger-model) | | `0 passed, 0 failed, 1 error` | [3. A collection error means nothing ran](#3-a-collection-error-means-nothing-ran) | Make `interface.module` and the declared signatures match what the tests import | | Everything suddenly worse than last week | [4. Check nothing is set that you have forgotten](#4-check-nothing-is-set-that-you-have-forgotten) | Run `qikly --trends --by week` and look for a setting that changed, such as `criteria_per_batch` | | The same test failing every iteration, no progress | [5. The same patch appearing over and over](#5-the-same-patch-appearing-over-and-over) | Move the restriction into `requirements`, or widen the criterion to match reality | | Stopped with a message naming two tests, each fix for one breaking the other | [6. Two generated tests disagree](#6-two-generated-tests-disagree) | Compare the two tests with the acceptance criteria. If one contradicts a criterion, run again without `--resume` so the suites are written and checked again, and leave a correct spec alone | | A stage spends its whole budget and never gets closer | [7. Check the criteria and requirements do not contradict each other](#7-check-the-criteria-and-requirements-do-not-contradict-each-other) | Run `qikly --check-criteria --tasks <MY_TASKS>` and correct whichever statement is wrong | | Tests check arbitrary values rather than the boundary | [8. Check the criteria name values, not adjectives](#8-check-the-criteria-name-values-not-adjectives) | Run `qikly --validate` and rewrite each flagged criterion as a value: "100 is accepted and 101 is rejected" | | A test passes whatever the code does | [9. Check your fixtures can reach every criterion](#9-check-your-fixtures-can-reach-every-criterion) | Run `propose_fixtures` and add the input rows it suggests | | Steady progress, then the budget ran out | [10. Give it more attempts](#10-give-it-more-attempts) | Raise `orchestrator.max_retries_per_stage` (default 10), only when the report shows progress | | **Nothing has run yet, and files you were told about are missing** | [Before a run](#before-a-run-where-did-my-files-go) | Read the first line of the command's output: it names the directory it worked in, and says when the project root is somewhere else | | **You are in a `demo/throwaway_<timestamp>/` folder** | [Before a run](#before-a-run-where-did-my-files-go) | That is a throwaway copy. Start your own project somewhere else | | `--score-code` says the suite does not pass | [Before a run](#before-a-run-where-did-my-files-go) | Usually pytest collected no tests at the path given to `--score-tests` | | **On macOS, no patch ever applies, on any task** | [12. On macOS, no patch ever applies](#12-on-macos-no-patch-ever-applies) | `brew install gpatch`. The system `patch` is BSD and rejects the options qikly sends | | Integration and system pass, unit does not | [11. Expect the unit stage to be where it fails](#11-expect-the-unit-stage-to-be-where-it-fails) | Expected. Accept it, or leave the unit stage out with `orchestrator.test_order` | --- ## 1. Try a stronger model This moves convergence more than anything else here, and it is one environment variable. ```bash export LLM_MODEL=gpt-4o # or a larger model on your provider qikly --tasks <MY_TASKS> # e.g. --tasks CALC_TAX,MERGE_SALES ``` Every convergence figure in this project was measured on `gemini-3.5-flash-lite`, a deliberately small and cheap model chosen so that sweeps of hundreds of runs were affordable. Treat those figures as a floor. An entry-level model on any provider may stall on a task a mid-tier one clears comfortably. If you are evaluating qikly, evaluate it on a model you would actually ship behind. **What a bigger model buys, and what it costs.** On one CALC_TAX run, `claude-sonnet-5` generated 44 tests against `gemini-3.5-flash-lite`'s 24, and its suite rejected code that Gemini's suite accepted, on five tests, while Gemini's suite accepted its code entirely. A stricter bar, in other words. It also took 403 seconds against 31, and cost \$0.81 against \$0.005. That trade is worth making deliberately rather than by default. A reasoning model produces thinking tokens you are billed for and wait on, which is why the cheapest model is the default here and why every published figure was measured on it: a 400-run sweep costs about \$3 on the default and roughly \$320 on a reasoning model. **Use the cheap model to measure and the expensive one to work.** If you need a convergence rate, take it on the default. If you need the strictest bar for one important specification, pay for it once. One run of each is an anecdote, not a comparison; the figures above are a single run per model. ### The default model is not the same size on every provider <a id="provider-defaults"></a>Set no `LLM_MODEL` and each provider gets its own default, and they are not the same class of model. This is the first thing to check when a run takes far longer on one provider than another: | Provider | Default model | What that means for a run | |---|---|---| | Gemini | `gemini-3.5-flash-lite` | Small, cheap, no reasoning step. Every published figure here was measured on it. A demo task runs in well under a minute | | OpenAI | `gpt-4o` | Mid-tier. Slower and dearer than the Gemini default | | Anthropic | `claude-sonnet-5` | A reasoning model. qikly sends no thinking configuration, and on this model that means adaptive thinking runs by default, so every call thinks before it answers | The Anthropic default is the one that surprises people. The same demo task that finishes in under a minute on the Gemini default has taken around sixteen minutes on it, for the same seven or so model calls. Nothing is wrong when that happens: you are watching a reasoning model think, and it produces a stricter suite for it. **For a like-for-like comparison with the Gemini default, name the model:** ```bash export LLM_MODEL=claude-haiku-4-5 # Windows PowerShell: $env:LLM_MODEL = "claude-haiku-4-5" ``` Haiku is the closest Anthropic analogue to a flash-lite class model, and qikly sends no thinking budget, which that model needs before it will think at all. So a run on it spends no time or tokens on reasoning. Keep `claude-sonnet-5` when you want the stricter bar, and expect the run to take minutes rather than seconds. What qikly does not yet expose is the middle setting: the API takes a reasoning effort level, and a way to ask for less of it without changing model would make this a dial rather than a switch. **If it looks stuck**, it probably is not. A single call can legitimately run for minutes on a reasoning model. After ten seconds of silence a run starts saying so, one line every fifteen: `[patch] still waiting on the model, 45s`. Those lines are the difference between slow and stalled, and `QIKLY_NO_PROGRESS=1` turns them off. One call is abandoned after `QIKLY_REQUEST_TIMEOUT` seconds, 300 by default, which was chosen when the slowest observed call was well under a minute; on a reasoning model consider raising it, or a slow-but-working call is thrown away and retried from scratch. ## 2. Read what actually blocked it Every run writes a timeline: ``` outputs/reports/iterations/<task>_<timestamp>_report.html ``` Open it in a browser. It shows every iteration, the FIX reasoning and the PATCH diff for each failure, and, most usefully, **which patches applied cleanly and changed nothing.** A run full of those is not a run that needs more attempts. It is [3](#3-a-collection-error-means-nothing-ran) or [5](#5-the-same-patch-appearing-over-and-over). ## 3. A collection error means nothing ran ``` [MY_TASK] [stage 1/3] [iteration 3] 0 passed, 0 failed, 1 error, 0 skipped ``` No test failed, because no test ran. The module could not be imported. This is a different problem from a wrong answer, and until it is fixed nothing else can be assessed. The FIX prompt is told this explicitly, and the report carries the underlying `ImportError` or `SyntaxError`. The usual cause is a mismatch between what `interface` declares and what the agent wrote, so check that `interface.module` and the declared function signatures are exactly what the tests should be importing. ## 4. Check nothing is set that you have forgotten ```bash qikly --trends --by week ``` Convergence per task over time, from the run summaries already on disk. Every period names the model and settings behind it, and a period where those changed is marked. This exists because of a specific, expensive mistake. `criteria_per_batch` in a settings file controls how many acceptance criteria a single test-generation call is shown. At `0` one call sees the whole bar. At `4` the bar is split into batches and each gets its own call, so a long bar produces roughly three times as many tests, and every run has three times as much to satisfy. Left set from an earlier experiment, it made convergence appear to collapse across nine tasks at once. Half a day went into diffing prompts, specs and provider parameters before anyone looked at the override. **A rate belongs to a tool, a model and a configuration together.** A rate that moved when the configuration moved is not a finding. ## 5. The same patch appearing over and over Identical diffs, not merely a repeated failure, is a specific signature: the model is fighting something it correctly knows about the world. Restrict a real-world field to an artificial subset, say three valid street suffixes, and the model will keep widening the restriction back. Not out of disobedience. Every piece of its training agrees that "Boulevard" is a street suffix, and your criterion is the outlier. At temperature zero this does not converge slowly; it does not converge at all. **Fix:** widen the criterion to match reality, or move the restriction into `requirements`, where the coding agent can read it and treat it as a given rather than as an error to correct. To confirm it, compare successive diffs under `outputs/logs/patches/<task>/<timestamp>/`. Byte-identical patches mean this. Different patches that never resolve the same test mean something else: a bug that needs more than the failure text to fix, which is [1](#1-try-a-stronger-model). ## 6. Two generated tests disagree ``` Stopped on stage 'system' after 4 attempts: the last three patches alternated between the same two diffs. [...] The failing tests alternate between test_run_headway_two_second_rule_warning and test_integration_pipeline_flow in the 'system' and 'integration' suites: each fix for one breaks the other [...] ``` Each patch makes one test pass and the other fail, because the two tests expect different results for the same input. No code can pass both, so more attempts cannot help, and the fault is in the tests rather than the specification. In the run behind this section, an integration test warned at exactly 2.00 seconds of headway and a system test did not, against a criterion saying exactly 2.00 seconds raises no warning. The same mistake can also be made identically in both suites. Then they agree with each other and still contradict the criterion, the run fails without this message, and the place to look is the same: each test's comparison at every limit the criteria state. An opt-in check, `check_suites: true` under `test_generation` in settings, looks for these before any code is written and rewrites a suite once. Measured on one task it found every wrong suite but did not raise convergence, and a wrong finding once led a correct suite to be rewritten wrong, so it is off by default. **Fix:** compare the two named tests with the acceptance criteria. If one contradicts a criterion, run again without `--resume`, so the suites are written and checked again. Do not change a specification that is already right: this is the one stall where the spec is not the problem. To check suites already on disk without a run, one model call per task: ```bash python -m qikly.orchestrator.tuning.check_suites --tasks <MY_TASKS> ``` ## 7. Check the criteria and requirements do not contradict each other ```bash qikly --check-criteria --tasks <MY_TASKS> ``` One model call per task, and it changes nothing. A criterion that no implementation could satisfy alongside the requirements produces a stage that spends its entire budget discovering that the slow way. It exits non-zero on a contradiction, so a pipeline can gate on it. ## 8. Check the criteria name values, not adjectives ```bash qikly --validate ``` Free, offline, and it flags criteria written as adjectives. > "Reject large amounts" invites a test at some arbitrary large number. > "100 is accepted and 101 is rejected" forces a test at the boundary. This is the highest-leverage habit in writing a bar. A suite that never tests a boundary cannot catch an error at that boundary, no matter how many other cases it covers, and off-by-one at a boundary is among the oldest defect classes in software. `--validate` also catches the quiet structural mistakes: `acceptance_criteria` written as one long string instead of a list, a `task_id` that disagrees with its filename, and fixture paths that do not resolve. Each of those otherwise surfaces twenty minutes and several dollars into a run. ## 9. Check your fixtures can reach every criterion ```bash qikly --propose-fixtures --tasks <MY_TASKS> ``` A criterion that no input row can trigger produces a test that passes whatever the code does. The bar is not lower; part of it is absent. Eight of the ten tasks bundled with qikly had at least one before this was run on them, from one in `CALC_CALENDAR` to seven of thirteen in `MERGE_CONTACTS`. Assume yours do too. It writes proposals to a file and never edits your data. ## 10. Give it more attempts `orchestrator.max_retries_per_stage` in `config/settings.yaml`, default 10. Worth raising when the report shows steady progress that simply ran out of room. Not worth raising when it shows the same patch repeating: that run will fail identically with a hundred attempts, and cost ten times as much doing it. ## 11. Expect the unit stage to be where it fails About twenty points of the gap between "passes integration and system" and "passes everything" is the unit stage, consistently, across every sweep this project has run. The reason is structural rather than mysterious: the unit suite is the largest, runs last, and is the only one written with sight of the implementation. [design_2_performance.md](https://github.com/gal-a/qikly/blob/main/docs/design_2_performance.md#nearly-the-whole-gap-between-those-two-numbers-is-the-unit-stage) has the full explanation. If behavioural verification is what you need, `orchestrator.test_order` in settings can leave it out. ## 12. On macOS, no patch ever applies Every generated diff fails, on every task, from the first iteration, for a reason that reads like the model's fault and is not. The system `patch` on macOS is BSD, and it rejects the options qikly sends. Install GNU patch and the same run goes through: ```bash brew install gpatch ``` Nothing else changes. If patches apply on one machine and fail on all of them on another, this is the first thing to check. --- --- ## One run is an artifact, not a rate The same task with the same seed converges on some runs and not others. Before concluding anything about a task, a model or a setting: ```bash python -m qikly.orchestrator.run_all --tasks <MY_TASKS> --repeat 10 \ --skip-eval --skip-refine ``` That writes an aggregate with a confidence interval instead of a pass count. Ten runs is usually enough to tell a real difference from noise, and it is worth knowing that at n=10 the intervals are wide: this project has measured the same unchanged task at 72% and then 90% on consecutive sweeps. If a change looks like an improvement after one run, it is not yet evidence of anything. --- ## Before a run: where did my files go? None of the sections above apply if nothing has run yet, and this is the question that arrives most often. ### It said it created files and they are not there They almost certainly are, somewhere you did not look. qikly moves to a resolved **project root** when it starts, which can be a different directory from the one you are standing in, and `--init`, `--example` and `--scaffold` write relative to one of those two. Every command now opens with a line that settles it: ``` qikly 0.5.4 2026-09-27 11:27:18 run in C:\Users\you\my-project project root is elsewhere: C:\some\other\place (set QIKLY_PROJECT_ROOT to choose it, or cd there) ``` The second and third lines appear only when the two differ. If you see them, that is your answer. If you are on an older version that does not print them, upgrade, or search for one of the files by name: ```powershell Get-ChildItem $HOME -Recurse -Filter "MY_METRICS*" -ErrorAction SilentlyContinue | Select-Object FullName ``` To pin the project explicitly rather than let it be inferred, set `QIKLY_PROJECT_ROOT` to the directory you mean. ### You are standing in the demo's folder `qikly --demo` runs in a throwaway `demo/throwaway_<timestamp>/` directory so it cannot touch anything of yours, which also means **everything in it goes when you delete the folder, and nothing in it is yours**. Somebody who has just watched the demo work is standing in something that looks exactly like a working project, and the obvious next move is to start theirs there. qikly now refuses, names the folder, and says where to go instead. If you deliberately kept that directory and want to work in it, delete the `.qikly-demo` marker inside it and the refusal stops. ### The version you are running is not the version you installed Two things can disagree. `qikly --version` reports what actually runs; `pip show qikly` reports metadata that an interrupted or repeated upgrade can leave stale. Trust `--version`. If they disagree, clean it: ```powershell pip uninstall qikly -y pip install qikly ``` `qikly --version` also prints the package directory and the interpreter, which is what to check when a flag the documentation describes does not exist. ### `--score-code` says your suite does not pass It refuses to score a suite that does not pass your untouched code, because every planted fault would then fail for the reason the original does and the number would mean nothing. Two causes, likeliest first: - **pytest collected nothing.** The path given to `--score-tests` holds no tests, or none that match its discovery rules. Run pytest on that path yourself and read what it says. - **Your suite genuinely fails.** Fix that first, then score it. There is no reachability warning in this mode, unlike `--score-suite`: there is no task file, so there are no criteria to be unreachable. A fault that survives may still sit on a line no test executes at all, which is a gap in what your tests reach rather than in what they assert. --- ## Still stuck The run kept everything. `outputs/logs/transactions_<task>_<timestamp>.jsonl` is an append-only record of every test run, every FIX, every PATCH and every apply outcome, and it is the source of truth that the reports are rendered from. Issues and results are welcome: [github.com/gal-a/qikly/issues](https://github.com/gal-a/qikly/issues).
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.