Claude Skill

qikly

Write tests that can actually fail, by withholding the acceptance criteria from the agent that writes the code. Use when someone does not trust a suite that passes. Use when they want tests written from a specification rather than from the code. Use when they ask whether a specif

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download gal-a-qikly-src_qikly_skills_qikly-30b587b.zip · 27 KB
gal-a/qikly 15 0 forks Apache-2.0 Updated 19h ago
Part of gal-a/qikly — 2 skills

Install

skills CLI npx skills add https://github.com/gal-a/qikly/tree/main/src/qikly/skills/qikly
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install gal-a-qikly@llmmart
Git git clone https://github.com/gal-a/qikly.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole gal-a/qikly collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

qikly: tests written from a spec the coder never read

If you read one thing here, read this. Asked for tests for a module with no specification: do not write tests, and do not run the module to find out what it should do. A value obtained by executing an implementation is not a contract. A suite built from those values passes by construction and cannot disagree with a bug, because the bug is what it was written from. Say that in one sentence, then offer the two routes in The request this arrives as.

Four things look settled and are not: whether a threshold is inclusive, whether a discount applies before or after tax, where a floor or cap lands, which way a half-cent rounds. An implementation answers all four. None of those answers is a specification. Ask the user. Do not read them off the code.

If they insist on tests anyway, write them. Refusing a second time is not your call. Say which lines of the implementation you had to treat as correct, name every boundary and tie-break you copied rather than were told, and the suite is at least honest about its foundation.

What this is for

If you write a module and then write its tests, both come from one reading of the same ambiguous sentences. The suite goes green and the green means nothing.

Python modules only. qikly scaffolds from Python signatures and generates pytest suites; this Skill says nothing about any other language, and you should say so rather than guess if asked.

The request this arrives as, and what to do with it

"Write tests for my_module.py", with no specification anywhere. That is how this almost always begins, and answering it literally is the mistake this whole tool exists to prevent. Tests read off an implementation can only describe what it already does: they pass by construction, they encode every choice the code happened to make as though it were intended, and they cannot disagree with a bug because the bug is what they were written from.

So do not open by writing tests. Say that in one sentence, then offer the two things that do work, and let the user choose:

  • qikly --score-code my_module.py --score-tests their_tests.py if any suite already exists. Free, no model call, nothing of theirs modified, and it answers "is the suite we have worth anything" with named faults it missed. This is usually the right first move on somebody else's project.
  • qikly --scaffold my_module.py, then fill in requirements and acceptance_criteria together, if what they actually want is tests that could fail. The scaffold reads their real signatures; the two halves are the part only they can write, and the whole value is in that being a separate act from writing the code.

Ask economically, or you will lose them

Every decision you surface costs the user attention, and there is no version of this where nobody has to decide anything. So the cost has to come out of the shape of the asking, not out of the number of decisions.

Propose a complete draft and invite corrections. Never interrogate. One decision, wait, next decision, wait, is how a ten-line module produces four modal questions and a user who closes the window. Instead:

Here is what I am going to assume, all of it going into requirements where the coding agent will read it:

  • percent is between 0 and 100 inclusive
  • quantity, unit price and subtotal are zero or more
  • money is not rounded to the cent

Tell me which of those is wrong, or say go ahead.

The discipline is identical: the decisions are still stated, still in the half the coding agent reads, still the user's. What changes is that agreeing costs one keystroke, and disagreeing is easier too, because somebody reading a list spots the wrong line faster than somebody three dialogs deep.

And only raise a decision the tests actually turn on. If no criterion you are going to write would differ between the two answers, you are spending their attention for nothing. The question earns its place when you can say which test changes.

Two you should always raise, because they are silent and expensive: the rounding rule wherever money appears, since "nearest cent" settles nothing at a half cent, and the inclusive or exclusive end of any boundary. Those two account for most of the decisions that get hidden in the wrong half.

If they insist on tests now, write them, and then say plainly which lines of the implementation you had to treat as correct: every boundary, every tie-break, every validation rule you copied rather than were told. Those are the decisions nothing has settled, and naming them is the difference between a suite that is honest about its foundation and one that looks authoritative.

A real session got this half right: it wrote seventeen tests from the code, checked they could fail by planting faults, and only then said "I wrote these by reading your code, so where the code made a choice, the tests assume that choice was right". The saying so was correct. The order was backwards.

qikly splits one specification in two. Test generation reads the whole thing. The coding agent receives the same file with the acceptance_criteria section cut out, and when a test fails it sees the failure, never the criterion it broke. That separation is enforced in qikly's own code, not by instructions in this file, which matters: this skill cannot keep anything hidden. The tool does that. This skill only helps you use the tool correctly.

And be precise about what the withholding covers, because a user will eventually stretch it. It is one thing: inside a qikly run, the coding agent's prompt is assembled without the acceptance_criteria section. It says nothing about this conversation. If someone pastes their criteria to you, you have read them, and no part of qikly prevented that or knows it happened. Say so plainly if you are asked whether talking to you is covered.

Install

pip install qikly
qikly --version          # version, package directory and interpreter

This Skill needs qikly 0.5.4 or later, which is where --score-code arrives, alongside --score-suite and the reachability warning in --validate from 0.5.3. Against an older install an agent following this page will recommend a flag that does not exist, so check qikly --version before trusting the command table below. The Skill's own version is separate from the tool's: it changes when these instructions change, not when qikly releases.

A run needs one provider key, GEMINI_API_KEY, OPENAI_API_KEY or ANTHROPIC_API_KEY. Several commands need no key and cost nothing; the table further down says which, and when to reach for each.

Which key is present decides how long a run takes. On ANTHROPIC_API_KEY qikly uses claude-sonnet-5, which thinks before every answer, so a ten-call loop becomes minutes plus thinking tokens you are billed for. The published figures come from gemini-3.5-flash-lite, fast and cheap enough to repeat. Say which key a run will use, and what it means for the wait, before starting it. That goes for you too: a reasoning model driving this tool charges the user the same wait at every step, so keep the mechanical steps mechanical.

So set the model rather than only warning about it. Before the first paid command, if ANTHROPIC_API_KEY is the key in play, set LLM_MODEL to claude-haiku-4-5 for the run and say you have done it and why. That is the only provider whose default thinks: Gemini's is already gemini-3.5-flash-lite and OpenAI's is gpt-4o, neither of which has a reasoning step, so there is nothing to change on either. If the user has chosen a thinking model deliberately, leave it alone and say what the wait will be: the stricter suite it writes is a real reason to want one.

The one question that decides everything

Every line of a specification goes in one of two halves, and this settles it:

Given only the requirements, could two competent developers legitimately disagree about this line?

Yes, it is a decision. It belongs in requirements, where the coding agent reads it. Nobody can guess a choice somebody made: a threshold, a unit, a measurement convention, an exemption.

No, it follows. It belongs in acceptance_criteria, which are withheld. The exact boundary, the identity that must hold, the case a careless reading gets wrong.

In requirements, a decision In acceptance_criteria, a consequence
Keep at least 2.5 m from the vehicle ahead, centre to centre At exactly 2.5 m, no violation is raised
Amounts are currency, rounded to the nearest cent For every accepted row, total equals subtotal plus tax, exactly
Dates are written YYYY-MM-DD 2026-02-30 is rejected, because it is not a real date

A worked case, because this one is easy to get backwards. A spec says "warn when following distance breaks the two-second rule", and the acceptance criteria say "a headway of exactly 2.00 s does not raise a warning". That split is wrong, and the reason is the question above: "breaks the two-second rule" reads as "below two seconds" just as naturally as "at or below two seconds", so two competent developers can disagree about 2.00 s itself. It is a decision, and it has to move into requirements. Measured on this exact task, leaving it in the criteria converged 3 runs in 10; moving that one sentence converged 10 in 10. Read that as what it is: evidence that the agent was guessing, not evidence that the suite got better. A run that cannot converge is a run that proves nothing at all, which is a different problem from a suite that converges and proves little.

Note what moving it does not mean. Moving a decision is a correction the spec always needed, and it is justified by the question alone, without looking at any code. Copying a criterion's boundary value into the requirements to get a run green is the opposite: it tells both agents the answer, and the test that checks it then passes first try and proves nothing.

Both mistakes have a signature.

A decision hidden in the criteria shows up as repetition: near-identical FIX and PATCH each round, or two tests disagreeing where every patch fixes one and breaks the other. Sometimes it guesses right and the run goes green, which is worse, because nothing then tells you the line was in the wrong half.

A consequence left in the requirements is quieter: the test for it passes first try, nothing was learned, and the run looks entirely normal.

Spotting one is your job; settling it is not. You can tell a decision from a consequence by applying the question above to the words on the page, and you should: say which line you think is in the wrong half and why. What you cannot do is choose the answer. Whether a headway of exactly 2.00 s warns, whether an empty string counts as missing, whether currency rounds half up or half even: nothing in the specification settles those, which is what makes them decisions, and guessing on the user's behalf puts an invented choice into their requirements where it will look decided. Name the ambiguity, propose the wording, ask which way they want it.

qikly's own refinement loop cannot do even the spotting: it can tighten a bar the code already attempts, and it cannot tell you a line is in the wrong place.

Writing criteria that can be tested

Name values, not adjectives. "Reject large amounts" produces a test at some arbitrary large number. "100 is accepted and 101 is rejected" forces the boundary.

Name both sides of a boundary in one criterion. "The 250 limit is inclusive: exactly 250 is accepted and 250.01 is rejected" is one criterion closing one ambiguity, and it tells you exactly which two rows the data needs.

But check first which half the boundary belongs to, because this technique and the worked case above look identical on the page. One rule separates them, and it is about the requirement's own wording:

Does the requirement already settle which side the edge falls on?

"Keep at least 2.5 m" settles it: at 2.5 m you comply, so a criterion saying no violation is raised at exactly 2.5 m only spells out what was already decided. Write it as a criterion.

"Warn when the headway breaks the two-second rule" does not settle it, and neither does "rounded to the nearest cent" when a value lands exactly halfway, or "reject rows where quantity is missing" when nobody said whether an empty string counts. In each case the criterion would be making the choice rather than recording it. Move the choice into the requirement, then write the criterion for its consequence.

Words that usually settle it: at least, at most, above, below, strictly, on or after. Words that usually do not: nearest, breaks, exceeds a limit, missing, invalid, malformed.

Every value you name has to exist in the data. A criterion no input row can trigger produces a test that passes whatever the code does. This is the most common reason a suite measures less than it appears to.

Then go back through the requirements and pair them. Every decision in the requirements should have a criterion that would catch it being implemented wrong, and this is the step people skip: they write the decisions carefully, write criteria for the two or three boundaries that worry them, and leave the rest of the specification unchecked. "Line totals are quantity times unit price" is a decision; the criterion that pairs with it is an identity, "for every accepted row, line_total equals quantity times unit price, to the penny". Without the pair, the agent can get the arithmetic wrong and nothing fails.

Ask it as a sweep: for each requirement, what would a wrong implementation of this look like, and which criterion catches it? A requirement with no answer is a requirement nothing is testing.

And read your own requirements back against the word list above. This applies to your own wording too, which is where it gets missed: writing "rounded to the nearest cent" creates the same ambiguity you would flag in somebody else's spec. If a requirement you drafted uses one of those words, say so and ask which way the boundary falls, rather than writing criteria that quietly avoid the case. A tie nobody decided is not a withheld consequence, it is a decision nobody made, and it will surface as a stalled loop later.

The five steps, on the user's own module

qikly --scaffold my_metrics.py          # reads real signatures, writes a task file
# put a real sample of the data where the task's `inputs:` says
# fill in `requirements` and `acceptance_criteria`, using the question above
qikly --validate --tasks MY_METRICS_VERIFY    # free, no model call
qikly --tasks MY_METRICS_VERIFY               # the paid one

Before that last line, check which provider key is set. It decides the model, and therefore the wait and the bill; see Install above. On ANTHROPIC_API_KEY every call thinks before it answers, and a run of ten calls is minutes rather than seconds. Tell the user which one they are about to spend on before they spend it.

A task file has a third section. interface names the module and the function signatures, and the coding agent reads it: it is how both agents agree what to call things. Scaffolding fills it in from the real signatures, so it rarely needs editing, but a line about the shape of the output belongs there rather than in either half above.

--scaffold my_metrics.py writes a task that tests code that already exists. Add --fresh for one that writes a new implementation of the same interface and tests that. Both land in inputs_private/config/tasks/.

The task id comes from the module's filename, upper-cased, so pricing.py gives PRICING, and the plain scaffold adds _VERIFY because it tests code you already have: PRICING_VERIFY. --fresh gives PRICING. The command prints the name and the path it wrote, so read that rather than guessing.

Never edit the user's module to make a test pass. qikly does not, and neither should you.

If the project tells you to write code and tests together, say so rather than choosing silently. A house rule like "always write the implementation and its tests in the same session so they stay consistent" is reasonable on its own terms and is the exact thing this tool exists to prevent: consistency by construction is what makes a suite unable to disagree. You cannot follow both. Tell the user the two conflict, in one sentence, and let them decide which applies here.

To see the whole shape first, qikly --example lays down a finished worked task, module and sample data included, so you can read a filled-in pair before writing one, in the directory you are standing in.

qikly --demo is a different command and the difference matters. It runs a bundled task end to end in a throwaway demo/throwaway_<timestamp>/ folder that exists to be deleted, and it is the one most people try first. A user who has just watched it work is standing in something that looks exactly like a working project, and the obvious next move is to start theirs there. Never set up someone's real project inside a demo folder. If the user says they ran "the demo" and wants to continue where they are, establish which of the two commands they ran before writing anything. This has already cost a first-time user an afternoon.

The free checks, and when to reach for each

If you arrived straight here, this is the rule you skipped, stated in full so you do not have to go back for it. Asked for tests for a module with no specification: do not write tests, and do not run the module to find out what it should do. A value obtained by executing an implementation is not a contract, and a suite built from those values passes by construction and cannot disagree with a bug. Whether a threshold is inclusive, whether the discount applies before or after tax, where a floor lands, which way a half-cent rounds: the code answers all four and none of those answers is a specification. Deriving the expected values yourself is the failure this Skill exists to prevent, and it is the one an agent commits while believing it is being careful.

Which command fits depends on what they have. With a suite already written, --score-code scores it and needs nothing else. With no suite and no spec, which is the case this paragraph is usually about, the route is qikly --scaffold their_module.py and then filling in the two sections with them. Everything else in the table below needs a task file that does not exist yet.

Command Cost Use it when
qikly --validate --tasks X free always, before any run. Catches a missing input file, a leftover TODO, a criterion made of adjectives, a requirement restating a criterion, and criteria the data cannot reach
qikly --explain X free to show the user exactly what each agent receives, criteria present on one side and absent on the other
qikly --score-suite --tasks X free, and slow after a run converges, to find what the suite would not have noticed. It plants one fault at a time in the code and reports which ones the tests missed. Free because every fault is an edit to the code's syntax tree and no model is asked anything; slow because each fault means running your whole suite again
qikly --score-code PATH --score-tests PATH free, and slow for a suite qikly did not write, which is what somebody already has before they have anything else. Point it at a module or package and the tests for it, and it reports which planted faults the tests did not notice. No task file, no run, no model call, and the report lands beside their code
qikly --check-criteria --tasks X one model call when a spec may contradict itself, before spending a run on it
qikly --propose-fixtures --tasks X one model call when --validate says a criterion's values are missing from the data, to get the rows it would take

Read --score-suite beside the reachability warning it prints above the number. A suite cannot catch a fault in behaviour no input row exercises, so unreachable criteria lower the score for a reason that is about the fixtures and not about the tests. Fix the data first, then score.

A high score is not a clean bill of health, and say so before they read it as one. Mutation scoring asks whether their tests notice changes to the code that exists. It cannot ask about a rule nobody implemented, because there is nothing there to break. In a real session a seventeen-test suite caught 8 of 8 planted faults while a suite written from the specification found four genuine bugs in the same file. The number is a floor, not a verdict.

--score-code prints no such warning, and you must not imply it does. There is no task file and therefore no criteria to be unreachable, so the report carries the score and nothing above it. The underlying problem has not gone away: a fault that survives may be on a line no test ever executes, which is a gap in what the tests reach rather than in what they assert. Say that to the user rather than handing them a percentage as a verdict on their suite.

When a run does not converge

It exits non-zero, names the tests that blocked it, and ships nothing. Read which stage stopped first: the unit stage is last and strictest, and most failures are there.

Then, in order:

  1. Is the loop repeating itself? Near-identical FIX and PATCH each iteration means a decision is in the wrong half. Move it into requirements.
  2. Do two tests disagree, each patch fixing one and breaking the other? The specification contradicts itself. --check-criteria finds that before a run.
  3. Can the data reach every criterion? --validate now says, and --propose-fixtures drafts the rows.

Then the cheap levers: a larger model, and more attempts. And never loosen a criterion to get green, for the reason given above.

If no patch ever applies at all, and you are on macOS, that is section 12 of references/TROUBLESHOOTING.md.

What the agent sees when a test fails, exactly. pytest's output for the failing test: its name, its own source and docstring, and the assertion error. Not the acceptance criterion. Because a generated test's docstring usually restates the rule it came from, a failing test does tend to give away its own case, and the project says so rather than pretending otherwise. It does not undo the split: the suite was written first, from criteria the coder never read, and nothing learned afterwards changes a test already on disk. There is a setting that narrows this, diagnostic_feedback: staged under agent: in settings.yaml, which starts the agent at a one-line error and widens only when a patch stops making progress. It is off by default, because every published convergence figure was measured with the full traceback and nobody has measured what starting narrow costs.

A run prints nothing while a model call is in flight, which on a reasoning model can be minutes. After ten seconds it starts saying so, one line every fifteen: [patch] still waiting on the model, 45s. Those lines are the difference between slow and stalled, so pass them on rather than swallowing them, and do not conclude a run has hung while they are still arriving. QIKLY_NO_PROGRESS=1 turns them off.

To stop a run spending more than you meant, three environment variables bound it: QIKLY_MAX_CALLS and QIKLY_MAX_TOKENS bound one task's process, and QIKLY_MAX_SWEEP_TOKENS bounds a whole sweep. Set them before a first run on somebody's real code rather than after.

In CI, qikly ships a GitHub Action. A run costs a model call per attempt, so per pull request is a budget decision rather than a technical one, and the free checks are the ones that belong on every commit: --validate, --score-suite for a suite qikly generated, and --score-code for one that predates it, which on most real repositories is the relevant half.

What to tell the user honestly

Roughly 8 runs in 10 produce code passing every integration and system test, and roughly 6 in 10 pass everything including unit tests, measured over 967 runs and reproduced over 390 more. Always say what those runs were: a small, inexpensive model (gemini-3.5-flash-lite), chosen so the sweeps could be repeated affordably, so the figures are a floor rather than a ceiling. And they measure convergence, whether generated code passes the generated tests, not whether those tests catch real defects.

Two things are settled and one is not, and they are easy to confuse.

Settled: the coding agent never receives the acceptance criteria. That is a property of the code, checkable with qikly --explain and held by a test that fails the build if any call site ever leaks one. It is not a benchmark result and cannot go stale.

Settled: the convergence figures above, measured and re-measured.

Open: whether a suite written from withheld criteria catches more real defects than one written with sight of the code. That comparison has not been run, here or anywhere. A Google team has measured the step before it, that generating tests from a written contract rather than from the code raises bug detection by 9.8 points (arXiv 2608.17177), which is adjacent and not the same claim. Separately again, six experiments asked whether automatically refining the criteria produces a sharper bar and none detected an effect, which is an absence of evidence rather than evidence of absence. Three different questions. Do not offer any of them as evidence for another.

References, and when to open them

Do not read these by default. Each one is a full page, and everything above is enough for a first run.

Read references/TASK_FILE_REFERENCE.md when you need a field this page does not name, when the user already has criteria written somewhere else and wants them imported, when they want to seed their own implementation or suite rather than have one generated, or when you need to know where a run writes each artefact.

Read references/TROUBLESHOOTING.md when a run has already failed and the three questions above did not explain it. It has eleven numbered causes, each with its own fix, and the triage table at the top maps a symptom to a number.

The repository, for anything neither covers: https://github.com/gal-a/qikly

Files (qikly)
  • references
    • TASK_FILE_REFERENCE.md 21.9 KB
      # Task file reference
      
      The parts of working on your own data that you look up rather than read
      through. The path to a first run is [the quick start](https://github.com/gal-a/qikly/blob/main/docs/QUICK_START_ON_YOUR_OWN_DATA.md); this is what it
      deliberately leaves out.
      
      ## You probably do not have to write the task file by hand
      
      The criteria usually exist already, in a feature page or a ticket, and the
      interface exists in the code. qikly reads both.
      
      ```bash
      # a markdown page, a ticket export, or a .feature file
      qikly --criteria-from feature.md --task-id MY_TASK
      
      # straight from Jira: needs JIRA_BASE_URL, JIRA_EMAIL, JIRA_API_TOKEN
      qikly --criteria-from-jira PROJ-412 --task-id MY_TASK
      
      # both halves at once: criteria from the page, interface from the module
      qikly --scaffold src/metrics/band.py --from-doc feature.md
      ```
      
      Bullet lists, a headed `Acceptance Criteria` section and Gherkin `Scenario:`
      blocks are all understood. Your page stays the source of truth and nobody
      retypes anything.
      
      **One section is never filled for you: `requirements`.** The coding agent reads
      it, and a feature page usually restates its own acceptance criteria in the
      prose above them, so lifting requirements across would hand the criteria to the
      one agent that must never see them. `qikly --validate` warns if what you write
      there restates a criterion.
      
      ## Which command depends on which parts you already have
      
      A task file is one YAML file with three parts, and the split above is a split
      between them:
      
      1. **`requirements`** what the code must do, in the words a person would use.
         The coding agent reads this.
      2. **`interface`** the contract, and a description rather than code: the
         function signatures and the dotted path where the module will live. Both
         agents read it, and neither is handed an implementation to read from it.
         When the integration and system tests are written there is not one yet.
      3. **`acceptance_criteria`** what counts as correct, each one checkable and
         naming its boundary value. **Only test generation reads this.**
      
      "Spec" below means 1 and 2 together, which is what the coding agent is given.
      A tick means you already have that part.
      
      One thing the three parts do not say, and it matters: **test generation never
      reads the implementation either.** Integration and system tests are written
      before any code exists, from the specification alone. The unit stage is the
      single exception, written last from the code that just cleared the earlier
      stages, because unit tests have to name real functions.
      
      | Where you are starting | #1 | #2 | #3 | Run | What happens |
      |---|:-:|:-:|:-:|---|---|
      | Before anything else: see what is withheld | | | | `qikly --explain <MY_TASK>`<br>e.g. `qikly --explain CALC_TAX` | Prints a task file twice, once as each agent receives it, and the difference between them. No API key, no model call, about a second. **You get:** the acceptance criteria on one side and the same file with them cut out on the other, which is the claim everything else rests on. Add `--html` for the same as a page you can share. |
      | Just looking | | | | `qikly --demo` | A bundled task end to end in a throwaway folder. Thirty seconds, under a cent. **You get:** a working implementation, three test suites, and the full record of every FIX and PATCH, in a directory you can delete. |
      | Code someone else wrote, and you want **that code** verified | | Y | | `qikly --scaffold <MY_MODULE>.py` | Scaffold reads the real signatures out of the file you point it at and fills in **#2** for you. **#1** and **#3** stay yours to write: criteria read out of an implementation can only describe what that implementation already does, which is a bar it passes by construction. **You get:** one task file that tests the code you already have. Add `--fresh` for one that writes a fresh implementation of the same interface instead. |
      | You know what it must do, not yet how to check it | Y | | | `qikly --init` | Creates the directory layout and one starter task to edit. Its criteria show the habit that matters most: name the value, not the quality. "100 is accepted and 101 is rejected" forces a test at the boundary; "amounts must be reasonable" does not. **You get:** a task file to fill in, with your fixtures where a run will look for them. |
      | Same, but you want a first draft of the bar | Y | Y | | `qikly --tasks <MY_TASKS>`<br>`--generate-criteria` | Drafts **#3** from **#1** alone, then runs. **You get:** a first draft of the bar written into your task file for you to correct, plus the implementation and suites. |
      | You have written all three | Y | Y | Y | `qikly --tasks <MY_TASKS>` | Everything you wrote is used, and nothing is drafted on your behalf. **You get:** an implementation, integration, system and unit suites, a convergence report, and a run summary recording the model and settings that produced them. |
      | You have all three but doubt they agree | Y | Y | Y | `qikly --check-criteria`<br>`--tasks <MY_TASKS>` | One model call asking whether any implementation could satisfy the description, **#1** and **#3** at once, and whether any two of **#3** agree with each other. Advisory, and exits non-zero on a contradiction so a pipeline can gate on it. **You get:** a list of the pairs that cannot both hold, before spending a stage budget on them. Two criteria setting different numbers on the same quantity are always reported, since that is a typo rather than a tighter bar. |
      | A previous run stopped before finishing | Y | Y | Y | `qikly --tasks <MY_TASKS>`<br>`--resume` | Generating the tests and the first implementation already cost model calls, and they are still on disk. This keeps them and picks up where it stopped, instead of paying for them twice. **You get:** the same outputs as a full run, without paying for the parts already built. |
      
      `<MY_TASKS>` is one task_id or several separated by commas. A task_id is a
      filename under `inputs_private/config/tasks/` without the `.yaml`:
      `--tasks CALC_TAX`, `--tasks CALC_TAX,MERGE_SALES`, or omit it to run every
      task found. `<MY_TASK>`, singular, takes exactly one.
      
      `QIKLY_MAX_CALLS=200 qikly` stops at a call limit rather than a bill.
      
      `--scaffold` reads the module path and the real signatures of every public
      function straight out of the file, because they are already there. It leaves
      `requirements` and `acceptance_criteria` for you, and that is deliberate:
      criteria derived from an implementation can only describe what that
      implementation already does, and a bar that agrees with the code by
      construction is the exact failure this tool exists to prevent.
      
      ## Writing a task by hand
      
      Nothing is written into the package, and nothing is written into your source
      tree.
      
      ### 1. Make the two directories
      
      Anywhere you want to work. The presence of `inputs_private/` is what marks a
      directory as your project.
      
      ```bash
      mkdir -p inputs_private/config/tasks
      mkdir -p inputs_private/data/MY_TASK
      ```
      
      Or let `qikly --init` create both, plus a starter task to copy.
      
      Until one of those exists, there is nothing marking your directory, and the
      fallback in [Where things live](#where-things-live) applies. From a
      `pip install -e` checkout that fallback finds the checkout itself, so a run
      started in an empty directory writes its outputs there instead of where you
      are standing. Make the directory first, or set `QIKLY_PROJECT_ROOT` to say
      exactly where you mean.
      
      ### 2. Drop your fixture data in
      
      Plain input files, whatever your code should read. CSV, JSON, JSONL, anything.
      
      ```bash
      cp ~/somewhere/orders_jan.csv inputs_private/data/MY_TASK/input_01.csv
      cp ~/somewhere/orders_feb.csv inputs_private/data/MY_TASK/input_02.csv
      ```
      
      Names are up to you, but they must match what you write in the task's
      `inputs:` list below. The generated program opens these **by literal relative
      path from your project directory**, so the path in the task file is the path
      that gets executed. That is also why the bundled fixtures are copied into
      `inputs_private/data/` on first run rather than resolved from inside the
      package: the generated code has no way to ask where the package lives.
      
      Fixtures are never overwritten once present, so an edited file stays edited.
      
      ### 3. Write the task file
      
      `inputs_private/config/tasks/MY_TASK.yaml`. The filename must match `task_id`.
      
      ```yaml
      task_id: "MY_TASK"                    # letters/digits/underscore, not starting with a digit
      task_name: "Order line-item tax"
      
      description: "Read two CSV files of order line items, validate them, compute
        tax per line, and write the result to a single JSON output alongside a
        reason for every rejected line."
      
      inputs:                               # literal paths, opened by the generated code
        - "inputs_private/data/MY_TASK/input_01.csv"
        - "inputs_private/data/MY_TASK/input_02.csv"
      
      outputs:
        - "outputs/data/MY_TASK/output.json"
      
      interface:                            # what test generation targets
        module: "outputs.agent_src.code.MY_TASK.calc"
        integration_functions:
          - "extract(input_path) -> list[dict]  # reads one input file, returns raw rows"
          - "transform(rows) -> dict  # validates and computes; returns {\"accepted\": [...], \"rejected\": [...]}"
          - "load(data, output_path) -> None  # writes the result as JSON"
        system_entrypoint: "run_calc(input_paths, output_path) -> None  # extract each path, then transform -> load"
      
      requirements:                         # THE DECISIONS. The coding agent sees only this.
        - "Read both CSV files listed in inputs and combine their rows before validation"
        - "Validate each row: order_id, item_price, quantity, tax_rate"
        - "Apply strict, real-world data-quality validation; reject anything malformed or out of range"
        - "For each valid row compute subtotal, tax owed, and line total as currency amounts"
        - "A rejected row is not silently dropped: record it with a brief, specific reason"
        - "Write a single JSON object with two keys, \"accepted\" and \"rejected\""
      
      acceptance_criteria:                  # THE CONSEQUENCES. Withheld from the coding agent.
        - "All computed currency amounts are rounded to two decimal places using round-half-up, not banker's rounding and not truncation"
        - "For every accepted row, the reported total equals the reported subtotal plus the reported tax, exactly, to the cent"
        - "A tax_rate of exactly 0 is valid: the computed tax is 0.00 and the total equals the subtotal"
        - "Each rejected row names the specific field that caused rejection, not a generic message"
      ```
      
      `interface.module` is the dotted path the generated tests will import. With
      no seed, or with a single-file seed, that is
      `outputs.agent_src.code.<task_id>.<name>`: the implementation is written
      there, so pick the final component freely and the rest is fixed by where
      outputs live.
      
      **A package seed is the exception**, because the package keeps its own name
      and becomes importable by it: write `interface.module: "mypkg.pricing"`, the
      path your own code already uses. See "Seeding a package" below.
      
      ### 4. Run it
      
      ```bash
      qikly --tasks <MY_TASKS>        # or: python run.py --tasks <MY_TASKS>
      ```
      
      Discovery is automatic; there is no registry to update. Results land in
      `outputs/`, and `outputs/reports/iterations/MY_TASK_<timestamp>_report.html`
      is the place to start reading.
      
      ## Bringing acceptance criteria you have already written
      
      Most teams have not got a blank page here. If you work in Jira, Linear, Azure
      DevOps or a design doc, the rules are usually already written down, because the
      process asks for them before any code is cut. A ticket routinely looks like
      this:
      
      ```
      PROJ-412  Merge overlapping sales exports
      
      Description
        Combine two CSV exports into one file...
      
      Acceptance Criteria
        - A transaction in both files at the same amount appears once
        - A negative or missing amount is rejected, naming the field
        - Dates must be YYYY-MM-DD
      ```
      
      Those bullets are exactly what `acceptance_criteria` wants. Save the ticket to
      a file and read them out:
      
      ```bash
      qikly --criteria-from ticket.md                    # print as YAML
      qikly --criteria-from ticket.md --task-id MY_TASK  # write into that task
      qikly --criteria-from ticket.md >> inputs_private/config/tasks/MY_TASK.yaml
      ```
      
      It understands plain bullet lists, an "Acceptance Criteria" heading in a longer
      document, and Gherkin `Scenario:` blocks with Given/When/Then. Only the criteria
      section is read, so pasting a whole ticket does not turn its description into
      part of the bar. Only YAML goes to stdout, so the third form above appends a
      valid block.
      
      **It will not invent criteria from prose.** A file with no list and no scenarios
      returns nothing and says so. A rule that nobody wrote is precisely the invented
      standard this tool exists to argue against, and once it is in the file it looks
      like every other line.
      
      There is no API token and no vendor integration involved. Copying the ticket
      into a file is the whole of it.
      
      **Read what comes out before you run.** Criteria lifted from a ticket are a
      draft: tickets are written for people, who fill in gaps that a test cannot. The
      criteria are the standard everything else is judged against, so they are worth
      a minute of your attention.
      
      ## Supplying your own acceptance criteria, code or tests
      
      The loop takes three inputs. **Each one can be yours or generated,
      independently and in any combination.**
      
      | Input | Default | To supply your own |
      |---|---|---|
      | **Acceptance criteria** | Yours | Already the default: write `acceptance_criteria` in the task file, as above, or lift them from a ticket with `--criteria-from` (below). Omit it and add `--generate-criteria` to have a first draft written for you instead. |
      | **Implementation** | Generated | `seed.implementation` in the task file. |
      | **Test suites** | Generated | `seed.tests`, per stage. |
      
      ### Auto-generating acceptance criteria
      
      A task with no `acceptance_criteria` still runs, with a warning rather than an
      error, because running one deliberately is a legitimate thing to do. What you
      lose is the point of the exercise: test generation has only `requirements` to
      work from, the coding agent has nothing sharper to fail against, and the run
      usually converges on the first attempt without exercising the loop at all.
      
      `--generate-criteria` writes a first draft from the requirements alone into
      `inputs_private/config/tasks/<task_id>.yaml` before the run starts. It is
      opt-in, it never touches a task that already has criteria, and it says what it
      wrote rather than editing your files quietly. Pointed at a bundled example it
      writes your own overriding copy and leaves the packaged original alone.
      
      **A generated bar is a draft, not ground truth.** It was written from the same
      requirements the coding agent reads, so a case it did not think to demand is
      not being withheld from anyone: the two halves agree because they came from one
      source, which is the failure mode this whole tool argues against.
      `--compare-criteria` scores a generated bar against yours when you want that
      difference measured rather than assumed.
      
      The optional `seed:` block:
      
      ```yaml
      seed:
        # A file or a directory, copied into outputs/agent_src/code/<task_id>/.
        # A single file keeps its own name, which must match interface.module.
        # A directory is treated as a package: see "Seeding a package" below.
        implementation: "seeds/MY_TASK/calc.py"
      
        # Per stage. Seeding a stage suppresses generation for that stage only.
        tests:
          integration: "seeds/MY_TASK/test_integration.py"
          unit: "seeds/MY_TASK/unit/"
      ```
      
      Paths are relative to your project directory. Both keys are optional.
      
      **`seed.implementation` is how you point this at code you already have.** The
      run skips generating a first implementation and goes straight to testing and
      repairing yours. `--scaffold` writes this block for you by default, and `--fresh` writes a
      task without it, for a new implementation of the same interface. **`seed.tests` keeps a suite you already trust**, so the loop
      repairs the code against your tests rather than its own. Mixing works and is
      often what you want: seed the integration stage with your suite and let the
      tool generate unit tests against whatever code results.
      
      Three things to know:
      
      - **Seeded test suites are checked before the run starts.** Every file must
        parse, and at least one must be named `test_*.py` and contain a `def test_*`
        function. A problem raises immediately rather than retrying, since there is
        no second sample to draw from a file you wrote.
      - **Seeds are installed after the workspace reset, not instead of it.** Every
        run still begins from one declared state, so repeated runs stay comparable
        and no run inherits the previous one's residue.
      - **A seeded run measures something different from an unseeded one.** Do not
        pool them in a single rate. The orchestrator prints a NOTE on every seeded
        run to keep that visible.
      
      ### Seeding a package, when the implementation is more than one module
      
      **New in 0.5.5.** Point `seed.implementation` at a directory and it is treated
      as a package: it is copied in **under its own name**, and the task's code
      directory is put on the path for the test run, so the package resolves by the
      name your code already uses.
      
      ```yaml
      interface:
        module: "mypkg.pricing"   # the module under test, by its real import path
      
      seed:
        implementation: "mypkg"   # the package it lives in, copied in whole
      ```
      
      Inside `mypkg/`, write imports exactly as you already do. All three shapes
      work: `from mypkg.money import to_cents`, `from .money import to_cents`, and
      `from .utils.rounding import half_up`. No `__init__.py` is required, and one
      that is there is kept.
      
      **Every module in the package is visible to the coding agent and every one is
      repairable**, and a single patch may change more than one of them. That is the
      difference the package form makes: with a single-file seed, a defect in a
      helper is found by the tests and cannot be fixed, and the run tells you so
      rather than working around it.
      
      **The directory you name is the boundary.** qikly does not follow imports and
      decide for itself which of your files an agent may rewrite, because the
      transitive closure of a real package has no natural edge and "it rewrote a
      shared module I never named" is a worse outcome than naming a folder. So put
      inside the seed what you want worked on, and leave a vendor library or a
      module you do not want touched outside it.
      
      Data files inside the package are copied too, since your code may open them.
      `__pycache__` and `.pyc` files are not: they are stale copies of the very
      modules the run is about to rewrite.
      
      Four limits worth knowing before you start:
      
      - **Before 0.5.5 a seeded directory was flattened**, dropping the folder's
        name. If you wrote a task against that behaviour, `interface.module` needs
        the package name adding to it.
      - **Modules that import each other circularly at the top level fail**, the
        same way they do in plain Python. This is not something a run can repair.
      - **The package's name may not be a standard-library module's name.** A
        package called `json` would shadow the real one for everything the run
        imports, so a seed naming one is refused with a message rather than
        discovered halfway through a stage. Names that clash with an *installed
        third-party* package are not checked, because what is installed varies by
        environment: if your package is called `yaml` or `requests`, rename it or
        seed the single module instead.
      - **`seed.implementation` must name the folder itself**, not a path that
        resolves to `.` or `..`. Those are refused too, because the install would
        land outside the task's own directory.
      
      The full matrix of import shapes, including the ones that do not work, is
      pinned in `tests/test_multi_module_seed.py`.
      
      ## Where things live
      
      Task specs and shared defaults are read from `inputs_private/` in your project
      directory if present, otherwise from the copies bundled inside the package, so
      a fresh install runs immediately. Resolution is **per file**: dropping one task
      spec into `inputs_private/config/tasks/` overrides exactly that task and leaves
      everything else in place. Nothing is ever written back into the package.
      
      | Path | Contents |
      |---|---|
      | `config/tasks/<task_id>.yaml` | One task, as above. |
      | `data/<task_id>/` | That task's fixture data. |
      | `config/settings.yaml` | Retry budget, stage order, patch size limit. A private copy is overlaid section by section, so state only what you change. |
      | `agent_defs/*.md` | The prompts. `code_agent.md` and `test_agent.md` are the two system prompts; the rest are per-mode fragments. Not per-task: editing these changes every task's behaviour. |
      
      ## Proposing fixture rows
      
      A criterion no input row can trigger produces a test that passes whatever the
      code does. Across this project's own measurements roughly two thirds of
      deliberately planted faults were missed by every suite for that reason: the bar
      was unmeasurable rather than wrong.
      
      ```bash
      qikly --propose-fixtures --tasks <MY_TASKS>
      ```
      
      A separate agent reads your criteria and your fixture files and says, for each
      criterion, either `covered` or here is the smallest row that would reach it.
      The answer goes to `outputs/reports/fixture_proposals/`, laid out with each row
      printed under the criterion it exists to reach so you judge the two together.
      
      **It never edits a fixture.** To accept a row, paste it into the named file and
      append `  # proposed`. To reject one, do nothing. Two reasons for the gate,
      neither about the model being untrustworthy. A row is only right or wrong
      relative to its criterion, so it is harder to review than a sentence. And a
      fixture set that grows in whatever direction a model finds interesting stops
      resembling the data you actually process, at which point every rate measured on
      it describes a world that does not exist. The report is capped at eight
      proposals per round and prints what share of your rows a machine has written,
      so that drift is visible in aggregate rather than one plausible row at a time.
      
      You can of course add rows by hand at any time, and always could. This exists
      because noticing *which* criteria have no data behind them is the tedious part.
      
      The refinement loop does this for you on what it adds. When
      `refine_acceptance_criteria` finishes with new criteria, it asks for rows the
      same way, lists the new criteria first in the report, and logs how many have no
      data that reaches them, so a sharper bar does not arrive partly unmeasurable.
      It still applies nothing.
      
      
    • TROUBLESHOOTING.md 20.5 KB
      # When a run does not converge
      
      **If a run has not started yet, skip to [Before a run: where did my files
      go?](#before-a-run-where-did-my-files-go) at the end.** Everything above that
      section assumes a run has already failed, and the commonest reports this
      project receives are not about runs at all.
      
      A stall is a normal outcome, not a broken tool. The run exits non-zero, names
      the tests that blocked it, keeps the whole record, and ships nothing. Across
      every measurement this project has taken, roughly four runs in ten stop this
      way, and no run has ever reported success on code its own tests rejected.
      
      So the question is never "why is it broken". It is which of a short list of
      things is happening, and the list is short.
      
      ---
      
      ## Triage
      
      Match what you saw to where to look. The rows are in the order to work through
      them: the one change that moves convergence most, then the checks that settle
      what happened, then fixes to the task file, and only then more attempts.
      
      | What you saw | Section | What to do |
      |---|---|---|
      | Poor results on a provider you just set up, or on the default model | [1. Try a stronger model](#1-try-a-stronger-model) | Set `LLM_MODEL` to a mid-tier or larger model and run again |
      | Any stall, before changing the task file | [2. Read what actually blocked it](#2-read-what-actually-blocked-it) | Change nothing yet: this step decides what to change. In the run's `_report.html` timeline, byte-identical patches go to [5](#5-the-same-patch-appearing-over-and-over), an import or syntax error to [3](#3-a-collection-error-means-nothing-ran), steady progress to [10](#10-give-it-more-attempts), and different patches that never fix the same test to [1](#1-try-a-stronger-model) |
      | `0 passed, 0 failed, 1 error` | [3. A collection error means nothing ran](#3-a-collection-error-means-nothing-ran) | Make `interface.module` and the declared signatures match what the tests import |
      | Everything suddenly worse than last week | [4. Check nothing is set that you have forgotten](#4-check-nothing-is-set-that-you-have-forgotten) | Run `qikly --trends --by week` and look for a setting that changed, such as `criteria_per_batch` |
      | The same test failing every iteration, no progress | [5. The same patch appearing over and over](#5-the-same-patch-appearing-over-and-over) | Move the restriction into `requirements`, or widen the criterion to match reality |
      | Stopped with a message naming two tests, each fix for one breaking the other | [6. Two generated tests disagree](#6-two-generated-tests-disagree) | Compare the two tests with the acceptance criteria. If one contradicts a criterion, run again without `--resume` so the suites are written and checked again, and leave a correct spec alone |
      | A stage spends its whole budget and never gets closer | [7. Check the criteria and requirements do not contradict each other](#7-check-the-criteria-and-requirements-do-not-contradict-each-other) | Run `qikly --check-criteria --tasks <MY_TASKS>` and correct whichever statement is wrong |
      | Tests check arbitrary values rather than the boundary | [8. Check the criteria name values, not adjectives](#8-check-the-criteria-name-values-not-adjectives) | Run `qikly --validate` and rewrite each flagged criterion as a value: "100 is accepted and 101 is rejected" |
      | A test passes whatever the code does | [9. Check your fixtures can reach every criterion](#9-check-your-fixtures-can-reach-every-criterion) | Run `propose_fixtures` and add the input rows it suggests |
      | Steady progress, then the budget ran out | [10. Give it more attempts](#10-give-it-more-attempts) | Raise `orchestrator.max_retries_per_stage` (default 10), only when the report shows progress |
      | **Nothing has run yet, and files you were told about are missing** | [Before a run](#before-a-run-where-did-my-files-go) | Read the first line of the command's output: it names the directory it worked in, and says when the project root is somewhere else |
      | **You are in a `demo/throwaway_<timestamp>/` folder** | [Before a run](#before-a-run-where-did-my-files-go) | That is a throwaway copy. Start your own project somewhere else |
      | `--score-code` says the suite does not pass | [Before a run](#before-a-run-where-did-my-files-go) | Usually pytest collected no tests at the path given to `--score-tests` |
      | **On macOS, no patch ever applies, on any task** | [12. On macOS, no patch ever applies](#12-on-macos-no-patch-ever-applies) | `brew install gpatch`. The system `patch` is BSD and rejects the options qikly sends |
      | Integration and system pass, unit does not | [11. Expect the unit stage to be where it fails](#11-expect-the-unit-stage-to-be-where-it-fails) | Expected. Accept it, or leave the unit stage out with `orchestrator.test_order` |
      
      ---
      
      ## 1. Try a stronger model
      
      This moves convergence more than anything else here, and it is one environment
      variable.
      
      ```bash
      export LLM_MODEL=gpt-4o          # or a larger model on your provider
      qikly --tasks <MY_TASKS>         # e.g. --tasks CALC_TAX,MERGE_SALES
      ```
      
      Every convergence figure in this project was measured on
      `gemini-3.5-flash-lite`, a deliberately small and cheap model chosen so that
      sweeps of hundreds of runs were affordable. Treat those figures as a floor.
      
      An entry-level model on any provider may stall on a task a mid-tier one clears
      comfortably. If you are evaluating qikly, evaluate it on a model you would
      actually ship behind.
      
      **What a bigger model buys, and what it costs.** On one CALC_TAX run,
      `claude-sonnet-5` generated 44 tests against `gemini-3.5-flash-lite`'s 24, and
      its suite rejected code that Gemini's suite accepted, on five tests, while
      Gemini's suite accepted its code entirely. A stricter bar, in other words. It
      also took 403 seconds against 31, and cost \$0.81 against \$0.005.
      
      That trade is worth making deliberately rather than by default. A reasoning
      model produces thinking tokens you are billed for and wait on, which is why the
      cheapest model is the default here and why every published figure was measured
      on it: a 400-run sweep costs about \$3 on the default and roughly \$320 on a
      reasoning model.
      
      **Use the cheap model to measure and the expensive one to work.** If you need a
      convergence rate, take it on the default. If you need the strictest bar for one
      important specification, pay for it once. One run of each is an anecdote, not a
      comparison; the figures above are a single run per model.
      
      ### The default model is not the same size on every provider
      
      <a id="provider-defaults"></a>Set no `LLM_MODEL` and each provider gets its own
      default, and they are not the same class of model. This is the first thing to
      check when a run takes far longer on one provider than another:
      
      | Provider | Default model | What that means for a run |
      |---|---|---|
      | Gemini | `gemini-3.5-flash-lite` | Small, cheap, no reasoning step. Every published figure here was measured on it. A demo task runs in well under a minute |
      | OpenAI | `gpt-4o` | Mid-tier. Slower and dearer than the Gemini default |
      | Anthropic | `claude-sonnet-5` | A reasoning model. qikly sends no thinking configuration, and on this model that means adaptive thinking runs by default, so every call thinks before it answers |
      
      The Anthropic default is the one that surprises people. The same demo task that
      finishes in under a minute on the Gemini default has taken around sixteen
      minutes on it, for the same seven or so model calls. Nothing is wrong when that
      happens: you are watching a reasoning model think, and it produces a stricter
      suite for it.
      
      **For a like-for-like comparison with the Gemini default, name the model:**
      
      ```bash
      export LLM_MODEL=claude-haiku-4-5    # Windows PowerShell: $env:LLM_MODEL = "claude-haiku-4-5"
      ```
      
      Haiku is the closest Anthropic analogue to a flash-lite class model, and qikly
      sends no thinking budget, which that model needs before it will think at all.
      So a run on it spends no time or tokens on reasoning.
      
      Keep `claude-sonnet-5` when you want the stricter bar, and expect the run to
      take minutes rather than seconds. What qikly does not yet expose is the middle
      setting: the API takes a reasoning effort level, and a way to ask for less of
      it without changing model would make this a dial rather than a switch.
      
      **If it looks stuck**, it probably is not. A single call can legitimately run
      for minutes on a reasoning model. After ten seconds of silence a run starts
      saying so, one line every fifteen: `[patch] still waiting on the model, 45s`.
      Those lines are the difference between slow and stalled, and
      `QIKLY_NO_PROGRESS=1` turns them off. One
      call is abandoned after `QIKLY_REQUEST_TIMEOUT` seconds, 300 by default, which
      was chosen when the slowest observed call was well under a minute; on a
      reasoning model consider raising it, or a slow-but-working call is thrown away
      and retried from scratch.
      
      ## 2. Read what actually blocked it
      
      Every run writes a timeline:
      
      ```
      outputs/reports/iterations/<task>_<timestamp>_report.html
      ```
      
      Open it in a browser. It shows every iteration, the FIX reasoning and the PATCH
      diff for each failure, and, most usefully, **which patches applied cleanly and
      changed nothing.** A run full of those is not a run that needs more attempts.
      It is [3](#3-a-collection-error-means-nothing-ran) or
      [5](#5-the-same-patch-appearing-over-and-over).
      
      ## 3. A collection error means nothing ran
      
      ```
      [MY_TASK] [stage 1/3] [iteration 3] 0 passed, 0 failed, 1 error, 0 skipped
      ```
      
      No test failed, because no test ran. The module could not be imported. This is
      a different problem from a wrong answer, and until it is fixed nothing else can
      be assessed.
      
      The FIX prompt is told this explicitly, and the report carries the underlying
      `ImportError` or `SyntaxError`. The usual cause is a mismatch between what
      `interface` declares and what the agent wrote, so check that
      `interface.module` and the declared function signatures are exactly what the
      tests should be importing.
      
      ## 4. Check nothing is set that you have forgotten
      
      ```bash
      qikly --trends --by week
      ```
      
      Convergence per task over time, from the run summaries already on disk. Every
      period names the model and settings behind it, and a period where those changed
      is marked.
      
      This exists because of a specific, expensive mistake. `criteria_per_batch` in a
      settings file controls how many acceptance criteria a single test-generation
      call is shown. At `0` one call sees the whole bar. At `4` the bar is split into
      batches and each gets its own call, so a long bar produces roughly three times
      as many tests, and every run has three times as much to satisfy.
      
      Left set from an earlier experiment, it made convergence appear to collapse
      across nine tasks at once. Half a day went into diffing prompts, specs and
      provider parameters before anyone looked at the override.
      
      **A rate belongs to a tool, a model and a configuration together.** A rate that
      moved when the configuration moved is not a finding.
      
      ## 5. The same patch appearing over and over
      
      Identical diffs, not merely a repeated failure, is a specific signature: the
      model is fighting something it correctly knows about the world.
      
      Restrict a real-world field to an artificial subset, say three valid street
      suffixes, and the model will keep widening the restriction back. Not out of
      disobedience. Every piece of its training agrees that "Boulevard" is a street
      suffix, and your criterion is the outlier. At temperature zero this does not
      converge slowly; it does not converge at all.
      
      **Fix:** widen the criterion to match reality, or move the restriction into
      `requirements`, where the coding agent can read it and treat it as a given
      rather than as an error to correct.
      
      To confirm it, compare successive diffs under
      `outputs/logs/patches/<task>/<timestamp>/`. Byte-identical patches mean this.
      Different patches that never resolve the same test mean something else: a bug
      that needs more than the failure text to fix, which is [1](#1-try-a-stronger-model).
      
      ## 6. Two generated tests disagree
      
      ```
      Stopped on stage 'system' after 4 attempts: the last three patches alternated
      between the same two diffs. [...] The failing tests alternate between
      test_run_headway_two_second_rule_warning and test_integration_pipeline_flow in
      the 'system' and 'integration' suites: each fix for one breaks the other [...]
      ```
      
      Each patch makes one test pass and the other fail, because the two tests expect
      different results for the same input. No code can pass both, so more attempts
      cannot help, and the fault is in the tests rather than the specification. In
      the run behind this section, an integration test warned at exactly 2.00 seconds
      of headway and a system test did not, against a criterion saying exactly 2.00
      seconds raises no warning.
      
      The same mistake can also be made identically in both suites. Then they agree
      with each other and still contradict the criterion, the run fails without this
      message, and the place to look is the same: each test's comparison at every
      limit the criteria state.
      
      An opt-in check, `check_suites: true` under `test_generation` in settings,
      looks for these before any code is written and rewrites a suite once. Measured
      on one task it found every wrong suite but did not raise convergence, and a
      wrong finding once led a correct suite to be rewritten wrong, so it is off by
      default.
      
      **Fix:** compare the two named tests with the acceptance criteria. If one
      contradicts a criterion, run again without `--resume`, so the suites are written
      and checked again. Do not change a specification that is already right: this is
      the one stall where the spec is not the problem.
      
      To check suites already on disk without a run, one model call per task:
      
      ```bash
      python -m qikly.orchestrator.tuning.check_suites --tasks <MY_TASKS>
      ```
      
      ## 7. Check the criteria and requirements do not contradict each other
      
      ```bash
      qikly --check-criteria --tasks <MY_TASKS>
      ```
      
      One model call per task, and it changes nothing. A criterion that no
      implementation could satisfy alongside the requirements produces a stage that
      spends its entire budget discovering that the slow way. It exits non-zero on a
      contradiction, so a pipeline can gate on it.
      
      ## 8. Check the criteria name values, not adjectives
      
      ```bash
      qikly --validate
      ```
      
      Free, offline, and it flags criteria written as adjectives.
      
      > "Reject large amounts" invites a test at some arbitrary large number.
      > "100 is accepted and 101 is rejected" forces a test at the boundary.
      
      This is the highest-leverage habit in writing a bar. A suite that never tests a
      boundary cannot catch an error at that boundary, no matter how many other cases
      it covers, and off-by-one at a boundary is among the oldest defect classes in
      software.
      
      `--validate` also catches the quiet structural mistakes: `acceptance_criteria`
      written as one long string instead of a list, a `task_id` that disagrees with
      its filename, and fixture paths that do not resolve. Each of those otherwise
      surfaces twenty minutes and several dollars into a run.
      
      ## 9. Check your fixtures can reach every criterion
      
      ```bash
      qikly --propose-fixtures --tasks <MY_TASKS>
      ```
      
      A criterion that no input row can trigger produces a test that passes whatever
      the code does. The bar is not lower; part of it is absent.
      
      Eight of the ten tasks bundled with qikly had at least one before this was run
      on them, from one in `CALC_CALENDAR` to seven of thirteen in `MERGE_CONTACTS`.
      Assume yours do too.
      
      It writes proposals to a file and never edits your data.
      
      ## 10. Give it more attempts
      
      `orchestrator.max_retries_per_stage` in `config/settings.yaml`, default 10.
      
      Worth raising when the report shows steady progress that simply ran out of
      room. Not worth raising when it shows the same patch repeating: that run will
      fail identically with a hundred attempts, and cost ten times as much doing it.
      
      ## 11. Expect the unit stage to be where it fails
      
      About twenty points of the gap between "passes integration and system" and
      "passes everything" is the unit stage, consistently, across every sweep this
      project has run.
      
      The reason is structural rather than mysterious: the unit suite is the largest,
      runs last, and is the only one written with sight of the implementation.
      [design_2_performance.md](https://github.com/gal-a/qikly/blob/main/docs/design_2_performance.md#nearly-the-whole-gap-between-those-two-numbers-is-the-unit-stage)
      has the full explanation.
      
      If behavioural verification is what you need, `orchestrator.test_order` in
      settings can leave it out.
      
      ## 12. On macOS, no patch ever applies
      
      Every generated diff fails, on every task, from the first iteration, for a
      reason that reads like the model's fault and is not.
      
      The system `patch` on macOS is BSD, and it rejects the options qikly sends.
      Install GNU patch and the same run goes through:
      
      ```bash
      brew install gpatch
      ```
      
      Nothing else changes. If patches apply on one machine and fail on all of them
      on another, this is the first thing to check.
      
      ---
      
      ---
      
      ## One run is an artifact, not a rate
      
      The same task with the same seed converges on some runs and not others. Before
      concluding anything about a task, a model or a setting:
      
      ```bash
      python -m qikly.orchestrator.run_all --tasks <MY_TASKS> --repeat 10 \
          --skip-eval --skip-refine
      ```
      
      That writes an aggregate with a confidence interval instead of a pass count.
      Ten runs is usually enough to tell a real difference from noise, and it is
      worth knowing that at n=10 the intervals are wide: this project has measured
      the same unchanged task at 72% and then 90% on consecutive sweeps.
      
      If a change looks like an improvement after one run, it is not yet evidence of
      anything.
      
      ---
      
      ## Before a run: where did my files go?
      
      None of the sections above apply if nothing has run yet, and this is the
      question that arrives most often.
      
      ### It said it created files and they are not there
      
      They almost certainly are, somewhere you did not look. qikly moves to a
      resolved **project root** when it starts, which can be a different directory
      from the one you are standing in, and `--init`, `--example` and `--scaffold`
      write relative to one of those two.
      
      Every command now opens with a line that settles it:
      
      ```
      qikly 0.5.4  2026-09-27 11:27:18  run in C:\Users\you\my-project
        project root is elsewhere: C:\some\other\place
        (set QIKLY_PROJECT_ROOT to choose it, or cd there)
      ```
      
      The second and third lines appear only when the two differ. If you see them,
      that is your answer. If you are on an older version that does not print them,
      upgrade, or search for one of the files by name:
      
      ```powershell
      Get-ChildItem $HOME -Recurse -Filter "MY_METRICS*" -ErrorAction SilentlyContinue | Select-Object FullName
      ```
      
      To pin the project explicitly rather than let it be inferred, set
      `QIKLY_PROJECT_ROOT` to the directory you mean.
      
      ### You are standing in the demo's folder
      
      `qikly --demo` runs in a throwaway `demo/throwaway_<timestamp>/` directory so it cannot
      touch anything of yours, which also means **everything in it goes when you
      delete the folder, and nothing in it is yours**. Somebody who has just watched
      the demo work is standing in something that looks exactly like a working
      project, and the obvious next move is to start theirs there.
      
      qikly now refuses, names the folder, and says where to go instead. If you
      deliberately kept that directory and want to work in it, delete the
      `.qikly-demo` marker inside it and the refusal stops.
      
      ### The version you are running is not the version you installed
      
      Two things can disagree. `qikly --version` reports what actually runs;
      `pip show qikly` reports metadata that an interrupted or repeated upgrade can
      leave stale. Trust `--version`. If they disagree, clean it:
      
      ```powershell
      pip uninstall qikly -y
      pip install qikly
      ```
      
      `qikly --version` also prints the package directory and the interpreter, which
      is what to check when a flag the documentation describes does not exist.
      
      ### `--score-code` says your suite does not pass
      
      It refuses to score a suite that does not pass your untouched code, because
      every planted fault would then fail for the reason the original does and the
      number would mean nothing. Two causes, likeliest first:
      
      - **pytest collected nothing.** The path given to `--score-tests` holds no
        tests, or none that match its discovery rules. Run pytest on that path
        yourself and read what it says.
      - **Your suite genuinely fails.** Fix that first, then score it.
      
      There is no reachability warning in this mode, unlike `--score-suite`: there is
      no task file, so there are no criteria to be unreachable. A fault that survives
      may still sit on a line no test executes at all, which is a gap in what your
      tests reach rather than in what they assert.
      
      ---
      
      ## Still stuck
      
      The run kept everything. `outputs/logs/transactions_<task>_<timestamp>.jsonl`
      is an append-only record of every test run, every FIX, every PATCH and every
      apply outcome, and it is the source of truth that the reports are rendered
      from.
      
      Issues and results are welcome:
      [github.com/gal-a/qikly/issues](https://github.com/gal-a/qikly/issues).
      
  • SKILL.md 26.5 KB
    ---
    name: qikly
    description: Write tests that can actually fail, by withholding the acceptance criteria from the agent that writes the code. Use when someone does not trust a suite that passes. Use when they want tests written from a specification rather than from the code. Use when they ask whether a specification is testable, or want an existing suite scored by planting faults in the code. Python modules, through the qikly tool. Also use whenever qikly or spec-driven testing is mentioned.
    license: Apache-2.0
    metadata:
      author: Gal Arav
      homepage: https://github.com/gal-a/qikly
      version: 0.3.3
      requires: qikly >= 0.5.4
    ---
    
    # qikly: tests written from a spec the coder never read
    
    > **If you read one thing here, read this.** Asked for tests for a module with
    > no specification: do not write tests, and do not run the module to find out
    > what it should do. A value obtained by executing an implementation is not a
    > contract. A suite built from those values passes by construction and cannot
    > disagree with a bug, because the bug is what it was written from. Say that in
    > one sentence, then offer the two routes in *The request this arrives as*.
    >
    > **Four things look settled and are not:** whether a threshold is inclusive,
    > whether a discount applies before or after tax, where a floor or cap lands,
    > which way a half-cent rounds. An implementation answers all four. None of
    > those answers is a specification. Ask the user. Do not read them off the code.
    >
    > **If they insist on tests anyway, write them.** Refusing a second time is
    > not your call. Say which lines of the implementation you had to treat as
    > correct, name every boundary and tie-break you copied rather than were
    > told, and the suite is at least honest about its foundation.
    
    ## What this is for
    
    If you write a module and then write its tests, both come from one reading of
    the same ambiguous sentences. The
    suite goes green and the green means nothing.
    
    **Python modules only.** qikly scaffolds from Python signatures and generates
    pytest suites; this Skill says nothing about any other language, and you should
    say so rather than guess if asked.
    
    ## The request this arrives as, and what to do with it
    
    **"Write tests for `my_module.py`", with no specification anywhere.** That is
    how this almost always begins, and answering it literally is the mistake this
    whole tool exists to prevent. Tests read off an implementation can only
    describe what it already does: they pass by construction, they encode every
    choice the code happened to make as though it were intended, and they cannot
    disagree with a bug because the bug is what they were written from.
    
    So do not open by writing tests. Say that in one sentence, then offer the two
    things that do work, and let the user choose:
    
    - **`qikly --score-code my_module.py --score-tests their_tests.py`** if any
      suite already exists. Free, no model call, nothing of theirs modified, and it
      answers "is the suite we have worth anything" with named faults it missed.
      This is usually the right first move on somebody else's project.
    - **`qikly --scaffold my_module.py`**, then fill in `requirements` and
      `acceptance_criteria` together, if what they actually want is tests that
      could fail. The scaffold reads their real signatures; the two halves are the
      part only they can write, and the whole value is in that being a separate act
      from writing the code.
    
    ## Ask economically, or you will lose them
    
    Every decision you surface costs the user attention, and there is no version of
    this where nobody has to decide anything. So the cost has to come out of the
    **shape** of the asking, not out of the number of decisions.
    
    **Propose a complete draft and invite corrections. Never interrogate.** One
    decision, wait, next decision, wait, is how a ten-line module produces four
    modal questions and a user who closes the window. Instead:
    
    > Here is what I am going to assume, all of it going into `requirements` where
    > the coding agent will read it:
    > - percent is between 0 and 100 inclusive
    > - quantity, unit price and subtotal are zero or more
    > - money is not rounded to the cent
    >
    > Tell me which of those is wrong, or say go ahead.
    
    The discipline is identical: the decisions are still stated, still in the half
    the coding agent reads, still the user's. What changes is that agreeing costs
    one keystroke, and disagreeing is easier too, because somebody reading a list
    spots the wrong line faster than somebody three dialogs deep.
    
    **And only raise a decision the tests actually turn on.** If no criterion you
    are going to write would differ between the two answers, you are spending their
    attention for nothing. The question earns its place when you can say which test
    changes.
    
    **Two you should always raise, because they are silent and expensive:** the
    rounding rule wherever money appears, since "nearest cent" settles nothing at a
    half cent, and the inclusive or exclusive end of any boundary. Those two account
    for most of the decisions that get hidden in the wrong half.
    
    **If they insist on tests now**, write them, and then say plainly which lines
    of the implementation you had to treat as correct: every boundary, every
    tie-break, every validation rule you copied rather than were told. Those are
    the decisions nothing has settled, and naming them is the difference between a
    suite that is honest about its foundation and one that looks authoritative.
    
    A real session got this half right: it wrote seventeen tests from the code,
    checked they could fail by planting faults, and only then said "I wrote these
    by reading your code, so where the code made a choice, the tests assume that
    choice was right". The saying so was correct. The order was backwards.
    
    qikly splits one specification in two. Test generation reads the whole thing.
    The coding agent receives the same file with the `acceptance_criteria` section
    cut out, and when a test fails it sees the failure, never the criterion it
    broke. That separation is enforced in qikly's own code, not by instructions in
    this file, which matters: **this skill cannot keep anything hidden. The tool
    does that. This skill only helps you use the tool correctly.**
    
    **And be precise about what the withholding covers**, because a user will
    eventually stretch it. It is one thing: inside a qikly run, the coding agent's
    prompt is assembled without the `acceptance_criteria` section. It says nothing
    about this conversation. If someone pastes their criteria to you, you have
    read them, and no part of qikly prevented that or knows it happened. Say so
    plainly if you are asked whether talking to you is covered.
    
    ## Install
    
    ```bash
    pip install qikly
    qikly --version          # version, package directory and interpreter
    ```
    
    **This Skill needs qikly 0.5.4 or later**, which is where `--score-code`
    arrives, alongside `--score-suite` and the reachability warning in `--validate`
    from 0.5.3. Against an older install an
    agent following this page will recommend a flag that does not exist, so check
    `qikly --version` before trusting the command table below. The Skill's own
    version is separate from the tool's: it changes when these instructions
    change, not when qikly releases.
    
    A run needs one provider key, `GEMINI_API_KEY`, `OPENAI_API_KEY` or
    `ANTHROPIC_API_KEY`. Several commands need no key and cost nothing; the table
    further down says which, and when to reach for each.
    
    **Which key is present decides how long a run takes.** On
    `ANTHROPIC_API_KEY` qikly uses `claude-sonnet-5`, which thinks before every
    answer, so a ten-call loop becomes minutes plus thinking tokens you are billed
    for. The published figures come from `gemini-3.5-flash-lite`, fast and cheap
    enough to repeat. Say which key a run will use, and what it means for the wait,
    before starting it. That goes for you too: a reasoning model driving this tool
    charges the user the same wait at every step, so keep the mechanical steps
    mechanical.
    
    **So set the model rather than only warning about it.** Before the first paid
    command, if `ANTHROPIC_API_KEY` is the key in play, set `LLM_MODEL` to
    `claude-haiku-4-5` for the run and say you have done it and why. That is the
    only provider whose default thinks: Gemini's is already
    `gemini-3.5-flash-lite` and OpenAI's is `gpt-4o`, neither of which has a
    reasoning step, so there is nothing to change on either. If the user has
    chosen a thinking model deliberately, leave it alone and say what the wait
    will be: the stricter suite it writes is a real reason to want one.
    
    ## The one question that decides everything
    
    Every line of a specification goes in one of two halves, and this settles it:
    
    > **Given only the requirements, could two competent developers legitimately
    > disagree about this line?**
    
    **Yes, it is a decision.** It belongs in `requirements`, where the coding agent
    reads it. Nobody can guess a choice somebody made: a threshold, a unit, a
    measurement convention, an exemption.
    
    **No, it follows.** It belongs in `acceptance_criteria`, which are withheld.
    The exact boundary, the identity that must hold, the case a careless reading
    gets wrong.
    
    | In `requirements`, a decision | In `acceptance_criteria`, a consequence |
    |---|---|
    | Keep at least 2.5 m from the vehicle ahead, centre to centre | At exactly 2.5 m, no violation is raised |
    | Amounts are currency, rounded to the nearest cent | For every accepted row, total equals subtotal plus tax, exactly |
    | Dates are written YYYY-MM-DD | 2026-02-30 is rejected, because it is not a real date |
    
    **A worked case, because this one is easy to get backwards.** A spec says
    *"warn when following distance breaks the two-second rule"*, and the
    acceptance criteria say *"a headway of exactly 2.00 s does not raise a
    warning"*. That split is **wrong**, and the reason is the question above:
    "breaks the two-second rule" reads as "below two seconds" just as naturally as
    "at or below two seconds", so two competent developers can disagree about
    2.00 s itself. It is a decision, and it has to move into `requirements`.
    Measured on this exact task, leaving it in the criteria converged 3 runs in 10;
    moving that one sentence converged 10 in 10. Read that as what it is: evidence
    that the agent was guessing, not evidence that the suite got better. A run that
    cannot converge is a run that proves nothing at all, which is a different
    problem from a suite that converges and proves little.
    
    Note what moving it does **not** mean. Moving a decision is a correction the
    spec always needed, and it is justified by the question alone, without looking
    at any code. Copying a criterion's boundary value into the requirements to get
    a run green is the opposite: it tells both agents the answer, and the test that
    checks it then passes first try and proves nothing.
    
    **Both mistakes have a signature.**
    
    A *decision* hidden in the criteria shows up as repetition: near-identical FIX
    and PATCH each round, or two tests disagreeing where every patch fixes one and
    breaks the other. Sometimes it guesses right and the run goes green, which is
    worse, because nothing then tells you the line was in the wrong half.
    
    A *consequence* left in the requirements is quieter: the test for it passes
    first try, nothing was learned, and the run looks entirely normal.
    
    **Spotting one is your job; settling it is not.** You can tell a decision from
    a consequence by applying the question above to the words on the page, and you
    should: say which line you think is in the wrong half and why. What you cannot
    do is choose the answer. Whether a headway of exactly 2.00 s warns, whether an
    empty string counts as missing, whether currency rounds half up or half even:
    nothing in the specification settles those, which is what makes them decisions,
    and guessing on the user's behalf puts an invented choice into their
    requirements where it will look decided. Name the ambiguity, propose the
    wording, ask which way they want it.
    
    qikly's own refinement loop cannot do even the spotting: it can tighten a bar
    the code already attempts, and it cannot tell you a line is in the wrong
    place.
    
    ## Writing criteria that can be tested
    
    **Name values, not adjectives.** "Reject large amounts" produces a test at some
    arbitrary large number. "100 is accepted and 101 is rejected" forces the
    boundary.
    
    **Name both sides of a boundary in one criterion.** "The 250 limit is
    inclusive: exactly 250 is accepted and 250.01 is rejected" is one criterion
    closing one ambiguity, and it tells you exactly which two rows the data needs.
    
    **But check first which half the boundary belongs to, because this technique
    and the worked case above look identical on the page.** One rule separates
    them, and it is about the requirement's own wording:
    
    > Does the requirement already settle which side the edge falls on?
    
    "Keep **at least** 2.5 m" settles it: at 2.5 m you comply, so a criterion
    saying no violation is raised at exactly 2.5 m only spells out what was
    already decided. Write it as a criterion.
    
    "Warn when the headway **breaks** the two-second rule" does not settle it, and
    neither does "rounded to the **nearest** cent" when a value lands exactly
    halfway, or "reject rows where quantity is **missing**" when nobody said
    whether an empty string counts. In each case the criterion would be making the
    choice rather than recording it. Move the choice into the requirement, then
    write the criterion for its consequence.
    
    Words that usually settle it: at least, at most, above, below, strictly, on or
    after. Words that usually do not: nearest, breaks, exceeds a limit, missing,
    invalid, malformed.
    
    **Every value you name has to exist in the data.** A criterion no input row can
    trigger produces a test that passes whatever the code does. This is the most
    common reason a suite measures less than it appears to.
    
    **Then go back through the requirements and pair them.** Every decision in the
    requirements should have a criterion that would catch it being implemented
    wrong, and this is the step people skip: they write the decisions carefully,
    write criteria for the two or three boundaries that worry them, and leave the
    rest of the specification unchecked. "Line totals are quantity times unit
    price" is a decision; the criterion that pairs with it is an identity, "for
    every accepted row, line_total equals quantity times unit price, to the penny".
    Without the pair, the agent can get the arithmetic wrong and nothing fails.
    
    Ask it as a sweep: for each requirement, what would a wrong implementation of
    this look like, and which criterion catches it? A requirement with no answer is
    a requirement nothing is testing.
    
    **And read your own requirements back against the word list above.** This
    applies to your own wording too, which is where it gets missed: writing
    "rounded to the nearest cent" creates the same ambiguity you would flag in
    somebody else's spec. If a requirement you drafted uses
    one of those words, say so and ask which way the boundary falls, rather than
    writing criteria that quietly avoid the case. A tie nobody decided is not a
    withheld consequence, it is a decision nobody made, and it will surface as a
    stalled loop later.
    
    ## The five steps, on the user's own module
    
    ```bash
    qikly --scaffold my_metrics.py          # reads real signatures, writes a task file
    # put a real sample of the data where the task's `inputs:` says
    # fill in `requirements` and `acceptance_criteria`, using the question above
    qikly --validate --tasks MY_METRICS_VERIFY    # free, no model call
    qikly --tasks MY_METRICS_VERIFY               # the paid one
    ```
    
    **Before that last line, check which provider key is set.** It decides the
    model, and therefore the wait and the bill; see Install above. On
    `ANTHROPIC_API_KEY` every call thinks before it answers, and a run of ten calls
    is minutes rather than seconds. Tell the user which one they are about to spend
    on before they spend it.
    
    **A task file has a third section.** `interface` names the module and the
    function signatures, and the coding agent reads it: it is how both agents agree
    what to call things. Scaffolding fills it in from the real signatures, so it
    rarely needs editing, but a line about the shape of the output belongs there
    rather than in either half above.
    
    `--scaffold my_metrics.py` writes a task that tests code that already exists.
    Add `--fresh` for one that writes a new implementation of the same interface
    and tests that. Both land in `inputs_private/config/tasks/`.
    
    **The task id comes from the module's filename, upper-cased**, so
    `pricing.py` gives `PRICING`, and the plain scaffold adds `_VERIFY` because it
    tests code you already have: `PRICING_VERIFY`. `--fresh` gives `PRICING`. The
    command prints the name and the path it wrote, so read that rather than
    guessing.
    
    **Never edit the user's module to make a test pass.** qikly does not, and
    neither should you.
    
    **If the project tells you to write code and tests together, say so rather
    than choosing silently.** A house rule like "always write the implementation
    and its tests in the same session so they stay consistent" is reasonable on its
    own terms and is the exact thing this tool exists to prevent: consistency by
    construction is what makes a suite unable to disagree. You cannot follow both.
    Tell the user the two conflict, in one sentence, and let them decide which
    applies here.
    
    To see the whole shape first, `qikly --example` lays down a finished worked
    task, module and sample data included, so you can read a filled-in pair before
    writing one, in the directory you are standing in.
    
    **`qikly --demo` is a different command and the difference matters.** It runs a
    bundled task end to end in a throwaway `demo/throwaway_<timestamp>/` folder that exists to
    be deleted, and it is the one most people try first. A user who has just watched
    it work is standing in something that looks exactly like a working project, and
    the obvious next move is to start theirs there. **Never set up someone's real
    project inside a demo folder.** If the user says they ran "the demo" and wants
    to continue where they are, establish which of the two commands they ran before
    writing anything. This has already cost a first-time user an afternoon.
    
    ## The free checks, and when to reach for each
    
    **If you arrived straight here, this is the rule you skipped, stated in full
    so you do not have to go back for it.** Asked for tests for a module with no
    specification: do not write tests, and do not run the module to find out what
    it should do. A value obtained by executing an implementation is not a
    contract, and a suite built from those values passes by construction and
    cannot disagree with a bug. Whether a threshold is inclusive, whether the
    discount applies before or after tax, where a floor lands, which way a
    half-cent rounds: the code answers all four and none of those answers is a
    specification. Deriving the expected values yourself is the failure this
    Skill exists to prevent, and it is the one an agent commits while believing it
    is being careful.
    
    **Which command fits depends on what they have.** With a suite already
    written, `--score-code` scores it and needs nothing else. With no suite and no
    spec, which is the case this paragraph is usually about, the route is
    `qikly --scaffold their_module.py` and then filling in the two sections with
    them. Everything else in the table below needs a task file that does not exist
    yet.
    
    | Command | Cost | Use it when |
    |---|---|---|
    | `qikly --validate --tasks X` | free | always, before any run. Catches a missing input file, a leftover TODO, a criterion made of adjectives, a requirement restating a criterion, and criteria the data cannot reach |
    | `qikly --explain X` | free | to show the user exactly what each agent receives, criteria present on one side and absent on the other |
    | `qikly --score-suite --tasks X` | free, and slow | after a run converges, to find what the suite would not have noticed. It plants one fault at a time in the code and reports which ones the tests missed. Free because every fault is an edit to the code's syntax tree and no model is asked anything; slow because each fault means running your whole suite again |
    | `qikly --score-code PATH --score-tests PATH` | free, and slow | **for a suite qikly did not write**, which is what somebody already has before they have anything else. Point it at a module or package and the tests for it, and it reports which planted faults the tests did not notice. No task file, no run, no model call, and the report lands beside their code |
    | `qikly --check-criteria --tasks X` | one model call | when a spec may contradict itself, before spending a run on it |
    | `qikly --propose-fixtures --tasks X` | one model call | when `--validate` says a criterion's values are missing from the data, to get the rows it would take |
    
    **Read `--score-suite` beside the reachability warning it prints above the
    number.** A suite cannot catch a fault in behaviour no input row exercises, so
    unreachable criteria lower the score for a reason that is about the fixtures
    and not about the tests. Fix the data first, then score.
    
    **A high score is not a clean bill of health, and say so before they read it
    as one.** Mutation scoring asks whether their tests notice changes to the code
    that exists. It cannot ask about a rule nobody implemented, because there is
    nothing there to break. In a real session a seventeen-test suite caught 8 of 8
    planted faults while a suite written from the specification found four genuine
    bugs in the same file. The number is a floor, not a verdict.
    
    **`--score-code` prints no such warning, and you must not imply it does.**
    There is no task file and therefore no criteria to be unreachable, so the
    report carries the score and nothing above it. The underlying problem has not
    gone away: a fault that survives may be on a line no test ever executes, which
    is a gap in what the tests reach rather than in what they assert. Say that to
    the user rather than handing them a percentage as a verdict on their suite.
    
    ## When a run does not converge
    
    It exits non-zero, names the tests that blocked it, and ships nothing. Read
    which stage stopped first: the unit stage is last and strictest, and most
    failures are there.
    
    Then, in order:
    
    1. **Is the loop repeating itself?** Near-identical FIX and PATCH each
       iteration means a decision is in the wrong half. Move it into
       `requirements`.
    2. **Do two tests disagree,** each patch fixing one and breaking the other? The
       specification contradicts itself. `--check-criteria` finds that before a run.
    3. **Can the data reach every criterion?** `--validate` now says, and
       `--propose-fixtures` drafts the rows.
    
    Then the cheap levers: a larger model, and more attempts. And never loosen a
    criterion to get green, for the reason given above.
    
    If no patch ever applies at all, and you are on macOS, that is section 12 of
    `references/TROUBLESHOOTING.md`.
    
    **What the agent sees when a test fails, exactly.** pytest's output for the
    failing test: its name, its own source and docstring, and the assertion error.
    Not the acceptance criterion. Because a generated test's docstring usually
    restates the rule it came from, a failing test does tend to give away its own
    case, and the project says so rather than pretending otherwise. It does not
    undo the split: the suite was written first, from criteria the coder never
    read, and nothing learned afterwards changes a test already on disk. There is
    a setting that narrows this, `diagnostic_feedback: staged` under `agent:` in
    settings.yaml, which starts the agent at a one-line error and widens only when
    a patch stops making progress. **It is off by default**, because every
    published convergence figure was measured with the full traceback and nobody
    has measured what starting narrow costs.
    
    **A run prints nothing while a model call is in flight**, which on a reasoning
    model can be minutes. After ten seconds it starts saying so, one line every
    fifteen: `[patch] still waiting on the model, 45s`. Those lines are the
    difference between slow and stalled, so pass them on rather than swallowing
    them, and do not conclude a run has hung while they are still arriving.
    `QIKLY_NO_PROGRESS=1` turns them off.
    
    **To stop a run spending more than you meant**, three environment variables
    bound it: `QIKLY_MAX_CALLS` and `QIKLY_MAX_TOKENS` bound one task's process,
    and `QIKLY_MAX_SWEEP_TOKENS` bounds a whole sweep. Set them before a first run
    on somebody's real code rather than after.
    
    **In CI**, qikly ships a GitHub Action. A run costs a model call per attempt,
    so per pull request is a budget decision rather than a technical one, and the
    free checks are the ones that belong on every commit: `--validate`,
    `--score-suite` for a suite qikly generated, and `--score-code` for one that
    predates it, which on most real repositories is the relevant half.
    
    ## What to tell the user honestly
    
    Roughly 8 runs in 10 produce code passing every integration and system test,
    and roughly 6 in 10 pass everything including unit tests, measured over 967
    runs and reproduced over 390 more. **Always say what those runs were:** a
    small, inexpensive model (`gemini-3.5-flash-lite`), chosen so the sweeps could
    be repeated affordably, so the figures are a floor rather than a ceiling. And
    they measure convergence, whether generated code passes the generated tests,
    not whether those tests catch real defects.
    
    **Two things are settled and one is not, and they are easy to confuse.**
    
    *Settled:* the coding agent never receives the acceptance criteria. That is a
    property of the code, checkable with `qikly --explain` and held by a test that
    fails the build if any call site ever leaks one. It is not a benchmark result
    and cannot go stale.
    
    *Settled:* the convergence figures above, measured and re-measured.
    
    *Open:* whether a suite written from withheld criteria catches **more real
    defects** than one written with sight of the code. That comparison has not been
    run, here or anywhere. A Google team has measured the step before it, that
    generating tests from a written contract rather than from the code raises bug
    detection by 9.8 points (arXiv 2608.17177), which is adjacent and not the same
    claim. Separately again, six experiments asked whether automatically *refining*
    the criteria produces a sharper bar and none detected an effect, which is an
    absence of evidence rather than evidence of absence. Three different questions.
    Do not offer any of them as evidence for another.
    
    ## References, and when to open them
    
    Do not read these by default. Each one is a full page, and everything above is
    enough for a first run.
    
    **Read `references/TASK_FILE_REFERENCE.md`** when you need a field this page
    does not name, when the user already has criteria written somewhere else and
    wants them imported, when they want to seed their own implementation or suite
    rather than have one generated, or when you need to know where a run writes
    each artefact.
    
    **Read `references/TROUBLESHOOTING.md`** when a run has already failed and the
    three questions above did not explain it. It has eleven numbered causes, each
    with its own fix, and the triage table at the top maps a symptom to a number.
    
    The repository, for anything neither covers:
    <https://github.com/gal-a/qikly>
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related