qikly
Write tests that can actually fail, by withholding the acceptance criteria from the agent that writes the code. Use when someone does not trust a suite that passes. Use when they want tests written from a specification rather than from the code. Use when they ask whether a specif
Install
npx skills add https://github.com/gal-a/qikly/tree/main/src/qikly/skills/qikly
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install gal-a-qikly@llmmart
git clone https://github.com/gal-a/qikly.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole gal-a/qikly collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
qikly: tests written from a spec the coder never read
If you read one thing here, read this. Asked for tests for a module with no specification: do not write tests, and do not run the module to find out what it should do. A value obtained by executing an implementation is not a contract. A suite built from those values passes by construction and cannot disagree with a bug, because the bug is what it was written from. Say that in one sentence, then offer the two routes in The request this arrives as.
Four things look settled and are not: whether a threshold is inclusive, whether a discount applies before or after tax, where a floor or cap lands, which way a half-cent rounds. An implementation answers all four. None of those answers is a specification. Ask the user. Do not read them off the code.
If they insist on tests anyway, write them. Refusing a second time is not your call. Say which lines of the implementation you had to treat as correct, name every boundary and tie-break you copied rather than were told, and the suite is at least honest about its foundation.
What this is for
If you write a module and then write its tests, both come from one reading of the same ambiguous sentences. The suite goes green and the green means nothing.
Python modules only. qikly scaffolds from Python signatures and generates pytest suites; this Skill says nothing about any other language, and you should say so rather than guess if asked.
The request this arrives as, and what to do with it
"Write tests for my_module.py", with no specification anywhere. That is
how this almost always begins, and answering it literally is the mistake this
whole tool exists to prevent. Tests read off an implementation can only
describe what it already does: they pass by construction, they encode every
choice the code happened to make as though it were intended, and they cannot
disagree with a bug because the bug is what they were written from.
So do not open by writing tests. Say that in one sentence, then offer the two things that do work, and let the user choose:
qikly --score-code my_module.py --score-tests their_tests.pyif any suite already exists. Free, no model call, nothing of theirs modified, and it answers "is the suite we have worth anything" with named faults it missed. This is usually the right first move on somebody else's project.qikly --scaffold my_module.py, then fill inrequirementsandacceptance_criteriatogether, if what they actually want is tests that could fail. The scaffold reads their real signatures; the two halves are the part only they can write, and the whole value is in that being a separate act from writing the code.
Ask economically, or you will lose them
Every decision you surface costs the user attention, and there is no version of this where nobody has to decide anything. So the cost has to come out of the shape of the asking, not out of the number of decisions.
Propose a complete draft and invite corrections. Never interrogate. One decision, wait, next decision, wait, is how a ten-line module produces four modal questions and a user who closes the window. Instead:
Here is what I am going to assume, all of it going into
requirementswhere the coding agent will read it:
- percent is between 0 and 100 inclusive
- quantity, unit price and subtotal are zero or more
- money is not rounded to the cent
Tell me which of those is wrong, or say go ahead.
The discipline is identical: the decisions are still stated, still in the half the coding agent reads, still the user's. What changes is that agreeing costs one keystroke, and disagreeing is easier too, because somebody reading a list spots the wrong line faster than somebody three dialogs deep.
And only raise a decision the tests actually turn on. If no criterion you are going to write would differ between the two answers, you are spending their attention for nothing. The question earns its place when you can say which test changes.
Two you should always raise, because they are silent and expensive: the rounding rule wherever money appears, since "nearest cent" settles nothing at a half cent, and the inclusive or exclusive end of any boundary. Those two account for most of the decisions that get hidden in the wrong half.
If they insist on tests now, write them, and then say plainly which lines of the implementation you had to treat as correct: every boundary, every tie-break, every validation rule you copied rather than were told. Those are the decisions nothing has settled, and naming them is the difference between a suite that is honest about its foundation and one that looks authoritative.
A real session got this half right: it wrote seventeen tests from the code, checked they could fail by planting faults, and only then said "I wrote these by reading your code, so where the code made a choice, the tests assume that choice was right". The saying so was correct. The order was backwards.
qikly splits one specification in two. Test generation reads the whole thing.
The coding agent receives the same file with the acceptance_criteria section
cut out, and when a test fails it sees the failure, never the criterion it
broke. That separation is enforced in qikly's own code, not by instructions in
this file, which matters: this skill cannot keep anything hidden. The tool
does that. This skill only helps you use the tool correctly.
And be precise about what the withholding covers, because a user will
eventually stretch it. It is one thing: inside a qikly run, the coding agent's
prompt is assembled without the acceptance_criteria section. It says nothing
about this conversation. If someone pastes their criteria to you, you have
read them, and no part of qikly prevented that or knows it happened. Say so
plainly if you are asked whether talking to you is covered.
Install
pip install qikly
qikly --version # version, package directory and interpreter
This Skill needs qikly 0.5.4 or later, which is where --score-code
arrives, alongside --score-suite and the reachability warning in --validate
from 0.5.3. Against an older install an
agent following this page will recommend a flag that does not exist, so check
qikly --version before trusting the command table below. The Skill's own
version is separate from the tool's: it changes when these instructions
change, not when qikly releases.
A run needs one provider key, GEMINI_API_KEY, OPENAI_API_KEY or
ANTHROPIC_API_KEY. Several commands need no key and cost nothing; the table
further down says which, and when to reach for each.
Which key is present decides how long a run takes. On
ANTHROPIC_API_KEY qikly uses claude-sonnet-5, which thinks before every
answer, so a ten-call loop becomes minutes plus thinking tokens you are billed
for. The published figures come from gemini-3.5-flash-lite, fast and cheap
enough to repeat. Say which key a run will use, and what it means for the wait,
before starting it. That goes for you too: a reasoning model driving this tool
charges the user the same wait at every step, so keep the mechanical steps
mechanical.
So set the model rather than only warning about it. Before the first paid
command, if ANTHROPIC_API_KEY is the key in play, set LLM_MODEL to
claude-haiku-4-5 for the run and say you have done it and why. That is the
only provider whose default thinks: Gemini's is already
gemini-3.5-flash-lite and OpenAI's is gpt-4o, neither of which has a
reasoning step, so there is nothing to change on either. If the user has
chosen a thinking model deliberately, leave it alone and say what the wait
will be: the stricter suite it writes is a real reason to want one.
The one question that decides everything
Every line of a specification goes in one of two halves, and this settles it:
Given only the requirements, could two competent developers legitimately disagree about this line?
Yes, it is a decision. It belongs in requirements, where the coding agent
reads it. Nobody can guess a choice somebody made: a threshold, a unit, a
measurement convention, an exemption.
No, it follows. It belongs in acceptance_criteria, which are withheld.
The exact boundary, the identity that must hold, the case a careless reading
gets wrong.
In requirements, a decision |
In acceptance_criteria, a consequence |
|---|---|
| Keep at least 2.5 m from the vehicle ahead, centre to centre | At exactly 2.5 m, no violation is raised |
| Amounts are currency, rounded to the nearest cent | For every accepted row, total equals subtotal plus tax, exactly |
| Dates are written YYYY-MM-DD | 2026-02-30 is rejected, because it is not a real date |
A worked case, because this one is easy to get backwards. A spec says
"warn when following distance breaks the two-second rule", and the
acceptance criteria say "a headway of exactly 2.00 s does not raise a
warning". That split is wrong, and the reason is the question above:
"breaks the two-second rule" reads as "below two seconds" just as naturally as
"at or below two seconds", so two competent developers can disagree about
2.00 s itself. It is a decision, and it has to move into requirements.
Measured on this exact task, leaving it in the criteria converged 3 runs in 10;
moving that one sentence converged 10 in 10. Read that as what it is: evidence
that the agent was guessing, not evidence that the suite got better. A run that
cannot converge is a run that proves nothing at all, which is a different
problem from a suite that converges and proves little.
Note what moving it does not mean. Moving a decision is a correction the spec always needed, and it is justified by the question alone, without looking at any code. Copying a criterion's boundary value into the requirements to get a run green is the opposite: it tells both agents the answer, and the test that checks it then passes first try and proves nothing.
Both mistakes have a signature.
A decision hidden in the criteria shows up as repetition: near-identical FIX and PATCH each round, or two tests disagreeing where every patch fixes one and breaks the other. Sometimes it guesses right and the run goes green, which is worse, because nothing then tells you the line was in the wrong half.
A consequence left in the requirements is quieter: the test for it passes first try, nothing was learned, and the run looks entirely normal.
Spotting one is your job; settling it is not. You can tell a decision from a consequence by applying the question above to the words on the page, and you should: say which line you think is in the wrong half and why. What you cannot do is choose the answer. Whether a headway of exactly 2.00 s warns, whether an empty string counts as missing, whether currency rounds half up or half even: nothing in the specification settles those, which is what makes them decisions, and guessing on the user's behalf puts an invented choice into their requirements where it will look decided. Name the ambiguity, propose the wording, ask which way they want it.
qikly's own refinement loop cannot do even the spotting: it can tighten a bar the code already attempts, and it cannot tell you a line is in the wrong place.
Writing criteria that can be tested
Name values, not adjectives. "Reject large amounts" produces a test at some arbitrary large number. "100 is accepted and 101 is rejected" forces the boundary.
Name both sides of a boundary in one criterion. "The 250 limit is inclusive: exactly 250 is accepted and 250.01 is rejected" is one criterion closing one ambiguity, and it tells you exactly which two rows the data needs.
But check first which half the boundary belongs to, because this technique and the worked case above look identical on the page. One rule separates them, and it is about the requirement's own wording:
Does the requirement already settle which side the edge falls on?
"Keep at least 2.5 m" settles it: at 2.5 m you comply, so a criterion saying no violation is raised at exactly 2.5 m only spells out what was already decided. Write it as a criterion.
"Warn when the headway breaks the two-second rule" does not settle it, and neither does "rounded to the nearest cent" when a value lands exactly halfway, or "reject rows where quantity is missing" when nobody said whether an empty string counts. In each case the criterion would be making the choice rather than recording it. Move the choice into the requirement, then write the criterion for its consequence.
Words that usually settle it: at least, at most, above, below, strictly, on or after. Words that usually do not: nearest, breaks, exceeds a limit, missing, invalid, malformed.
Every value you name has to exist in the data. A criterion no input row can trigger produces a test that passes whatever the code does. This is the most common reason a suite measures less than it appears to.
Then go back through the requirements and pair them. Every decision in the requirements should have a criterion that would catch it being implemented wrong, and this is the step people skip: they write the decisions carefully, write criteria for the two or three boundaries that worry them, and leave the rest of the specification unchecked. "Line totals are quantity times unit price" is a decision; the criterion that pairs with it is an identity, "for every accepted row, line_total equals quantity times unit price, to the penny". Without the pair, the agent can get the arithmetic wrong and nothing fails.
Ask it as a sweep: for each requirement, what would a wrong implementation of this look like, and which criterion catches it? A requirement with no answer is a requirement nothing is testing.
And read your own requirements back against the word list above. This applies to your own wording too, which is where it gets missed: writing "rounded to the nearest cent" creates the same ambiguity you would flag in somebody else's spec. If a requirement you drafted uses one of those words, say so and ask which way the boundary falls, rather than writing criteria that quietly avoid the case. A tie nobody decided is not a withheld consequence, it is a decision nobody made, and it will surface as a stalled loop later.
The five steps, on the user's own module
qikly --scaffold my_metrics.py # reads real signatures, writes a task file
# put a real sample of the data where the task's `inputs:` says
# fill in `requirements` and `acceptance_criteria`, using the question above
qikly --validate --tasks MY_METRICS_VERIFY # free, no model call
qikly --tasks MY_METRICS_VERIFY # the paid one
Before that last line, check which provider key is set. It decides the
model, and therefore the wait and the bill; see Install above. On
ANTHROPIC_API_KEY every call thinks before it answers, and a run of ten calls
is minutes rather than seconds. Tell the user which one they are about to spend
on before they spend it.
A task file has a third section. interface names the module and the
function signatures, and the coding agent reads it: it is how both agents agree
what to call things. Scaffolding fills it in from the real signatures, so it
rarely needs editing, but a line about the shape of the output belongs there
rather than in either half above.
--scaffold my_metrics.py writes a task that tests code that already exists.
Add --fresh for one that writes a new implementation of the same interface
and tests that. Both land in inputs_private/config/tasks/.
The task id comes from the module's filename, upper-cased, so
pricing.py gives PRICING, and the plain scaffold adds _VERIFY because it
tests code you already have: PRICING_VERIFY. --fresh gives PRICING. The
command prints the name and the path it wrote, so read that rather than
guessing.
Never edit the user's module to make a test pass. qikly does not, and neither should you.
If the project tells you to write code and tests together, say so rather than choosing silently. A house rule like "always write the implementation and its tests in the same session so they stay consistent" is reasonable on its own terms and is the exact thing this tool exists to prevent: consistency by construction is what makes a suite unable to disagree. You cannot follow both. Tell the user the two conflict, in one sentence, and let them decide which applies here.
To see the whole shape first, qikly --example lays down a finished worked
task, module and sample data included, so you can read a filled-in pair before
writing one, in the directory you are standing in.
qikly --demo is a different command and the difference matters. It runs a
bundled task end to end in a throwaway demo/throwaway_<timestamp>/ folder that exists to
be deleted, and it is the one most people try first. A user who has just watched
it work is standing in something that looks exactly like a working project, and
the obvious next move is to start theirs there. Never set up someone's real
project inside a demo folder. If the user says they ran "the demo" and wants
to continue where they are, establish which of the two commands they ran before
writing anything. This has already cost a first-time user an afternoon.
The free checks, and when to reach for each
If you arrived straight here, this is the rule you skipped, stated in full so you do not have to go back for it. Asked for tests for a module with no specification: do not write tests, and do not run the module to find out what it should do. A value obtained by executing an implementation is not a contract, and a suite built from those values passes by construction and cannot disagree with a bug. Whether a threshold is inclusive, whether the discount applies before or after tax, where a floor lands, which way a half-cent rounds: the code answers all four and none of those answers is a specification. Deriving the expected values yourself is the failure this Skill exists to prevent, and it is the one an agent commits while believing it is being careful.
Which command fits depends on what they have. With a suite already
written, --score-code scores it and needs nothing else. With no suite and no
spec, which is the case this paragraph is usually about, the route is
qikly --scaffold their_module.py and then filling in the two sections with
them. Everything else in the table below needs a task file that does not exist
yet.
| Command | Cost | Use it when |
|---|---|---|
qikly --validate --tasks X |
free | always, before any run. Catches a missing input file, a leftover TODO, a criterion made of adjectives, a requirement restating a criterion, and criteria the data cannot reach |
qikly --explain X |
free | to show the user exactly what each agent receives, criteria present on one side and absent on the other |
qikly --score-suite --tasks X |
free, and slow | after a run converges, to find what the suite would not have noticed. It plants one fault at a time in the code and reports which ones the tests missed. Free because every fault is an edit to the code's syntax tree and no model is asked anything; slow because each fault means running your whole suite again |
qikly --score-code PATH --score-tests PATH |
free, and slow | for a suite qikly did not write, which is what somebody already has before they have anything else. Point it at a module or package and the tests for it, and it reports which planted faults the tests did not notice. No task file, no run, no model call, and the report lands beside their code |
qikly --check-criteria --tasks X |
one model call | when a spec may contradict itself, before spending a run on it |
qikly --propose-fixtures --tasks X |
one model call | when --validate says a criterion's values are missing from the data, to get the rows it would take |
Read --score-suite beside the reachability warning it prints above the
number. A suite cannot catch a fault in behaviour no input row exercises, so
unreachable criteria lower the score for a reason that is about the fixtures
and not about the tests. Fix the data first, then score.
A high score is not a clean bill of health, and say so before they read it as one. Mutation scoring asks whether their tests notice changes to the code that exists. It cannot ask about a rule nobody implemented, because there is nothing there to break. In a real session a seventeen-test suite caught 8 of 8 planted faults while a suite written from the specification found four genuine bugs in the same file. The number is a floor, not a verdict.
--score-code prints no such warning, and you must not imply it does.
There is no task file and therefore no criteria to be unreachable, so the
report carries the score and nothing above it. The underlying problem has not
gone away: a fault that survives may be on a line no test ever executes, which
is a gap in what the tests reach rather than in what they assert. Say that to
the user rather than handing them a percentage as a verdict on their suite.
When a run does not converge
It exits non-zero, names the tests that blocked it, and ships nothing. Read which stage stopped first: the unit stage is last and strictest, and most failures are there.
Then, in order:
- Is the loop repeating itself? Near-identical FIX and PATCH each
iteration means a decision is in the wrong half. Move it into
requirements. - Do two tests disagree, each patch fixing one and breaking the other? The
specification contradicts itself.
--check-criteriafinds that before a run. - Can the data reach every criterion?
--validatenow says, and--propose-fixturesdrafts the rows.
Then the cheap levers: a larger model, and more attempts. And never loosen a criterion to get green, for the reason given above.
If no patch ever applies at all, and you are on macOS, that is section 12 of
references/TROUBLESHOOTING.md.
What the agent sees when a test fails, exactly. pytest's output for the
failing test: its name, its own source and docstring, and the assertion error.
Not the acceptance criterion. Because a generated test's docstring usually
restates the rule it came from, a failing test does tend to give away its own
case, and the project says so rather than pretending otherwise. It does not
undo the split: the suite was written first, from criteria the coder never
read, and nothing learned afterwards changes a test already on disk. There is
a setting that narrows this, diagnostic_feedback: staged under agent: in
settings.yaml, which starts the agent at a one-line error and widens only when
a patch stops making progress. It is off by default, because every
published convergence figure was measured with the full traceback and nobody
has measured what starting narrow costs.
A run prints nothing while a model call is in flight, which on a reasoning
model can be minutes. After ten seconds it starts saying so, one line every
fifteen: [patch] still waiting on the model, 45s. Those lines are the
difference between slow and stalled, so pass them on rather than swallowing
them, and do not conclude a run has hung while they are still arriving.
QIKLY_NO_PROGRESS=1 turns them off.
To stop a run spending more than you meant, three environment variables
bound it: QIKLY_MAX_CALLS and QIKLY_MAX_TOKENS bound one task's process,
and QIKLY_MAX_SWEEP_TOKENS bounds a whole sweep. Set them before a first run
on somebody's real code rather than after.
In CI, qikly ships a GitHub Action. A run costs a model call per attempt,
so per pull request is a budget decision rather than a technical one, and the
free checks are the ones that belong on every commit: --validate,
--score-suite for a suite qikly generated, and --score-code for one that
predates it, which on most real repositories is the relevant half.
What to tell the user honestly
Roughly 8 runs in 10 produce code passing every integration and system test,
and roughly 6 in 10 pass everything including unit tests, measured over 967
runs and reproduced over 390 more. Always say what those runs were: a
small, inexpensive model (gemini-3.5-flash-lite), chosen so the sweeps could
be repeated affordably, so the figures are a floor rather than a ceiling. And
they measure convergence, whether generated code passes the generated tests,
not whether those tests catch real defects.
Two things are settled and one is not, and they are easy to confuse.
Settled: the coding agent never receives the acceptance criteria. That is a
property of the code, checkable with qikly --explain and held by a test that
fails the build if any call site ever leaks one. It is not a benchmark result
and cannot go stale.
Settled: the convergence figures above, measured and re-measured.
Open: whether a suite written from withheld criteria catches more real defects than one written with sight of the code. That comparison has not been run, here or anywhere. A Google team has measured the step before it, that generating tests from a written contract rather than from the code raises bug detection by 9.8 points (arXiv 2608.17177), which is adjacent and not the same claim. Separately again, six experiments asked whether automatically refining the criteria produces a sharper bar and none detected an effect, which is an absence of evidence rather than evidence of absence. Three different questions. Do not offer any of them as evidence for another.
References, and when to open them
Do not read these by default. Each one is a full page, and everything above is enough for a first run.
Read references/TASK_FILE_REFERENCE.md when you need a field this page
does not name, when the user already has criteria written somewhere else and
wants them imported, when they want to seed their own implementation or suite
rather than have one generated, or when you need to know where a run writes
each artefact.
Read references/TROUBLESHOOTING.md when a run has already failed and the
three questions above did not explain it. It has eleven numbered causes, each
with its own fix, and the triage table at the top maps a symptom to a number.
The repository, for anything neither covers: https://github.com/gal-a/qikly
Files (qikly)
-
references
-
TASK_FILE_REFERENCE.md 21.9 KB
# Task file reference The parts of working on your own data that you look up rather than read through. The path to a first run is [the quick start](https://github.com/gal-a/qikly/blob/main/docs/QUICK_START_ON_YOUR_OWN_DATA.md); this is what it deliberately leaves out. ## You probably do not have to write the task file by hand The criteria usually exist already, in a feature page or a ticket, and the interface exists in the code. qikly reads both. ```bash # a markdown page, a ticket export, or a .feature file qikly --criteria-from feature.md --task-id MY_TASK # straight from Jira: needs JIRA_BASE_URL, JIRA_EMAIL, JIRA_API_TOKEN qikly --criteria-from-jira PROJ-412 --task-id MY_TASK # both halves at once: criteria from the page, interface from the module qikly --scaffold src/metrics/band.py --from-doc feature.md ``` Bullet lists, a headed `Acceptance Criteria` section and Gherkin `Scenario:` blocks are all understood. Your page stays the source of truth and nobody retypes anything. **One section is never filled for you: `requirements`.** The coding agent reads it, and a feature page usually restates its own acceptance criteria in the prose above them, so lifting requirements across would hand the criteria to the one agent that must never see them. `qikly --validate` warns if what you write there restates a criterion. ## Which command depends on which parts you already have A task file is one YAML file with three parts, and the split above is a split between them: 1. **`requirements`** what the code must do, in the words a person would use. The coding agent reads this. 2. **`interface`** the contract, and a description rather than code: the function signatures and the dotted path where the module will live. Both agents read it, and neither is handed an implementation to read from it. When the integration and system tests are written there is not one yet. 3. **`acceptance_criteria`** what counts as correct, each one checkable and naming its boundary value. **Only test generation reads this.** "Spec" below means 1 and 2 together, which is what the coding agent is given. A tick means you already have that part. One thing the three parts do not say, and it matters: **test generation never reads the implementation either.** Integration and system tests are written before any code exists, from the specification alone. The unit stage is the single exception, written last from the code that just cleared the earlier stages, because unit tests have to name real functions. | Where you are starting | #1 | #2 | #3 | Run | What happens | |---|:-:|:-:|:-:|---|---| | Before anything else: see what is withheld | | | | `qikly --explain <MY_TASK>`<br>e.g. `qikly --explain CALC_TAX` | Prints a task file twice, once as each agent receives it, and the difference between them. No API key, no model call, about a second. **You get:** the acceptance criteria on one side and the same file with them cut out on the other, which is the claim everything else rests on. Add `--html` for the same as a page you can share. | | Just looking | | | | `qikly --demo` | A bundled task end to end in a throwaway folder. Thirty seconds, under a cent. **You get:** a working implementation, three test suites, and the full record of every FIX and PATCH, in a directory you can delete. | | Code someone else wrote, and you want **that code** verified | | Y | | `qikly --scaffold <MY_MODULE>.py` | Scaffold reads the real signatures out of the file you point it at and fills in **#2** for you. **#1** and **#3** stay yours to write: criteria read out of an implementation can only describe what that implementation already does, which is a bar it passes by construction. **You get:** one task file that tests the code you already have. Add `--fresh` for one that writes a fresh implementation of the same interface instead. | | You know what it must do, not yet how to check it | Y | | | `qikly --init` | Creates the directory layout and one starter task to edit. Its criteria show the habit that matters most: name the value, not the quality. "100 is accepted and 101 is rejected" forces a test at the boundary; "amounts must be reasonable" does not. **You get:** a task file to fill in, with your fixtures where a run will look for them. | | Same, but you want a first draft of the bar | Y | Y | | `qikly --tasks <MY_TASKS>`<br>`--generate-criteria` | Drafts **#3** from **#1** alone, then runs. **You get:** a first draft of the bar written into your task file for you to correct, plus the implementation and suites. | | You have written all three | Y | Y | Y | `qikly --tasks <MY_TASKS>` | Everything you wrote is used, and nothing is drafted on your behalf. **You get:** an implementation, integration, system and unit suites, a convergence report, and a run summary recording the model and settings that produced them. | | You have all three but doubt they agree | Y | Y | Y | `qikly --check-criteria`<br>`--tasks <MY_TASKS>` | One model call asking whether any implementation could satisfy the description, **#1** and **#3** at once, and whether any two of **#3** agree with each other. Advisory, and exits non-zero on a contradiction so a pipeline can gate on it. **You get:** a list of the pairs that cannot both hold, before spending a stage budget on them. Two criteria setting different numbers on the same quantity are always reported, since that is a typo rather than a tighter bar. | | A previous run stopped before finishing | Y | Y | Y | `qikly --tasks <MY_TASKS>`<br>`--resume` | Generating the tests and the first implementation already cost model calls, and they are still on disk. This keeps them and picks up where it stopped, instead of paying for them twice. **You get:** the same outputs as a full run, without paying for the parts already built. | `<MY_TASKS>` is one task_id or several separated by commas. A task_id is a filename under `inputs_private/config/tasks/` without the `.yaml`: `--tasks CALC_TAX`, `--tasks CALC_TAX,MERGE_SALES`, or omit it to run every task found. `<MY_TASK>`, singular, takes exactly one. `QIKLY_MAX_CALLS=200 qikly` stops at a call limit rather than a bill. `--scaffold` reads the module path and the real signatures of every public function straight out of the file, because they are already there. It leaves `requirements` and `acceptance_criteria` for you, and that is deliberate: criteria derived from an implementation can only describe what that implementation already does, and a bar that agrees with the code by construction is the exact failure this tool exists to prevent. ## Writing a task by hand Nothing is written into the package, and nothing is written into your source tree. ### 1. Make the two directories Anywhere you want to work. The presence of `inputs_private/` is what marks a directory as your project. ```bash mkdir -p inputs_private/config/tasks mkdir -p inputs_private/data/MY_TASK ``` Or let `qikly --init` create both, plus a starter task to copy. Until one of those exists, there is nothing marking your directory, and the fallback in [Where things live](#where-things-live) applies. From a `pip install -e` checkout that fallback finds the checkout itself, so a run started in an empty directory writes its outputs there instead of where you are standing. Make the directory first, or set `QIKLY_PROJECT_ROOT` to say exactly where you mean. ### 2. Drop your fixture data in Plain input files, whatever your code should read. CSV, JSON, JSONL, anything. ```bash cp ~/somewhere/orders_jan.csv inputs_private/data/MY_TASK/input_01.csv cp ~/somewhere/orders_feb.csv inputs_private/data/MY_TASK/input_02.csv ``` Names are up to you, but they must match what you write in the task's `inputs:` list below. The generated program opens these **by literal relative path from your project directory**, so the path in the task file is the path that gets executed. That is also why the bundled fixtures are copied into `inputs_private/data/` on first run rather than resolved from inside the package: the generated code has no way to ask where the package lives. Fixtures are never overwritten once present, so an edited file stays edited. ### 3. Write the task file `inputs_private/config/tasks/MY_TASK.yaml`. The filename must match `task_id`. ```yaml task_id: "MY_TASK" # letters/digits/underscore, not starting with a digit task_name: "Order line-item tax" description: "Read two CSV files of order line items, validate them, compute tax per line, and write the result to a single JSON output alongside a reason for every rejected line." inputs: # literal paths, opened by the generated code - "inputs_private/data/MY_TASK/input_01.csv" - "inputs_private/data/MY_TASK/input_02.csv" outputs: - "outputs/data/MY_TASK/output.json" interface: # what test generation targets module: "outputs.agent_src.code.MY_TASK.calc" integration_functions: - "extract(input_path) -> list[dict] # reads one input file, returns raw rows" - "transform(rows) -> dict # validates and computes; returns {\"accepted\": [...], \"rejected\": [...]}" - "load(data, output_path) -> None # writes the result as JSON" system_entrypoint: "run_calc(input_paths, output_path) -> None # extract each path, then transform -> load" requirements: # THE DECISIONS. The coding agent sees only this. - "Read both CSV files listed in inputs and combine their rows before validation" - "Validate each row: order_id, item_price, quantity, tax_rate" - "Apply strict, real-world data-quality validation; reject anything malformed or out of range" - "For each valid row compute subtotal, tax owed, and line total as currency amounts" - "A rejected row is not silently dropped: record it with a brief, specific reason" - "Write a single JSON object with two keys, \"accepted\" and \"rejected\"" acceptance_criteria: # THE CONSEQUENCES. Withheld from the coding agent. - "All computed currency amounts are rounded to two decimal places using round-half-up, not banker's rounding and not truncation" - "For every accepted row, the reported total equals the reported subtotal plus the reported tax, exactly, to the cent" - "A tax_rate of exactly 0 is valid: the computed tax is 0.00 and the total equals the subtotal" - "Each rejected row names the specific field that caused rejection, not a generic message" ``` `interface.module` is the dotted path the generated tests will import. With no seed, or with a single-file seed, that is `outputs.agent_src.code.<task_id>.<name>`: the implementation is written there, so pick the final component freely and the rest is fixed by where outputs live. **A package seed is the exception**, because the package keeps its own name and becomes importable by it: write `interface.module: "mypkg.pricing"`, the path your own code already uses. See "Seeding a package" below. ### 4. Run it ```bash qikly --tasks <MY_TASKS> # or: python run.py --tasks <MY_TASKS> ``` Discovery is automatic; there is no registry to update. Results land in `outputs/`, and `outputs/reports/iterations/MY_TASK_<timestamp>_report.html` is the place to start reading. ## Bringing acceptance criteria you have already written Most teams have not got a blank page here. If you work in Jira, Linear, Azure DevOps or a design doc, the rules are usually already written down, because the process asks for them before any code is cut. A ticket routinely looks like this: ``` PROJ-412 Merge overlapping sales exports Description Combine two CSV exports into one file... Acceptance Criteria - A transaction in both files at the same amount appears once - A negative or missing amount is rejected, naming the field - Dates must be YYYY-MM-DD ``` Those bullets are exactly what `acceptance_criteria` wants. Save the ticket to a file and read them out: ```bash qikly --criteria-from ticket.md # print as YAML qikly --criteria-from ticket.md --task-id MY_TASK # write into that task qikly --criteria-from ticket.md >> inputs_private/config/tasks/MY_TASK.yaml ``` It understands plain bullet lists, an "Acceptance Criteria" heading in a longer document, and Gherkin `Scenario:` blocks with Given/When/Then. Only the criteria section is read, so pasting a whole ticket does not turn its description into part of the bar. Only YAML goes to stdout, so the third form above appends a valid block. **It will not invent criteria from prose.** A file with no list and no scenarios returns nothing and says so. A rule that nobody wrote is precisely the invented standard this tool exists to argue against, and once it is in the file it looks like every other line. There is no API token and no vendor integration involved. Copying the ticket into a file is the whole of it. **Read what comes out before you run.** Criteria lifted from a ticket are a draft: tickets are written for people, who fill in gaps that a test cannot. The criteria are the standard everything else is judged against, so they are worth a minute of your attention. ## Supplying your own acceptance criteria, code or tests The loop takes three inputs. **Each one can be yours or generated, independently and in any combination.** | Input | Default | To supply your own | |---|---|---| | **Acceptance criteria** | Yours | Already the default: write `acceptance_criteria` in the task file, as above, or lift them from a ticket with `--criteria-from` (below). Omit it and add `--generate-criteria` to have a first draft written for you instead. | | **Implementation** | Generated | `seed.implementation` in the task file. | | **Test suites** | Generated | `seed.tests`, per stage. | ### Auto-generating acceptance criteria A task with no `acceptance_criteria` still runs, with a warning rather than an error, because running one deliberately is a legitimate thing to do. What you lose is the point of the exercise: test generation has only `requirements` to work from, the coding agent has nothing sharper to fail against, and the run usually converges on the first attempt without exercising the loop at all. `--generate-criteria` writes a first draft from the requirements alone into `inputs_private/config/tasks/<task_id>.yaml` before the run starts. It is opt-in, it never touches a task that already has criteria, and it says what it wrote rather than editing your files quietly. Pointed at a bundled example it writes your own overriding copy and leaves the packaged original alone. **A generated bar is a draft, not ground truth.** It was written from the same requirements the coding agent reads, so a case it did not think to demand is not being withheld from anyone: the two halves agree because they came from one source, which is the failure mode this whole tool argues against. `--compare-criteria` scores a generated bar against yours when you want that difference measured rather than assumed. The optional `seed:` block: ```yaml seed: # A file or a directory, copied into outputs/agent_src/code/<task_id>/. # A single file keeps its own name, which must match interface.module. # A directory is treated as a package: see "Seeding a package" below. implementation: "seeds/MY_TASK/calc.py" # Per stage. Seeding a stage suppresses generation for that stage only. tests: integration: "seeds/MY_TASK/test_integration.py" unit: "seeds/MY_TASK/unit/" ``` Paths are relative to your project directory. Both keys are optional. **`seed.implementation` is how you point this at code you already have.** The run skips generating a first implementation and goes straight to testing and repairing yours. `--scaffold` writes this block for you by default, and `--fresh` writes a task without it, for a new implementation of the same interface. **`seed.tests` keeps a suite you already trust**, so the loop repairs the code against your tests rather than its own. Mixing works and is often what you want: seed the integration stage with your suite and let the tool generate unit tests against whatever code results. Three things to know: - **Seeded test suites are checked before the run starts.** Every file must parse, and at least one must be named `test_*.py` and contain a `def test_*` function. A problem raises immediately rather than retrying, since there is no second sample to draw from a file you wrote. - **Seeds are installed after the workspace reset, not instead of it.** Every run still begins from one declared state, so repeated runs stay comparable and no run inherits the previous one's residue. - **A seeded run measures something different from an unseeded one.** Do not pool them in a single rate. The orchestrator prints a NOTE on every seeded run to keep that visible. ### Seeding a package, when the implementation is more than one module **New in 0.5.5.** Point `seed.implementation` at a directory and it is treated as a package: it is copied in **under its own name**, and the task's code directory is put on the path for the test run, so the package resolves by the name your code already uses. ```yaml interface: module: "mypkg.pricing" # the module under test, by its real import path seed: implementation: "mypkg" # the package it lives in, copied in whole ``` Inside `mypkg/`, write imports exactly as you already do. All three shapes work: `from mypkg.money import to_cents`, `from .money import to_cents`, and `from .utils.rounding import half_up`. No `__init__.py` is required, and one that is there is kept. **Every module in the package is visible to the coding agent and every one is repairable**, and a single patch may change more than one of them. That is the difference the package form makes: with a single-file seed, a defect in a helper is found by the tests and cannot be fixed, and the run tells you so rather than working around it. **The directory you name is the boundary.** qikly does not follow imports and decide for itself which of your files an agent may rewrite, because the transitive closure of a real package has no natural edge and "it rewrote a shared module I never named" is a worse outcome than naming a folder. So put inside the seed what you want worked on, and leave a vendor library or a module you do not want touched outside it. Data files inside the package are copied too, since your code may open them. `__pycache__` and `.pyc` files are not: they are stale copies of the very modules the run is about to rewrite. Four limits worth knowing before you start: - **Before 0.5.5 a seeded directory was flattened**, dropping the folder's name. If you wrote a task against that behaviour, `interface.module` needs the package name adding to it. - **Modules that import each other circularly at the top level fail**, the same way they do in plain Python. This is not something a run can repair. - **The package's name may not be a standard-library module's name.** A package called `json` would shadow the real one for everything the run imports, so a seed naming one is refused with a message rather than discovered halfway through a stage. Names that clash with an *installed third-party* package are not checked, because what is installed varies by environment: if your package is called `yaml` or `requests`, rename it or seed the single module instead. - **`seed.implementation` must name the folder itself**, not a path that resolves to `.` or `..`. Those are refused too, because the install would land outside the task's own directory. The full matrix of import shapes, including the ones that do not work, is pinned in `tests/test_multi_module_seed.py`. ## Where things live Task specs and shared defaults are read from `inputs_private/` in your project directory if present, otherwise from the copies bundled inside the package, so a fresh install runs immediately. Resolution is **per file**: dropping one task spec into `inputs_private/config/tasks/` overrides exactly that task and leaves everything else in place. Nothing is ever written back into the package. | Path | Contents | |---|---| | `config/tasks/<task_id>.yaml` | One task, as above. | | `data/<task_id>/` | That task's fixture data. | | `config/settings.yaml` | Retry budget, stage order, patch size limit. A private copy is overlaid section by section, so state only what you change. | | `agent_defs/*.md` | The prompts. `code_agent.md` and `test_agent.md` are the two system prompts; the rest are per-mode fragments. Not per-task: editing these changes every task's behaviour. | ## Proposing fixture rows A criterion no input row can trigger produces a test that passes whatever the code does. Across this project's own measurements roughly two thirds of deliberately planted faults were missed by every suite for that reason: the bar was unmeasurable rather than wrong. ```bash qikly --propose-fixtures --tasks <MY_TASKS> ``` A separate agent reads your criteria and your fixture files and says, for each criterion, either `covered` or here is the smallest row that would reach it. The answer goes to `outputs/reports/fixture_proposals/`, laid out with each row printed under the criterion it exists to reach so you judge the two together. **It never edits a fixture.** To accept a row, paste it into the named file and append ` # proposed`. To reject one, do nothing. Two reasons for the gate, neither about the model being untrustworthy. A row is only right or wrong relative to its criterion, so it is harder to review than a sentence. And a fixture set that grows in whatever direction a model finds interesting stops resembling the data you actually process, at which point every rate measured on it describes a world that does not exist. The report is capped at eight proposals per round and prints what share of your rows a machine has written, so that drift is visible in aggregate rather than one plausible row at a time. You can of course add rows by hand at any time, and always could. This exists because noticing *which* criteria have no data behind them is the tedious part. The refinement loop does this for you on what it adds. When `refine_acceptance_criteria` finishes with new criteria, it asks for rows the same way, lists the new criteria first in the report, and logs how many have no data that reaches them, so a sharper bar does not arrive partly unmeasurable. It still applies nothing. -
TROUBLESHOOTING.md 20.5 KB
# When a run does not converge **If a run has not started yet, skip to [Before a run: where did my files go?](#before-a-run-where-did-my-files-go) at the end.** Everything above that section assumes a run has already failed, and the commonest reports this project receives are not about runs at all. A stall is a normal outcome, not a broken tool. The run exits non-zero, names the tests that blocked it, keeps the whole record, and ships nothing. Across every measurement this project has taken, roughly four runs in ten stop this way, and no run has ever reported success on code its own tests rejected. So the question is never "why is it broken". It is which of a short list of things is happening, and the list is short. --- ## Triage Match what you saw to where to look. The rows are in the order to work through them: the one change that moves convergence most, then the checks that settle what happened, then fixes to the task file, and only then more attempts. | What you saw | Section | What to do | |---|---|---| | Poor results on a provider you just set up, or on the default model | [1. Try a stronger model](#1-try-a-stronger-model) | Set `LLM_MODEL` to a mid-tier or larger model and run again | | Any stall, before changing the task file | [2. Read what actually blocked it](#2-read-what-actually-blocked-it) | Change nothing yet: this step decides what to change. In the run's `_report.html` timeline, byte-identical patches go to [5](#5-the-same-patch-appearing-over-and-over), an import or syntax error to [3](#3-a-collection-error-means-nothing-ran), steady progress to [10](#10-give-it-more-attempts), and different patches that never fix the same test to [1](#1-try-a-stronger-model) | | `0 passed, 0 failed, 1 error` | [3. A collection error means nothing ran](#3-a-collection-error-means-nothing-ran) | Make `interface.module` and the declared signatures match what the tests import | | Everything suddenly worse than last week | [4. Check nothing is set that you have forgotten](#4-check-nothing-is-set-that-you-have-forgotten) | Run `qikly --trends --by week` and look for a setting that changed, such as `criteria_per_batch` | | The same test failing every iteration, no progress | [5. The same patch appearing over and over](#5-the-same-patch-appearing-over-and-over) | Move the restriction into `requirements`, or widen the criterion to match reality | | Stopped with a message naming two tests, each fix for one breaking the other | [6. Two generated tests disagree](#6-two-generated-tests-disagree) | Compare the two tests with the acceptance criteria. If one contradicts a criterion, run again without `--resume` so the suites are written and checked again, and leave a correct spec alone | | A stage spends its whole budget and never gets closer | [7. Check the criteria and requirements do not contradict each other](#7-check-the-criteria-and-requirements-do-not-contradict-each-other) | Run `qikly --check-criteria --tasks <MY_TASKS>` and correct whichever statement is wrong | | Tests check arbitrary values rather than the boundary | [8. Check the criteria name values, not adjectives](#8-check-the-criteria-name-values-not-adjectives) | Run `qikly --validate` and rewrite each flagged criterion as a value: "100 is accepted and 101 is rejected" | | A test passes whatever the code does | [9. Check your fixtures can reach every criterion](#9-check-your-fixtures-can-reach-every-criterion) | Run `propose_fixtures` and add the input rows it suggests | | Steady progress, then the budget ran out | [10. Give it more attempts](#10-give-it-more-attempts) | Raise `orchestrator.max_retries_per_stage` (default 10), only when the report shows progress | | **Nothing has run yet, and files you were told about are missing** | [Before a run](#before-a-run-where-did-my-files-go) | Read the first line of the command's output: it names the directory it worked in, and says when the project root is somewhere else | | **You are in a `demo/throwaway_<timestamp>/` folder** | [Before a run](#before-a-run-where-did-my-files-go) | That is a throwaway copy. Start your own project somewhere else | | `--score-code` says the suite does not pass | [Before a run](#before-a-run-where-did-my-files-go) | Usually pytest collected no tests at the path given to `--score-tests` | | **On macOS, no patch ever applies, on any task** | [12. On macOS, no patch ever applies](#12-on-macos-no-patch-ever-applies) | `brew install gpatch`. The system `patch` is BSD and rejects the options qikly sends | | Integration and system pass, unit does not | [11. Expect the unit stage to be where it fails](#11-expect-the-unit-stage-to-be-where-it-fails) | Expected. Accept it, or leave the unit stage out with `orchestrator.test_order` | --- ## 1. Try a stronger model This moves convergence more than anything else here, and it is one environment variable. ```bash export LLM_MODEL=gpt-4o # or a larger model on your provider qikly --tasks <MY_TASKS> # e.g. --tasks CALC_TAX,MERGE_SALES ``` Every convergence figure in this project was measured on `gemini-3.5-flash-lite`, a deliberately small and cheap model chosen so that sweeps of hundreds of runs were affordable. Treat those figures as a floor. An entry-level model on any provider may stall on a task a mid-tier one clears comfortably. If you are evaluating qikly, evaluate it on a model you would actually ship behind. **What a bigger model buys, and what it costs.** On one CALC_TAX run, `claude-sonnet-5` generated 44 tests against `gemini-3.5-flash-lite`'s 24, and its suite rejected code that Gemini's suite accepted, on five tests, while Gemini's suite accepted its code entirely. A stricter bar, in other words. It also took 403 seconds against 31, and cost \$0.81 against \$0.005. That trade is worth making deliberately rather than by default. A reasoning model produces thinking tokens you are billed for and wait on, which is why the cheapest model is the default here and why every published figure was measured on it: a 400-run sweep costs about \$3 on the default and roughly \$320 on a reasoning model. **Use the cheap model to measure and the expensive one to work.** If you need a convergence rate, take it on the default. If you need the strictest bar for one important specification, pay for it once. One run of each is an anecdote, not a comparison; the figures above are a single run per model. ### The default model is not the same size on every provider <a id="provider-defaults"></a>Set no `LLM_MODEL` and each provider gets its own default, and they are not the same class of model. This is the first thing to check when a run takes far longer on one provider than another: | Provider | Default model | What that means for a run | |---|---|---| | Gemini | `gemini-3.5-flash-lite` | Small, cheap, no reasoning step. Every published figure here was measured on it. A demo task runs in well under a minute | | OpenAI | `gpt-4o` | Mid-tier. Slower and dearer than the Gemini default | | Anthropic | `claude-sonnet-5` | A reasoning model. qikly sends no thinking configuration, and on this model that means adaptive thinking runs by default, so every call thinks before it answers | The Anthropic default is the one that surprises people. The same demo task that finishes in under a minute on the Gemini default has taken around sixteen minutes on it, for the same seven or so model calls. Nothing is wrong when that happens: you are watching a reasoning model think, and it produces a stricter suite for it. **For a like-for-like comparison with the Gemini default, name the model:** ```bash export LLM_MODEL=claude-haiku-4-5 # Windows PowerShell: $env:LLM_MODEL = "claude-haiku-4-5" ``` Haiku is the closest Anthropic analogue to a flash-lite class model, and qikly sends no thinking budget, which that model needs before it will think at all. So a run on it spends no time or tokens on reasoning. Keep `claude-sonnet-5` when you want the stricter bar, and expect the run to take minutes rather than seconds. What qikly does not yet expose is the middle setting: the API takes a reasoning effort level, and a way to ask for less of it without changing model would make this a dial rather than a switch. **If it looks stuck**, it probably is not. A single call can legitimately run for minutes on a reasoning model. After ten seconds of silence a run starts saying so, one line every fifteen: `[patch] still waiting on the model, 45s`. Those lines are the difference between slow and stalled, and `QIKLY_NO_PROGRESS=1` turns them off. One call is abandoned after `QIKLY_REQUEST_TIMEOUT` seconds, 300 by default, which was chosen when the slowest observed call was well under a minute; on a reasoning model consider raising it, or a slow-but-working call is thrown away and retried from scratch. ## 2. Read what actually blocked it Every run writes a timeline: ``` outputs/reports/iterations/<task>_<timestamp>_report.html ``` Open it in a browser. It shows every iteration, the FIX reasoning and the PATCH diff for each failure, and, most usefully, **which patches applied cleanly and changed nothing.** A run full of those is not a run that needs more attempts. It is [3](#3-a-collection-error-means-nothing-ran) or [5](#5-the-same-patch-appearing-over-and-over). ## 3. A collection error means nothing ran ``` [MY_TASK] [stage 1/3] [iteration 3] 0 passed, 0 failed, 1 error, 0 skipped ``` No test failed, because no test ran. The module could not be imported. This is a different problem from a wrong answer, and until it is fixed nothing else can be assessed. The FIX prompt is told this explicitly, and the report carries the underlying `ImportError` or `SyntaxError`. The usual cause is a mismatch between what `interface` declares and what the agent wrote, so check that `interface.module` and the declared function signatures are exactly what the tests should be importing. ## 4. Check nothing is set that you have forgotten ```bash qikly --trends --by week ``` Convergence per task over time, from the run summaries already on disk. Every period names the model and settings behind it, and a period where those changed is marked. This exists because of a specific, expensive mistake. `criteria_per_batch` in a settings file controls how many acceptance criteria a single test-generation call is shown. At `0` one call sees the whole bar. At `4` the bar is split into batches and each gets its own call, so a long bar produces roughly three times as many tests, and every run has three times as much to satisfy. Left set from an earlier experiment, it made convergence appear to collapse across nine tasks at once. Half a day went into diffing prompts, specs and provider parameters before anyone looked at the override. **A rate belongs to a tool, a model and a configuration together.** A rate that moved when the configuration moved is not a finding. ## 5. The same patch appearing over and over Identical diffs, not merely a repeated failure, is a specific signature: the model is fighting something it correctly knows about the world. Restrict a real-world field to an artificial subset, say three valid street suffixes, and the model will keep widening the restriction back. Not out of disobedience. Every piece of its training agrees that "Boulevard" is a street suffix, and your criterion is the outlier. At temperature zero this does not converge slowly; it does not converge at all. **Fix:** widen the criterion to match reality, or move the restriction into `requirements`, where the coding agent can read it and treat it as a given rather than as an error to correct. To confirm it, compare successive diffs under `outputs/logs/patches/<task>/<timestamp>/`. Byte-identical patches mean this. Different patches that never resolve the same test mean something else: a bug that needs more than the failure text to fix, which is [1](#1-try-a-stronger-model). ## 6. Two generated tests disagree ``` Stopped on stage 'system' after 4 attempts: the last three patches alternated between the same two diffs. [...] The failing tests alternate between test_run_headway_two_second_rule_warning and test_integration_pipeline_flow in the 'system' and 'integration' suites: each fix for one breaks the other [...] ``` Each patch makes one test pass and the other fail, because the two tests expect different results for the same input. No code can pass both, so more attempts cannot help, and the fault is in the tests rather than the specification. In the run behind this section, an integration test warned at exactly 2.00 seconds of headway and a system test did not, against a criterion saying exactly 2.00 seconds raises no warning. The same mistake can also be made identically in both suites. Then they agree with each other and still contradict the criterion, the run fails without this message, and the place to look is the same: each test's comparison at every limit the criteria state. An opt-in check, `check_suites: true` under `test_generation` in settings, looks for these before any code is written and rewrites a suite once. Measured on one task it found every wrong suite but did not raise convergence, and a wrong finding once led a correct suite to be rewritten wrong, so it is off by default. **Fix:** compare the two named tests with the acceptance criteria. If one contradicts a criterion, run again without `--resume`, so the suites are written and checked again. Do not change a specification that is already right: this is the one stall where the spec is not the problem. To check suites already on disk without a run, one model call per task: ```bash python -m qikly.orchestrator.tuning.check_suites --tasks <MY_TASKS> ``` ## 7. Check the criteria and requirements do not contradict each other ```bash qikly --check-criteria --tasks <MY_TASKS> ``` One model call per task, and it changes nothing. A criterion that no implementation could satisfy alongside the requirements produces a stage that spends its entire budget discovering that the slow way. It exits non-zero on a contradiction, so a pipeline can gate on it. ## 8. Check the criteria name values, not adjectives ```bash qikly --validate ``` Free, offline, and it flags criteria written as adjectives. > "Reject large amounts" invites a test at some arbitrary large number. > "100 is accepted and 101 is rejected" forces a test at the boundary. This is the highest-leverage habit in writing a bar. A suite that never tests a boundary cannot catch an error at that boundary, no matter how many other cases it covers, and off-by-one at a boundary is among the oldest defect classes in software. `--validate` also catches the quiet structural mistakes: `acceptance_criteria` written as one long string instead of a list, a `task_id` that disagrees with its filename, and fixture paths that do not resolve. Each of those otherwise surfaces twenty minutes and several dollars into a run. ## 9. Check your fixtures can reach every criterion ```bash qikly --propose-fixtures --tasks <MY_TASKS> ``` A criterion that no input row can trigger produces a test that passes whatever the code does. The bar is not lower; part of it is absent. Eight of the ten tasks bundled with qikly had at least one before this was run on them, from one in `CALC_CALENDAR` to seven of thirteen in `MERGE_CONTACTS`. Assume yours do too. It writes proposals to a file and never edits your data. ## 10. Give it more attempts `orchestrator.max_retries_per_stage` in `config/settings.yaml`, default 10. Worth raising when the report shows steady progress that simply ran out of room. Not worth raising when it shows the same patch repeating: that run will fail identically with a hundred attempts, and cost ten times as much doing it. ## 11. Expect the unit stage to be where it fails About twenty points of the gap between "passes integration and system" and "passes everything" is the unit stage, consistently, across every sweep this project has run. The reason is structural rather than mysterious: the unit suite is the largest, runs last, and is the only one written with sight of the implementation. [design_2_performance.md](https://github.com/gal-a/qikly/blob/main/docs/design_2_performance.md#nearly-the-whole-gap-between-those-two-numbers-is-the-unit-stage) has the full explanation. If behavioural verification is what you need, `orchestrator.test_order` in settings can leave it out. ## 12. On macOS, no patch ever applies Every generated diff fails, on every task, from the first iteration, for a reason that reads like the model's fault and is not. The system `patch` on macOS is BSD, and it rejects the options qikly sends. Install GNU patch and the same run goes through: ```bash brew install gpatch ``` Nothing else changes. If patches apply on one machine and fail on all of them on another, this is the first thing to check. --- --- ## One run is an artifact, not a rate The same task with the same seed converges on some runs and not others. Before concluding anything about a task, a model or a setting: ```bash python -m qikly.orchestrator.run_all --tasks <MY_TASKS> --repeat 10 \ --skip-eval --skip-refine ``` That writes an aggregate with a confidence interval instead of a pass count. Ten runs is usually enough to tell a real difference from noise, and it is worth knowing that at n=10 the intervals are wide: this project has measured the same unchanged task at 72% and then 90% on consecutive sweeps. If a change looks like an improvement after one run, it is not yet evidence of anything. --- ## Before a run: where did my files go? None of the sections above apply if nothing has run yet, and this is the question that arrives most often. ### It said it created files and they are not there They almost certainly are, somewhere you did not look. qikly moves to a resolved **project root** when it starts, which can be a different directory from the one you are standing in, and `--init`, `--example` and `--scaffold` write relative to one of those two. Every command now opens with a line that settles it: ``` qikly 0.5.4 2026-09-27 11:27:18 run in C:\Users\you\my-project project root is elsewhere: C:\some\other\place (set QIKLY_PROJECT_ROOT to choose it, or cd there) ``` The second and third lines appear only when the two differ. If you see them, that is your answer. If you are on an older version that does not print them, upgrade, or search for one of the files by name: ```powershell Get-ChildItem $HOME -Recurse -Filter "MY_METRICS*" -ErrorAction SilentlyContinue | Select-Object FullName ``` To pin the project explicitly rather than let it be inferred, set `QIKLY_PROJECT_ROOT` to the directory you mean. ### You are standing in the demo's folder `qikly --demo` runs in a throwaway `demo/throwaway_<timestamp>/` directory so it cannot touch anything of yours, which also means **everything in it goes when you delete the folder, and nothing in it is yours**. Somebody who has just watched the demo work is standing in something that looks exactly like a working project, and the obvious next move is to start theirs there. qikly now refuses, names the folder, and says where to go instead. If you deliberately kept that directory and want to work in it, delete the `.qikly-demo` marker inside it and the refusal stops. ### The version you are running is not the version you installed Two things can disagree. `qikly --version` reports what actually runs; `pip show qikly` reports metadata that an interrupted or repeated upgrade can leave stale. Trust `--version`. If they disagree, clean it: ```powershell pip uninstall qikly -y pip install qikly ``` `qikly --version` also prints the package directory and the interpreter, which is what to check when a flag the documentation describes does not exist. ### `--score-code` says your suite does not pass It refuses to score a suite that does not pass your untouched code, because every planted fault would then fail for the reason the original does and the number would mean nothing. Two causes, likeliest first: - **pytest collected nothing.** The path given to `--score-tests` holds no tests, or none that match its discovery rules. Run pytest on that path yourself and read what it says. - **Your suite genuinely fails.** Fix that first, then score it. There is no reachability warning in this mode, unlike `--score-suite`: there is no task file, so there are no criteria to be unreachable. A fault that survives may still sit on a line no test executes at all, which is a gap in what your tests reach rather than in what they assert. --- ## Still stuck The run kept everything. `outputs/logs/transactions_<task>_<timestamp>.jsonl` is an append-only record of every test run, every FIX, every PATCH and every apply outcome, and it is the source of truth that the reports are rendered from. Issues and results are welcome: [github.com/gal-a/qikly/issues](https://github.com/gal-a/qikly/issues).
-
-
SKILL.md 26.5 KB
--- name: qikly description: Write tests that can actually fail, by withholding the acceptance criteria from the agent that writes the code. Use when someone does not trust a suite that passes. Use when they want tests written from a specification rather than from the code. Use when they ask whether a specification is testable, or want an existing suite scored by planting faults in the code. Python modules, through the qikly tool. Also use whenever qikly or spec-driven testing is mentioned. license: Apache-2.0 metadata: author: Gal Arav homepage: https://github.com/gal-a/qikly version: 0.3.3 requires: qikly >= 0.5.4 --- # qikly: tests written from a spec the coder never read > **If you read one thing here, read this.** Asked for tests for a module with > no specification: do not write tests, and do not run the module to find out > what it should do. A value obtained by executing an implementation is not a > contract. A suite built from those values passes by construction and cannot > disagree with a bug, because the bug is what it was written from. Say that in > one sentence, then offer the two routes in *The request this arrives as*. > > **Four things look settled and are not:** whether a threshold is inclusive, > whether a discount applies before or after tax, where a floor or cap lands, > which way a half-cent rounds. An implementation answers all four. None of > those answers is a specification. Ask the user. Do not read them off the code. > > **If they insist on tests anyway, write them.** Refusing a second time is > not your call. Say which lines of the implementation you had to treat as > correct, name every boundary and tie-break you copied rather than were > told, and the suite is at least honest about its foundation. ## What this is for If you write a module and then write its tests, both come from one reading of the same ambiguous sentences. The suite goes green and the green means nothing. **Python modules only.** qikly scaffolds from Python signatures and generates pytest suites; this Skill says nothing about any other language, and you should say so rather than guess if asked. ## The request this arrives as, and what to do with it **"Write tests for `my_module.py`", with no specification anywhere.** That is how this almost always begins, and answering it literally is the mistake this whole tool exists to prevent. Tests read off an implementation can only describe what it already does: they pass by construction, they encode every choice the code happened to make as though it were intended, and they cannot disagree with a bug because the bug is what they were written from. So do not open by writing tests. Say that in one sentence, then offer the two things that do work, and let the user choose: - **`qikly --score-code my_module.py --score-tests their_tests.py`** if any suite already exists. Free, no model call, nothing of theirs modified, and it answers "is the suite we have worth anything" with named faults it missed. This is usually the right first move on somebody else's project. - **`qikly --scaffold my_module.py`**, then fill in `requirements` and `acceptance_criteria` together, if what they actually want is tests that could fail. The scaffold reads their real signatures; the two halves are the part only they can write, and the whole value is in that being a separate act from writing the code. ## Ask economically, or you will lose them Every decision you surface costs the user attention, and there is no version of this where nobody has to decide anything. So the cost has to come out of the **shape** of the asking, not out of the number of decisions. **Propose a complete draft and invite corrections. Never interrogate.** One decision, wait, next decision, wait, is how a ten-line module produces four modal questions and a user who closes the window. Instead: > Here is what I am going to assume, all of it going into `requirements` where > the coding agent will read it: > - percent is between 0 and 100 inclusive > - quantity, unit price and subtotal are zero or more > - money is not rounded to the cent > > Tell me which of those is wrong, or say go ahead. The discipline is identical: the decisions are still stated, still in the half the coding agent reads, still the user's. What changes is that agreeing costs one keystroke, and disagreeing is easier too, because somebody reading a list spots the wrong line faster than somebody three dialogs deep. **And only raise a decision the tests actually turn on.** If no criterion you are going to write would differ between the two answers, you are spending their attention for nothing. The question earns its place when you can say which test changes. **Two you should always raise, because they are silent and expensive:** the rounding rule wherever money appears, since "nearest cent" settles nothing at a half cent, and the inclusive or exclusive end of any boundary. Those two account for most of the decisions that get hidden in the wrong half. **If they insist on tests now**, write them, and then say plainly which lines of the implementation you had to treat as correct: every boundary, every tie-break, every validation rule you copied rather than were told. Those are the decisions nothing has settled, and naming them is the difference between a suite that is honest about its foundation and one that looks authoritative. A real session got this half right: it wrote seventeen tests from the code, checked they could fail by planting faults, and only then said "I wrote these by reading your code, so where the code made a choice, the tests assume that choice was right". The saying so was correct. The order was backwards. qikly splits one specification in two. Test generation reads the whole thing. The coding agent receives the same file with the `acceptance_criteria` section cut out, and when a test fails it sees the failure, never the criterion it broke. That separation is enforced in qikly's own code, not by instructions in this file, which matters: **this skill cannot keep anything hidden. The tool does that. This skill only helps you use the tool correctly.** **And be precise about what the withholding covers**, because a user will eventually stretch it. It is one thing: inside a qikly run, the coding agent's prompt is assembled without the `acceptance_criteria` section. It says nothing about this conversation. If someone pastes their criteria to you, you have read them, and no part of qikly prevented that or knows it happened. Say so plainly if you are asked whether talking to you is covered. ## Install ```bash pip install qikly qikly --version # version, package directory and interpreter ``` **This Skill needs qikly 0.5.4 or later**, which is where `--score-code` arrives, alongside `--score-suite` and the reachability warning in `--validate` from 0.5.3. Against an older install an agent following this page will recommend a flag that does not exist, so check `qikly --version` before trusting the command table below. The Skill's own version is separate from the tool's: it changes when these instructions change, not when qikly releases. A run needs one provider key, `GEMINI_API_KEY`, `OPENAI_API_KEY` or `ANTHROPIC_API_KEY`. Several commands need no key and cost nothing; the table further down says which, and when to reach for each. **Which key is present decides how long a run takes.** On `ANTHROPIC_API_KEY` qikly uses `claude-sonnet-5`, which thinks before every answer, so a ten-call loop becomes minutes plus thinking tokens you are billed for. The published figures come from `gemini-3.5-flash-lite`, fast and cheap enough to repeat. Say which key a run will use, and what it means for the wait, before starting it. That goes for you too: a reasoning model driving this tool charges the user the same wait at every step, so keep the mechanical steps mechanical. **So set the model rather than only warning about it.** Before the first paid command, if `ANTHROPIC_API_KEY` is the key in play, set `LLM_MODEL` to `claude-haiku-4-5` for the run and say you have done it and why. That is the only provider whose default thinks: Gemini's is already `gemini-3.5-flash-lite` and OpenAI's is `gpt-4o`, neither of which has a reasoning step, so there is nothing to change on either. If the user has chosen a thinking model deliberately, leave it alone and say what the wait will be: the stricter suite it writes is a real reason to want one. ## The one question that decides everything Every line of a specification goes in one of two halves, and this settles it: > **Given only the requirements, could two competent developers legitimately > disagree about this line?** **Yes, it is a decision.** It belongs in `requirements`, where the coding agent reads it. Nobody can guess a choice somebody made: a threshold, a unit, a measurement convention, an exemption. **No, it follows.** It belongs in `acceptance_criteria`, which are withheld. The exact boundary, the identity that must hold, the case a careless reading gets wrong. | In `requirements`, a decision | In `acceptance_criteria`, a consequence | |---|---| | Keep at least 2.5 m from the vehicle ahead, centre to centre | At exactly 2.5 m, no violation is raised | | Amounts are currency, rounded to the nearest cent | For every accepted row, total equals subtotal plus tax, exactly | | Dates are written YYYY-MM-DD | 2026-02-30 is rejected, because it is not a real date | **A worked case, because this one is easy to get backwards.** A spec says *"warn when following distance breaks the two-second rule"*, and the acceptance criteria say *"a headway of exactly 2.00 s does not raise a warning"*. That split is **wrong**, and the reason is the question above: "breaks the two-second rule" reads as "below two seconds" just as naturally as "at or below two seconds", so two competent developers can disagree about 2.00 s itself. It is a decision, and it has to move into `requirements`. Measured on this exact task, leaving it in the criteria converged 3 runs in 10; moving that one sentence converged 10 in 10. Read that as what it is: evidence that the agent was guessing, not evidence that the suite got better. A run that cannot converge is a run that proves nothing at all, which is a different problem from a suite that converges and proves little. Note what moving it does **not** mean. Moving a decision is a correction the spec always needed, and it is justified by the question alone, without looking at any code. Copying a criterion's boundary value into the requirements to get a run green is the opposite: it tells both agents the answer, and the test that checks it then passes first try and proves nothing. **Both mistakes have a signature.** A *decision* hidden in the criteria shows up as repetition: near-identical FIX and PATCH each round, or two tests disagreeing where every patch fixes one and breaks the other. Sometimes it guesses right and the run goes green, which is worse, because nothing then tells you the line was in the wrong half. A *consequence* left in the requirements is quieter: the test for it passes first try, nothing was learned, and the run looks entirely normal. **Spotting one is your job; settling it is not.** You can tell a decision from a consequence by applying the question above to the words on the page, and you should: say which line you think is in the wrong half and why. What you cannot do is choose the answer. Whether a headway of exactly 2.00 s warns, whether an empty string counts as missing, whether currency rounds half up or half even: nothing in the specification settles those, which is what makes them decisions, and guessing on the user's behalf puts an invented choice into their requirements where it will look decided. Name the ambiguity, propose the wording, ask which way they want it. qikly's own refinement loop cannot do even the spotting: it can tighten a bar the code already attempts, and it cannot tell you a line is in the wrong place. ## Writing criteria that can be tested **Name values, not adjectives.** "Reject large amounts" produces a test at some arbitrary large number. "100 is accepted and 101 is rejected" forces the boundary. **Name both sides of a boundary in one criterion.** "The 250 limit is inclusive: exactly 250 is accepted and 250.01 is rejected" is one criterion closing one ambiguity, and it tells you exactly which two rows the data needs. **But check first which half the boundary belongs to, because this technique and the worked case above look identical on the page.** One rule separates them, and it is about the requirement's own wording: > Does the requirement already settle which side the edge falls on? "Keep **at least** 2.5 m" settles it: at 2.5 m you comply, so a criterion saying no violation is raised at exactly 2.5 m only spells out what was already decided. Write it as a criterion. "Warn when the headway **breaks** the two-second rule" does not settle it, and neither does "rounded to the **nearest** cent" when a value lands exactly halfway, or "reject rows where quantity is **missing**" when nobody said whether an empty string counts. In each case the criterion would be making the choice rather than recording it. Move the choice into the requirement, then write the criterion for its consequence. Words that usually settle it: at least, at most, above, below, strictly, on or after. Words that usually do not: nearest, breaks, exceeds a limit, missing, invalid, malformed. **Every value you name has to exist in the data.** A criterion no input row can trigger produces a test that passes whatever the code does. This is the most common reason a suite measures less than it appears to. **Then go back through the requirements and pair them.** Every decision in the requirements should have a criterion that would catch it being implemented wrong, and this is the step people skip: they write the decisions carefully, write criteria for the two or three boundaries that worry them, and leave the rest of the specification unchecked. "Line totals are quantity times unit price" is a decision; the criterion that pairs with it is an identity, "for every accepted row, line_total equals quantity times unit price, to the penny". Without the pair, the agent can get the arithmetic wrong and nothing fails. Ask it as a sweep: for each requirement, what would a wrong implementation of this look like, and which criterion catches it? A requirement with no answer is a requirement nothing is testing. **And read your own requirements back against the word list above.** This applies to your own wording too, which is where it gets missed: writing "rounded to the nearest cent" creates the same ambiguity you would flag in somebody else's spec. If a requirement you drafted uses one of those words, say so and ask which way the boundary falls, rather than writing criteria that quietly avoid the case. A tie nobody decided is not a withheld consequence, it is a decision nobody made, and it will surface as a stalled loop later. ## The five steps, on the user's own module ```bash qikly --scaffold my_metrics.py # reads real signatures, writes a task file # put a real sample of the data where the task's `inputs:` says # fill in `requirements` and `acceptance_criteria`, using the question above qikly --validate --tasks MY_METRICS_VERIFY # free, no model call qikly --tasks MY_METRICS_VERIFY # the paid one ``` **Before that last line, check which provider key is set.** It decides the model, and therefore the wait and the bill; see Install above. On `ANTHROPIC_API_KEY` every call thinks before it answers, and a run of ten calls is minutes rather than seconds. Tell the user which one they are about to spend on before they spend it. **A task file has a third section.** `interface` names the module and the function signatures, and the coding agent reads it: it is how both agents agree what to call things. Scaffolding fills it in from the real signatures, so it rarely needs editing, but a line about the shape of the output belongs there rather than in either half above. `--scaffold my_metrics.py` writes a task that tests code that already exists. Add `--fresh` for one that writes a new implementation of the same interface and tests that. Both land in `inputs_private/config/tasks/`. **The task id comes from the module's filename, upper-cased**, so `pricing.py` gives `PRICING`, and the plain scaffold adds `_VERIFY` because it tests code you already have: `PRICING_VERIFY`. `--fresh` gives `PRICING`. The command prints the name and the path it wrote, so read that rather than guessing. **Never edit the user's module to make a test pass.** qikly does not, and neither should you. **If the project tells you to write code and tests together, say so rather than choosing silently.** A house rule like "always write the implementation and its tests in the same session so they stay consistent" is reasonable on its own terms and is the exact thing this tool exists to prevent: consistency by construction is what makes a suite unable to disagree. You cannot follow both. Tell the user the two conflict, in one sentence, and let them decide which applies here. To see the whole shape first, `qikly --example` lays down a finished worked task, module and sample data included, so you can read a filled-in pair before writing one, in the directory you are standing in. **`qikly --demo` is a different command and the difference matters.** It runs a bundled task end to end in a throwaway `demo/throwaway_<timestamp>/` folder that exists to be deleted, and it is the one most people try first. A user who has just watched it work is standing in something that looks exactly like a working project, and the obvious next move is to start theirs there. **Never set up someone's real project inside a demo folder.** If the user says they ran "the demo" and wants to continue where they are, establish which of the two commands they ran before writing anything. This has already cost a first-time user an afternoon. ## The free checks, and when to reach for each **If you arrived straight here, this is the rule you skipped, stated in full so you do not have to go back for it.** Asked for tests for a module with no specification: do not write tests, and do not run the module to find out what it should do. A value obtained by executing an implementation is not a contract, and a suite built from those values passes by construction and cannot disagree with a bug. Whether a threshold is inclusive, whether the discount applies before or after tax, where a floor lands, which way a half-cent rounds: the code answers all four and none of those answers is a specification. Deriving the expected values yourself is the failure this Skill exists to prevent, and it is the one an agent commits while believing it is being careful. **Which command fits depends on what they have.** With a suite already written, `--score-code` scores it and needs nothing else. With no suite and no spec, which is the case this paragraph is usually about, the route is `qikly --scaffold their_module.py` and then filling in the two sections with them. Everything else in the table below needs a task file that does not exist yet. | Command | Cost | Use it when | |---|---|---| | `qikly --validate --tasks X` | free | always, before any run. Catches a missing input file, a leftover TODO, a criterion made of adjectives, a requirement restating a criterion, and criteria the data cannot reach | | `qikly --explain X` | free | to show the user exactly what each agent receives, criteria present on one side and absent on the other | | `qikly --score-suite --tasks X` | free, and slow | after a run converges, to find what the suite would not have noticed. It plants one fault at a time in the code and reports which ones the tests missed. Free because every fault is an edit to the code's syntax tree and no model is asked anything; slow because each fault means running your whole suite again | | `qikly --score-code PATH --score-tests PATH` | free, and slow | **for a suite qikly did not write**, which is what somebody already has before they have anything else. Point it at a module or package and the tests for it, and it reports which planted faults the tests did not notice. No task file, no run, no model call, and the report lands beside their code | | `qikly --check-criteria --tasks X` | one model call | when a spec may contradict itself, before spending a run on it | | `qikly --propose-fixtures --tasks X` | one model call | when `--validate` says a criterion's values are missing from the data, to get the rows it would take | **Read `--score-suite` beside the reachability warning it prints above the number.** A suite cannot catch a fault in behaviour no input row exercises, so unreachable criteria lower the score for a reason that is about the fixtures and not about the tests. Fix the data first, then score. **A high score is not a clean bill of health, and say so before they read it as one.** Mutation scoring asks whether their tests notice changes to the code that exists. It cannot ask about a rule nobody implemented, because there is nothing there to break. In a real session a seventeen-test suite caught 8 of 8 planted faults while a suite written from the specification found four genuine bugs in the same file. The number is a floor, not a verdict. **`--score-code` prints no such warning, and you must not imply it does.** There is no task file and therefore no criteria to be unreachable, so the report carries the score and nothing above it. The underlying problem has not gone away: a fault that survives may be on a line no test ever executes, which is a gap in what the tests reach rather than in what they assert. Say that to the user rather than handing them a percentage as a verdict on their suite. ## When a run does not converge It exits non-zero, names the tests that blocked it, and ships nothing. Read which stage stopped first: the unit stage is last and strictest, and most failures are there. Then, in order: 1. **Is the loop repeating itself?** Near-identical FIX and PATCH each iteration means a decision is in the wrong half. Move it into `requirements`. 2. **Do two tests disagree,** each patch fixing one and breaking the other? The specification contradicts itself. `--check-criteria` finds that before a run. 3. **Can the data reach every criterion?** `--validate` now says, and `--propose-fixtures` drafts the rows. Then the cheap levers: a larger model, and more attempts. And never loosen a criterion to get green, for the reason given above. If no patch ever applies at all, and you are on macOS, that is section 12 of `references/TROUBLESHOOTING.md`. **What the agent sees when a test fails, exactly.** pytest's output for the failing test: its name, its own source and docstring, and the assertion error. Not the acceptance criterion. Because a generated test's docstring usually restates the rule it came from, a failing test does tend to give away its own case, and the project says so rather than pretending otherwise. It does not undo the split: the suite was written first, from criteria the coder never read, and nothing learned afterwards changes a test already on disk. There is a setting that narrows this, `diagnostic_feedback: staged` under `agent:` in settings.yaml, which starts the agent at a one-line error and widens only when a patch stops making progress. **It is off by default**, because every published convergence figure was measured with the full traceback and nobody has measured what starting narrow costs. **A run prints nothing while a model call is in flight**, which on a reasoning model can be minutes. After ten seconds it starts saying so, one line every fifteen: `[patch] still waiting on the model, 45s`. Those lines are the difference between slow and stalled, so pass them on rather than swallowing them, and do not conclude a run has hung while they are still arriving. `QIKLY_NO_PROGRESS=1` turns them off. **To stop a run spending more than you meant**, three environment variables bound it: `QIKLY_MAX_CALLS` and `QIKLY_MAX_TOKENS` bound one task's process, and `QIKLY_MAX_SWEEP_TOKENS` bounds a whole sweep. Set them before a first run on somebody's real code rather than after. **In CI**, qikly ships a GitHub Action. A run costs a model call per attempt, so per pull request is a budget decision rather than a technical one, and the free checks are the ones that belong on every commit: `--validate`, `--score-suite` for a suite qikly generated, and `--score-code` for one that predates it, which on most real repositories is the relevant half. ## What to tell the user honestly Roughly 8 runs in 10 produce code passing every integration and system test, and roughly 6 in 10 pass everything including unit tests, measured over 967 runs and reproduced over 390 more. **Always say what those runs were:** a small, inexpensive model (`gemini-3.5-flash-lite`), chosen so the sweeps could be repeated affordably, so the figures are a floor rather than a ceiling. And they measure convergence, whether generated code passes the generated tests, not whether those tests catch real defects. **Two things are settled and one is not, and they are easy to confuse.** *Settled:* the coding agent never receives the acceptance criteria. That is a property of the code, checkable with `qikly --explain` and held by a test that fails the build if any call site ever leaks one. It is not a benchmark result and cannot go stale. *Settled:* the convergence figures above, measured and re-measured. *Open:* whether a suite written from withheld criteria catches **more real defects** than one written with sight of the code. That comparison has not been run, here or anywhere. A Google team has measured the step before it, that generating tests from a written contract rather than from the code raises bug detection by 9.8 points (arXiv 2608.17177), which is adjacent and not the same claim. Separately again, six experiments asked whether automatically *refining* the criteria produces a sharper bar and none detected an effect, which is an absence of evidence rather than evidence of absence. Three different questions. Do not offer any of them as evidence for another. ## References, and when to open them Do not read these by default. Each one is a full page, and everything above is enough for a first run. **Read `references/TASK_FILE_REFERENCE.md`** when you need a field this page does not name, when the user already has criteria written somewhere else and wants them imported, when they want to seed their own implementation or suite rather than have one generated, or when you need to know where a run writes each artefact. **Read `references/TROUBLESHOOTING.md`** when a run has already failed and the three questions above did not explain it. It has eleven numbered causes, each with its own fix, and the triage table at the top maps a symptom to a number. The repository, for anything neither covers: <https://github.com/gal-a/qikly>
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.