Claude Skill

wit

Apply WIT (Writing Is Thinking) as a human–LLM collaborative scientific reasoning skill for scientific question formulation, finding-driven research planning, next-experiment selection, Results or Discussion review, claim–evidence and reviewer stress tests, manuscript logic audit

LLM Mart · 0 points · 1 views 2 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download deltadbu-wit-skill-wit-4931ff0.zip · 100 KB

Install

skills CLI npx skills add https://github.com/deltadbu/WIT-skill/tree/main/wit
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install deltadbu-wit-skill@llmmart
Git git clone https://github.com/deltadbu/WIT-skill.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole deltadbu/wit-skill collection as a plugin from our marketplace. Git is the plain clone.

README

This directory contains the installable WIT Agent Skill. See SKILL.md for agent instructions and references/ for the full framework.

Skill manifest

WIT

Use WIT to turn writing into scientific decision-making. Answer the user's actual question and preserve its scope.

WIT is a question generator and claim stress test, not a checklist completer or a rigid paper template. Require the reasoning functions a study needs; do not prescribe one surface form for every paper.

Preserve human scientific agency

WIT has a dual objective:

Advance the research. Grow the researcher.

Do not optimize only for producing a paper or completing the task. Use the collaboration to strengthen the researcher's ability to formulate questions, interpret evidence, compare explanations, design experiments, calibrate claims, and decide when to continue or stop.

Automate labor; augment judgment.

  • Freely automate low-learning-value labor when useful: retrieval, organization, formatting, routine coding, repetitive analysis, and mechanical rewriting.
  • Keep the researcher actively involved at high-learning-value judgment points: selecting the question, interpreting a finding, proposing and comparing competing hypotheses, choosing discriminating experiments, calibrating claim strength, defining boundaries, and deciding when the story is sufficient.
  • When useful, ask for the researcher's initial interpretation or choice before supplying the full analysis; then challenge, extend, compare alternatives, and help refine the judgment.
  • Do not turn collaboration into unnecessary interrogation. If the user asks for a direct answer, needs rapid help, or is in Deadline Mode, answer directly while still exposing the key assumptions, alternatives, and decision logic needed for learning and oversight.
  • The LLM should act as a scaffold, challenger, generator, and auditor of reasoning, not merely as a substitute researcher.

Load only the material the task needs

Read one complete authoritative workflow before applying WIT:

Load supporting material only when it is relevant to the current task.

Tests: assess WIT itself

Use materials in tests/ when evaluating, stress-testing, or refining WIT itself. Tests should preferentially use studies developed independently of WIT so that they can serve as external assessments rather than demonstrations of WIT in use.

  • Read Assess-WIT-using-AlphaGo.md, or its Chinese version, for an external assessment of WIT using a study developed independently of WIT, especially when testing whether WIT can accommodate pipeline-driven Results and non-formulaic Discussion.

Case studies: illustrate WIT in practice

Use materials in case-studies/ when an example of applying WIT to a real scientific project would improve the current reasoning or explanation.

  • Read Applying-WIT-to-MSFold.md, or its Chinese version, for a real-world example of applying WIT to scientific research and writing, including representation, search, sampling, ranking, manuscript logic, and next-step decisions.

Treat these resource types differently:

  • Tests assess WIT itself.
  • Case studies illustrate how WIT is applied.
  • Do not treat a case study as independent validation of WIT.

Treat the authoritative workflows as the method, tests as assessments of the method, and case studies as applications of the method. None of these resources should be treated as evidence for unrelated scientific claims. Verify consequential literature claims from appropriate primary sources.

Select the requested mode

Use only the mode or combination needed; do not dump the full framework by default.

  • Open a question: turn a vague idea into a researchable question space.
  • Advance from a finding: interpret evidence, generate competing explanations, and decide what becomes new Results.
  • Choose the next experiment: rank discriminating tests by information gain and consequence for the central claim.
  • Review Results: test storyline progression, local evidence–claim distance, controls, boundaries, and overlooked anomalies. For method or system papers, also audit the role of the early figures: Figure 1 does not have to summarize the whole paper; it may instead introduce the problem, representation, or motivation when that orientation is needed (as in AlphaDev). What matters is that an early overview figure—often, but not necessarily, Figure 1— lets the reader grasp the central idea and high-level operation of the proposed method. For Results figures, audit the captions as part of the evidence chain: a caption should primarily state the Fact shown by the figure and may add a restrained 1-hop Opinion as the immediate take-home message; avoid 2-hop interpretation or broad abstraction that belongs in the Results text or Discussion. Also audit whether mathematical or algorithmic methods first state the basic idea in natural language and then use a minimal concrete walkthrough with actual values or a tiny input so the reader can mentally execute the method; whether a compact worked case shows how the method operates, why it succeeds, why existing methods fail on the same case, and what mechanism creates the difference; and whether benchmark gains are localized through difference-focused case/subset analysis rather than reported only as aggregate metrics.
  • Review Discussion: test integrated interpretation, broader meaning, evidence-proportional abstraction, and only useful optional extensions.
  • Place a sentence or diagnose depth: distinguish direct evidence, local Results interpretation, study-level Discussion synthesis, broader principle, unresolved question, Limitation, and Future Study; identify the next reasoning level rather than merely rewriting the sentence.
  • Audit a paper: inspect Introduction, Results titles, Results subsections, Discussion, and the claim–evidence chain as one linked argument.
  • Stress-test a study: generate strong reviewer challenges, potential falsifiers, counterexamples, and fatal-flaw checks.
  • Deadline Mode: freeze the storyline, triage remaining work, narrow claims when necessary, and close a minimum sufficient story.

Run REWRITE as a decision loop

Start from the central question, central claim, available evidence, known constraints, and the user's immediate decision.

  1. Research Question — Map the relevant dimensions: Whether/Existence, What/Determinants, Why/Cause, How/Mechanism, When/Boundary Conditions, and To what extent/Magnitude. Use them to find omissions, not to force six answers.
  2. Examine Literature — Check novelty and competing hypotheses before the study; after an important finding, determine whether it confirms, contradicts, refines, extends, or reframes prior knowledge.
  3. Work / Experiment — Keep the link Question → Test → Data → Finding. Do not recommend an experiment merely because it is conventional.
  4. Read Finding — Separate Data, Finding/Fact, and restrained 1-hop Opinion. Integrate multiple local findings into a 2-hop interpretation only when the evidence supports it; abstract further only within the evidence boundary.
  5. Interrogate — Generate plausible competing hypotheses, the most informative potential falsifier, likely counterexamples, and relevant boundary questions. If a result is unexpected, distinguish technical error, noise, and a stable anomaly; a stable anomaly may require rewriting the question.
  6. Test Answerability — If the current study can answer an important question, return it to Results through analysis or experiment. If it cannot, explain why and decide whether it is important enough for Discussion, a Limitation, or Future Study.
  7. Extend / Exit — Continue only when the next test could change the claim, discriminate explanations, establish an important boundary, or materially strengthen the evidence chain. Otherwise apply the stop rule. Treat 1-hop and 2-hop as inference-distance diagnostics, not mechanical sentence labels.

Preserve these reasoning invariants

  • Advance the research. Grow the researcher. Scientific progress and researcher growth are both objectives of the interaction.
  • Automate labor; augment judgment. Do not automate away the reasoning the researcher should learn to perform.
  • Reasoning structure is not surface prose structure. Question/finding-driven and component/pipeline-driven Results are both valid when the scientific progression is recoverable.
  • Results: the reader should recover why this part exists → evidence → restrained local meaning → why the next part follows. Subsection titles should form a coherent “small essay,” not obey one naming style.
  • Method-paper explanatory evidence: use the early figures to orient the reader. Figure 1 may introduce the problem, representation, or motivation rather than the whole solution; if so, a subsequent early overview figure should make the central idea and high-level operation clear. Do not require a fixed figure number—require the explanatory function. For a Results figure, treat the caption as part of the scientific reasoning: Fact first, optional restrained 1-hop takeaway second; do not use the caption for 2-hop interpretation or general principles. If the core is mathematical or algorithmic, first explain the basic idea in natural language, then use the smallest concrete example that exposes the mechanism—instantiate key formula terms with actual values or walk through key algorithmic steps on a tiny input—so the reader can mentally execute the method. Benchmark statistics establish whether and by how much the method works; a compact worked case should explain why the proposed method succeeds and why existing methods fail; and difference-focused case/subset analysis should reveal where the aggregate gain comes from. Prefer mechanism-revealing examples and systematic differential patterns over cherry-picked cases.
  • Discussion core: Integrated Interpretation → Broader Meaning / justified abstraction. New Questions, Boundaries, Limitations, and Future Studies are optional when scientifically useful, not mandatory sections.
  • Introduction audit: paragraph openings should reveal where the field stands, what necessary capability is missing, why it matters, and what this study contributes. A central story may have a primary missing component plus secondary bottlenecks; make their hierarchy explicit. This is a logic test, not a required paragraph count.
  • Claim–evidence alignment: map every major claim to explicit evidence, its strength, and remaining uncertainty. Add evidence, narrow the claim, or remove it when the mapping fails.
  • Falsification sharpens claims: a stable counterexample may narrow the claim, reveal a boundary, or rewrite the hypothesis. Failure to find one adds support but never proves the claim.
  • Reviewer criticism is a stress test, not automatically a limitation. Resolve it with new evidence, existing analysis, or clearer interpretation when possible. A fatal flaw requires redesign or a narrower central claim.
  • The next experiment is not necessarily the easiest. Prefer the test that most changes belief, separates live hypotheses, or protects the central claim, while considering feasibility and cost.

Shape the output to the decision

  • For a question, return the central formulation, relevant dimension map, highest-value unresolved questions, and their answerability.
  • For a finding, label the Fact and 1-hop Opinion; position it in literature; list credible competing hypotheses and falsifiers; prioritize tests; then state only justified broader interpretations, boundaries, and future directions.
  • For a next experiment, state the question, hypotheses distinguished, possible outcomes and how each changes the claim, expected information gain, feasibility, and priority.
  • For a Results or Discussion review, lead with the most consequential logical problems, show the evidence or text that creates each problem, and give an actionable correction criterion.
  • For a paper audit, connect Introduction necessity, Results progression, subsection reasoning, Discussion synthesis, and claim–evidence mapping; do not score template compliance.
  • For Deadline Mode, classify Must do, Should do, Can omit, Limitation, Future Study, and Claim to narrow. Prioritize rather than enumerate everything that could be asked. Distinguish observation, inference, uncertainty, and proposal. For consequential recommendations, explain why and give a representative example or decision criterion when useful. Never invent data, citations, manuscript content, or certainty.

Stop rule

Freeze the storyline as:

Central Question → Central Claim → 3–5 Key Findings → Broader Meaning / General Principle

Stop expanding the current study when the central question is credibly answered, major competing explanations and reviewer risks are handled to a reasonable extent, the most important boundaries have adequate evidence, and remaining questions require work outside the present scope. Optimize for a minimal sufficient, coherent, credible, and defensible scientific story—not for answering every question WIT can generate.

Files (wit-skill)
  • case-studies
    • Applying-WIT-to-MSFold-cn.md 32.5 KB
      # 应用 WIT 于 MSFold:一个回顾性案例
      
      **论文:** *Sampling in structure-token space enables accurate prediction of multiple protein conformations*  
      **审阅版本:** 用户提供的最新版 `MSFold (3)(1).pdf`  
      **主题:** 利用 ESM3 structure-token space 与 parallel tempering 进行蛋白质多构象预测  
      **WIT 角色:** 应用案例,而不是对 WIT 的独立检验。MSFold 的研究与写作过程发生在 WIT 尚未成熟定型的阶段,因此这个案例更适合回顾性展示:**WIT 如何用于真实科研项目,组织 scientific questions、解释 findings、选择高信息量的下一步、校准 claims,并形成论文的 scientific story。**
      
      ---
      
      ## 1. WIT 在这个案例中如何发挥作用
      
      把 WIT 应用于 MSFold,可以把研究推进与论文写作重新组织成一条连续的 scientific reasoning chain,而不是把 WIT 只当成文章完成后的审稿工具:
      
      > **WIT 把 MSFold 的 research decisions 与 writing decisions 连接起来:central question、experiment、finding、follow-up question、claim strength 与 paper structure 可以作为同一个 reasoning process 来处理。**
      
      用 WIT 继续推进当前 manuscript,会自然暴露出几个仍值得加强的 research / writing 问题:
      
      1. **Introduction 中“missing component = sampling”的主线很强,但 ranking 也是明确 bottleneck,需要把主次关系说得更清楚。**
      2. **Results 2.6 从有限的 representation comparison 上升到“sampling depends on representation”时,已经接近 2-hop / general principle,需要进一步限制 claim 的边界。**
      3. **SLL 的 benchmark-level improvement 相对有限,因此 “effective criterion” 的措辞应与 evidence strength 对齐。**
      4. **由于 multi-conformation success 采用 best-of-ensemble 思路,baseline comparison 是否匹配 sampling budget / ensemble size 是一个高优先级 reviewer question;当前主文正文没有把这一点交代清楚。**
      5. **Boundary-condition analysis 仍可加强:哪些 proteins / transition types 上成功或失败?这是目前信息增益很高、而且可能成本较低的分析。**
      
      因此,这个案例的目的不是证明 MSFold “把 WIT 全部打勾”,而是展示 WIT 在真实项目中的工作方式:
      
      > **Finding → Question → Competing Explanations → Discriminating Test → Interpretation → Next Step → Scientific Story**
      
      即使 story 已经很强,WIT 仍会继续追问:哪些 claim 超出了 evidence,哪些 competing explanations 尚未排除,哪些 boundary conditions 仍不清楚。
      
      ---
      
      # 2. Introduction:Final Goal → Missing Component → This Study
      
      ## 2.1 Final Goal
      
      论文首先建立了明确的最终目标:
      
      > **从单一蛋白质序列预测多个 functionally relevant conformations,而不是只预测一个 dominant structure。**
      
      这个目标不是纯粹算法任务。论文把它连接到 ligand recognition、catalysis、transport 和 allosteric regulation 等功能过程,并指出实验结构数据库只能提供 conformational equilibria 的部分视图。
      
      因此 Introduction 的第一层功能非常清楚:
      
      > **Why does multi-conformation prediction matter?**
      
      ---
      
      ## 2.2 Necessary Components
      
      文章随后把 accurate multi-conformation prediction 拆成两个必要能力:
      
      > **efficient sampling + precise selection**
      
      也就是说:
      
      1. 必须能够探索 conformational space,找到 alternative states;
      2. 找到以后,还必须从 ensemble 中识别 plausible conformations。
      
      这个拆分很好,因为它避免把“生成很多结构”误认为“解决了多构象预测”。
      
      ---
      
      ## 2.3 Established Components
      
      Previous studies 已经建立了多类组件:
      
      - MD 可以探索 dynamics,但受 sampling cost、timescale 和 force field 限制;
      - AlphaFold / RoseTTAFold 等方法擅长 dominant conformation;
      - MSA sampling、clustering、diffusion、flow matching 可以增加 diversity;
      - ESM3 提供了结构的 discrete token representation 与 decoder。
      
      特别重要的是 ESM3:
      
      > **3D structure → residue-level structure-token sequence → 3D structure**
      
      这意味着 representation 与 decoding machinery 已经存在。
      
      ---
      
      ## 2.4 Missing Component
      
      文章给出的核心判断是:
      
      > **an effective sampling strategy has remained elusive.**
      
      标准 ESM3 decoding,包括 argmax 和 temperature-controlled randomization,本质上仍偏向 local search,难以跨越 barrier 到达 alternative basins。
      
      MSFold 因此补上:
      
      > **parallel tempering in structure-token space**
      
      并同时提出 SLL 用于 ranking。
      
      这与 WIT 的 Introduction 逻辑高度一致:
      
      > **Final Goal → Necessary Components → Established Components → Missing Component → This Study**
      
      ### WIT 应用
      
      **强。**
      
      但这里存在一个值得进一步精炼的逻辑点:
      
      前一段已经明确说:
      
      > **ranking is a further challenge**
      
      随后却说:
      
      > **the missing component is effective sampling**
      
      而 MSFold 又同时提出 sampling + SLL ranking。
      
      因此最好明确层级:
      
      > **Primary missing component: effective global sampling**  
      > **Secondary unresolved bottleneck: reliable ranking / selection**
      
      这样 central story 会更集中,也能避免给读者一种“刚说有两个缺口,下一段突然只剩一个缺口”的感觉。
      
      ---
      
      # 3. 六维 Scientific Question Space
      
      WIT 将“把问题打开”定义为六个 scientific dimensions:
      
      > **Existence → Determinants → Cause → Mechanism → Boundary Conditions → Magnitude**
      
      MSFold 几乎覆盖了全部六维,但覆盖程度并不相同。
      
      ---
      
      ## 3.1 Whether → Existence
      
      问题:
      
      > **MSFold 是否真的能够恢复多个 experimentally determined conformations?**
      
      核心 benchmark:
      
      - 312 proteins;
      - 成功定义:Fold 1 和 Fold 2 都达到 TM-score ≥ 0.75;
      - MSFold:161 / 312,success rate = **0.516**;
      - AlphaFold3:0.378;
      - AlphaFold2:0.360;
      - 其他方法更低。
      
      因此 Existence 得到非常直接的回答:
      
      > **MSFold 的 multi-conformation recovery advantage 确实存在。**
      
      ---
      
      ## 3.2 What → Determinants
      
      问题:
      
      > **哪些因素决定 multi-conformation prediction 的成功?**
      
      文章实际上逐步识别出多个 determinants:
      
      - sampling strategy;
      - replica exchange;
      - temperature regime;
      - representation of the search space;
      - ranking criterion;
      - sampling budget / parameter setting。
      
      尤其:
      
      > **2.3 Sampling strategy is critical for multi-conformation prediction**
      
      直接回答了一个 determinant question。
      
      随后 2.6 又把 determinant 从 algorithm 推向 representation:
      
      > **the representation of the search space also matters.**
      
      因此 What 维度比较充分。
      
      ---
      
      ## 3.3 Why → Cause
      
      问题:
      
      > **为什么标准 prediction / decoding 方法容易漏掉 alternative conformations?**
      
      论文给出的主要 causal explanation 是:
      
      > **standard decoding and local sampling are biased toward the dominant local basin / dominant mode.**
      
      最有力的证据来自一个相对干净的 comparison:
      
      - 都使用 ESM3 structure-token representation;
      - argmax / iterative sampling 与 MSFold 的主要区别是 sampling strategy;
      - success rate 分别约 0.330 / 0.370 / 0.516。
      
      这使得:
      
      > **sampling accessibility**
      
      成为 performance gap 的一个有力 causal explanation。
      
      ---
      
      ## 3.4 How → Mechanism
      
      问题:
      
      > **parallel tempering 如何使 alternative states 变得可达?**
      
      论文不是停留在“parallel tempering works”,而是继续分析:
      
      - low-T replicas:local refinement;
      - high-T replicas:broader exploration;
      - replica exchange:将两者结合;
      - trajectory 显示 basin-to-basin transition;
      - disabling exchange 会降低 transition ability;
      - single-temperature low T 被困在 dominant basin;
      - single-temperature high T 产生 distorted structures。
      
      因此可以形成清楚的 mechanism chain:
      
      > **local-search trapping**  
      > → **temperature-separated exploration/refinement**  
      > → **replica exchange**  
      > → **barrier crossing / basin transitions**  
      > → **alternative-state recovery**
      
      这是 Why 与 How 区分得非常好的例子:
      
      > **Why:问题来自 local-search bias。**  
      > **How:parallel tempering 通过 high-T exploration + low-T refinement + exchange 解决它。**
      
      ---
      
      ## 3.5 When → Boundary Conditions
      
      这一维度已经有一些重要 evidence:
      
      - diverse conformational changes;
      - temporal holdout;
      - pre-cutoff / post-cutoff structures;
      - 多个 case studies;
      - high / intermediate / low temperature;
      - parameter sensitivity。
      
      特别是 temporal evaluation:
      
      > 在 ESM3 training cutoff 之后发布结构的 32 proteins 上,MSFold success rate = **0.594**。
      
      这说明 performance improvement 不容易被“结构记忆”简单解释。
      
      但是 WIT 会继续问:
      
      > **什么时候 MSFold 成功,什么时候失败?**
      
      目前仍缺少更系统的 boundary map,例如:
      
      - apo–holo、fold switching、secondary-structure change 分别如何?
      - small vs large proteins?
      - single-domain vs multi-domain?
      - conformational distance 增大时 success rate 如何变化?
      - ESM3 对某区域 token entropy 很低时是否更容易失败?
      - 需要多少 replicas / steps 后性能开始饱和?
      
      这类分析很可能具有:
      
      > **low experimental cost + high information gain**
      
      因此是 WIT 会优先建议的下一步分析之一。
      
      ---
      
      ## 3.6 To what extent → Magnitude
      
      论文提供了多个层面的 magnitude:
      
      - overall success:0.516;
      - Fold 1 average TM-score:0.821;
      - Fold 2 average TM-score:0.740;
      - temporal holdout success:0.594;
      - LAOBP 两种状态:TM-score 约 0.99 / 0.98;
      - SLL / pTM / pLDDT 的 correlation 和 ranking comparison。
      
      所以文章并没有停留在:
      
      > **“有效。”**
      
      而是继续回答:
      
      > **“有效到什么程度?主要改善哪里?”**
      
      特别是 Fold 1 / Fold 2 decomposition 很有价值,因为它揭示:
      
      > **gain 主要来自 alternative-state recovery,而不是通过牺牲 dominant-state accuracy 换来的。**
      
      ---
      
      # 4. Results Titles:是否能组成一篇“小 Essay”?
      
      最新版 Results titles 是:
      
      1. **MSFold recovers multiple protein conformations across diverse proteins**
      2. **MSFold generalizes beyond known structures**
      3. **Sampling strategy is critical for multi-conformation prediction**
      4. **MSFold traverses conformational landscapes to access alternative states**
      5. **Parallel tempering enables exploration beyond local basins**
      6. **The discrete structure-token space reshapes conformational sampling**
      7. **Sequence log-likelihood improves selection of plausible conformations**
      
      把标题连起来,可以还原成一条非常清楚的问题链:
      
      > **Does it work?**  
      > → **Does it generalize beyond memorization?**  
      > → **What causes the gain?**  
      > → **What does successful exploration look like?**  
      > → **How does parallel tempering create transitions?**  
      > → **Why does representation matter?**  
      > → **After exploration, how do we select useful outputs?**
      
      ### WIT 应用
      
      **非常强。**
      
      这是典型的 **Finding-driven Results**。
      
      对 MSFold 而言,WIT 在这里的实际作用是:
      
      > **Results should expose the logical progression of the scientific story.**
      
      MSFold 采用的是 finding-driven progression。WIT 并不要求固定标题格式,而是要求读者能够从 Results 的顺序恢复出科学推理如何一步步推进。
      
      ---
      
      # 5. Results Subsection:Fact → Restrained 1-hop Opinion
      
      ## 5.1 Section 2.1:Performance
      
      ### Fact
      
      - MSFold success rate = 0.516;
      - Fold 1 accuracy 与 AlphaFold3 / AlphaFold2 接近;
      - Fold 2 accuracy 明显更高。
      
      ### 1-hop Opinion
      
      文章总结:
      
      > MSFold 的主要 improvement 来自 enhanced Fold 2 recovery,同时保持 Fold 1 accuracy。
      
      这是非常典型的:
      
      > **Fact → 1-hop Opinion**
      
      没有过早上升到“representation + search”的 general principle。
      
      ### WIT 应用
      
      **非常好。**
      
      ---
      
      ## 5.2 Section 2.2:Generalization vs Memorization
      
      ### Question
      
      > Improvement 是 genuine generalization,还是 training-data exposure / memorization?
      
      ### Test
      
      > post-ESM3-cutoff temporal subset。
      
      ### Fact
      
      > MSFold success rate = 0.594,与 full benchmark 0.516 相当,并继续优于多数方法。
      
      ### 1-hop Opinion
      
      > advantage is unlikely to be explained by simple memorization.
      
      这一措辞是合适的,因为它说的是:
      
      > **unlikely**
      
      而不是:
      
      > **memorization has been completely ruled out.**
      
      ### 一个 borderline inference
      
      第二个 temporal analysis 中,AlphaFold2 即使在两种 endpoint structures 都早于 cutoff 时,仍低于 MSFold。文章随后说:
      
      > the major limitation likely lies in the prediction objective and inference procedure.
      
      这个 conclusion 比前面的 Fact 多走了一步。
      
      `likely` 已经起到 restraint 的作用,但从 WIT 看,它仍接近:
      
      > **1.5-hop**
      
      因为“不是 exposure 就是 objective/inference”并没有完全排除其他 competing explanations。
      
      更稳妥的理解是:
      
      > **prior exposure alone is insufficient; objective/inference remain plausible major contributors.**
      
      ---
      
      # 6. Section 2.3:Competing Hypotheses 与 Discriminating Experiment
      
      这是整篇文章里很符合 WIT 的一节。
      
      可能解释:
      
      > **H1:performance gain 来自 ESM3 的 representation。**  
      > **H2:performance gain 来自 better sampling。**
      
      实验设计:
      
      > 固定 ESM3 structure-token representation,只改变 sampling / decoding strategy。
      
      结果:
      
      - argmax:0.330;
      - iterative sampling:0.370;
      - MSFold:0.516。
      
      因此:
      
      > **sampling strategy is a critical determinant**
      
      得到比较强的 discrimination。
      
      ### WIT 应用
      
      **这一节比普通 baseline comparison 更有科学价值。**
      
      因为它不是问:
      
      > “谁分数高?”
      
      而是问:
      
      > **“是什么造成分数差异?”**
      
      这正是:
      
      > **Competing Hypotheses → Discriminating Experiment**
      
      ---
      
      # 7. Sections 2.4–2.5:从 Finding 到 Mechanism
      
      ## 7.1 Section 2.4:What happens during successful sampling?
      
      LAOBP trajectory 显示:
      
      - local refinement periods;
      - Fold 1 / Fold 2 similarity curves crossing;
      - dominant ↔ alternative basin transitions;
      - UMAP ensemble 出现两个主要 clusters;
      - additional case studies 显示类似 behavior。
      
      这一节首先建立:
      
      > **MSFold 不只是生成两个 endpoint,而是在 sampling trajectory 中访问不同 basins。**
      
      它主要回答:
      
      > **What does the process look like?**
      
      ---
      
      ## 7.2 Section 2.5:Why / How do transitions happen?
      
      接下来才真正问:
      
      > **什么 mechanism 支持 basin transitions?**
      
      证据:
      
      - high-T replica 更 diverse;
      - low-T replica 更 concentrated;
      - disable exchange 后 transition ability 下降;
      - single-temperature low T 陷入 dominant basin;
      - high T 产生 distorted conformations。
      
      因此形成:
      
      > **high-T exploration + low-T refinement + replica exchange**
      
      这一 mechanism explanation 比单纯 trajectory visualization 更有力。
      
      ### WIT 应用
      
      2.4 与 2.5 看似接近,其实 reasoning function 不同:
      
      > **2.4 = phenomenon / process observation**  
      > **2.5 = mechanism test**
      
      因此目前分成两节是合理的,而不是重复。
      
      ---
      
      # 8. Section 2.6:Representation 的结论是否走得太远?
      
      这一节很重要,因为它推动文章从:
      
      > **sampling algorithm matters**
      
      进一步走向:
      
      > **representation + sampling jointly matter**
      
      证据包括:
      
      - 单次 token update 可改变很大比例 residues;
      - large-scale token changes 对应 substantial conformational moves;
      - 同一个 structural basin 可以对应不同 token sequences;
      - 10-ns Cartesian-space MD 中,多数测试 proteins 没有发生 state transition。
      
      文章由此写道:
      
      > **the success of sampling algorithms depends not only on the algorithm itself, but also on the representation of the search space**
      
      这是一个非常漂亮的 principle。
      
      但从 WIT 的 inference-distance 看,它也是全文最需要小心的一处。
      
      原因是:
      
      > **局部实验是在 ESM3 discrete token sampling 与有限时间 Cartesian MD 之间进行的。**
      
      而 conclusion 已经接近:
      
      > **sampling algorithms in general depend on representation**
      
      这已经从 subsection-level 1-hop 往 Discussion-level abstraction 走。
      
      ### WIT 建议
      
      Results 中可以更 restrained:
      
      > **Within the tested settings, the discrete structure-token representation changes the scale and connectivity of accessible conformational moves, facilitating broader exploration.**
      
      然后在 Discussion 中再上升为:
      
      > **conformational prediction is a joint problem of representation and search.**
      
      这样:
      
      > **Fact → 1-hop → 2-hop / General Principle**
      
      层级会更干净。
      
      最新版 Discussion 已经对 10-ns MD 的解释加了 boundary:
      
      > 只能视为 tested conditions 下的 sampling accessibility difference,而不能作为 MD 一般性 limitation 的证据。
      
      这一点非常重要,也符合 WIT。
      
      ---
      
      # 9. Section 2.7:SLL 的 Claim–Evidence Alignment
      
      SLL 的 scientific motivation 很自然:
      
      > **broader exploration 成功以后,下一个 bottleneck 是 selection。**
      
      这很好地体现:
      
      > **Finding → New Question**
      
      SLL 在 LAOBP case 上:
      
      - Spearman ρ = 0.71;
      - pTM = 0.65;
      - pLDDT = 0.68;
      - 两个高质量 structures 被 SLL 排进 top 1%,而 pTM / pLDDT 排名较差。
      
      但在 312-protein benchmark 上,top-ranked average TM-score 的 improvement 相对 modest:
      
      - Fold 1:SLL 0.720 vs pTM 0.715 vs pLDDT 0.697;
      - Fold 2:SLL 0.611 vs pTM 0.607 vs pLDDT 0.593。
      
      因此:
      
      > **“SLL improves selection” 有 evidence。**
      
      但:
      
      > **“SLL solves ranking” 没有 evidence。**
      
      当前 Discussion 也明确承认:
      
      > **reliable selection of plausible alternative conformations remains unresolved.**
      
      ### WIT 应用
      
      整体 claim 已经比较克制。
      
      但 Results 最后的:
      
      > **SLL provides an effective criterion**
      
      可以考虑进一步校准为:
      
      > **SLL provides a complementary reference-free criterion and modestly improves ranking on the benchmark, with larger gains in selected challenging cases.**
      
      这会使 claim strength 与 evidence strength 更贴合。
      
      ---
      
      # 10. Discussion:应用 WIT 的 Core Functions
      
      WIT 修正后的 Discussion 不要求固定六段式,而关注两个 core functions:
      
      > **Integrated Interpretation → Broader Meaning / justified abstraction**
      
      MSFold 在这一点上做得很好。
      
      ---
      
      ## 10.1 Integrated Interpretation
      
      第一段综合了:
      
      - benchmark improvement;
      - Fold 2 recovery;
      - temporal holdout;
      - standard decoding comparison。
      
      然后得到:
      
      > pretrained model 的 multi-conformation capability 不只取决于“学到了什么”,还取决于 inference 时“如何访问这些信息”。
      
      这明显不是重复 Results,而是:
      
      > **multiple local findings → integrated 2-hop interpretation**
      
      ### 一个需要注意的 inference
      
      文章说:
      
      > pretrained model already encodes structural information relevant to multiple conformations.
      
      这不是被直接观测到的 Fact,而是由:
      
      - same representation;
      - different sampling;
      - alternative-state recovery;
      - post-cutoff generalization
      
      共同支持的 interpretation。
      
      这正适合放在 Discussion,而不是 Results。
      
      ---
      
      ## 10.2 Broader Meaning / General Principle
      
      第二段提出:
      
      > **conformational prediction should be viewed as a joint problem of representation and search.**
      
      这个 principle 有多组 Results 支撑:
      
      - 2.3:sampling matters;
      - 2.5:search mechanism matters;
      - 2.6:representation matters。
      
      因此它不是空泛拔高。
      
      这是很好的:
      
      > **multiple 1-hop findings → 2-hop interpretation → abstraction**
      
      ---
      
      ## 10.3 Optional Extensions:Limitations / Future Studies
      
      对 MSFold 而言,显式 limitations 很有必要,因为几个 unresolved bottlenecks 本身就是 scientific story 的重要组成部分。
      
      当前 Discussion 列出:
      
      - ESM3 encoded information 的上限;
      - quantization / decoding error;
      - computation cost;
      - 10-ns MD timescale boundary;
      - ranking remains unresolved。
      
      这些 limitations 大部分都有明确 scientific origin,不是模板化的:
      
      > **Finding → unresolved question → current constraint**
      
      尤其 ranking:
      
      > broader exploration → more conformations → harder selection
      
      这是从 Results 自然生长出来的 limitation。
      
      因此 Future Study:
      
      > better exploration + better selection
      
      也不是愿望清单。
      
      ### WIT 应用
      
      **这是很好的 optional-extension 使用场景。**
      
      WIT 并不把 limitations / future work 当作 mandatory surface sections;但在 MSFold 中:
      
      > **当 unresolved bottleneck 对 central story 很重要时,就应该明确写。**
      
      ---
      
      # 11. 用 WIT 做 Falsification / Counterexample Check
      
      对 MSFold 的 central claims,可以主动问:
      
      ## 11.1 “MSFold improvement 只是 memorization”
      
      潜在 falsifier:
      
      > post-cutoff proteins 上 improvement 消失。
      
      实际结果:
      
      > 没消失。
      
      因此这个 competing explanation 被明显削弱。
      
      ---
      
      ## 11.2 “只要 ESM3 representation 好,sampling 无所谓”
      
      潜在 falsifier:
      
      > 同一 representation 下 argmax / iterative / MSFold 表现接近。
      
      实际结果:
      
      > 差异明显。
      
      因此 sampling claim strengthened。
      
      ---
      
      ## 11.3 “只要提高 temperature 就能找到 alternative states”
      
      潜在 falsifier:
      
      > single high-T sampling 与 MSFold 同样有效。
      
      实际观察:
      
      > high T 产生大量 distorted structures;low T 被困在 dominant basin。
      
      因此:
      
      > **high temperature alone is insufficient.**
      
      支持 replica-exchange mechanism。
      
      ---
      
      ## 11.4 “replica exchange 并不重要”
      
      潜在 falsifier:
      
      > disable exchange 后 behavior 不变。
      
      实际:
      
      > transitions 与 exploration ability 下降。
      
      因此 exchange 的 mechanism claim strengthened。
      
      ---
      
      ## 11.5 “SLL 是可靠的 ranking solution”
      
      这里反而出现一个有价值的“软 counterexample”:
      
      > benchmark-level improvement 相对 modest,而且 Discussion 自己承认 reliable ranking remains unresolved。
      
      因此 WIT 的处理不是“删除 SLL”,而是:
      
      > **Counterexample / weak gain → narrow the claim → better claim**
      
      即:
      
      > SLL 是一个 useful / complementary reference-free ranking signal,而不是 ranking problem 的完整解决方案。
      
      这正好体现:
      
      > **A successful falsification sharpens the claim.**
      
      ---
      
      # 12. Claim–Evidence Mapping
      
      | Major Claim | Main Evidence | WIT 判断 |
      |---|---|---|
      | MSFold improves multi-conformation recovery | 312-protein benchmark;0.516 success rate | 强支持 |
      | Gain mainly comes from Fold 2 recovery | Fold 1 / Fold 2 decomposition | 强支持 |
      | Improvement is not simple memorization | post-ESM3-cutoff temporal subset | 有力支持,但不是对所有 memorization mechanisms 的绝对排除 |
      | Sampling is a critical determinant | same ESM3 representation + different sampling strategies | 强 discriminating evidence |
      | Parallel tempering enables basin transitions | trajectories + exchange ablation + single-T comparison | 强机制支持 |
      | Representation reshapes accessible conformational moves | token-update analysis + MD comparison | 支持,但 generalization 必须受 tested conditions 约束 |
      | SLL improves ranking | case study + 312-protein benchmark | 有支持,但整体 improvement modest |
      | Pretrained PLM encodes more conformational information than standard decoding reveals | synthesis of sampling and temporal evidence | 合理 2-hop interpretation,不是直接 observation |
      | Prediction depends jointly on representation and search | synthesis of 2.3 / 2.5 / 2.6 | 合理 General Principle |
      
      ---
      
      # 13. Reviewer Stress Test
      
      WIT 要求投稿前问:
      
      > **最强 reviewer 会攻击哪三件事?**
      
      对当前版本,我认为至少有以下四个高价值问题。
      
      ## 13.1 Sampling Budget / Ensemble Size 是否公平?
      
      MSFold 生成:
      
      > **40 replicas × 500 steps = 20,000 conformations**
      
      而 success metric 使用 sampled ensemble 中的 best matching conformations。
      
      因此 reviewer 很自然会问:
      
      > **baseline 是否使用可比的 sample count / compute budget?**
      
      如果某些 baseline 只产生远少于 20,000 structures,那么:
      
      > **best-of-N 本身会带来优势。**
      
      当前主文正文没有把这个 control 讲清楚。
      
      如果 Supplementary 已严格控制:
      
      > **建议在主文中明确指出。**
      
      如果没有:
      
      > **这是高优先级 discriminating control。**
      
      因为它可能直接改变 central performance claim 的解释。
      
      ---
      
      ## 13.2 Representation vs MD Comparison 是否足够公平?
      
      当前 10-ns MD comparison 很容易被 reviewer 问:
      
      > **短时间 MD 没 transition,是否只是 timescale 不够?**
      
      Discussion 已经正确收缩 claim,这是优点。
      
      但如果要进一步支持:
      
      > **representation reshapes sampling**
      
      最好寻找更 matched 的 comparison,而不是把主要证据压在短 MD 上。
      
      ---
      
      ## 13.3 Boundary Conditions 在哪里?
      
      312-protein overall score 很强,但 reviewer 可能继续问:
      
      > **哪些 proteins 成功?哪些失败?为什么?**
      
      如果现有数据允许,建议做:
      
      - performance vs conformational-change type;
      - performance vs protein size;
      - performance vs Fold1–Fold2 structural distance;
      - performance vs token entropy / diversity;
      - performance vs domain architecture。
      
      这是很典型的:
      
      > **When → Boundary Conditions**
      
      而且很可能是低成本、高 information gain 的 analysis。
      
      ---
      
      ## 13.4 Ranking Claim 是否过强?
      
      SLL 有优势,但 benchmark average gain 不大。
      
      因此:
      
      > **“improves” 没问题;“solves / reliable / substantially better” 要谨慎。**
      
      这不是 fatal flaw,而是 claim calibration。
      
      ---
      
      # 14. Information Gain:用 WIT 选择下一步
      
      如果只根据 WIT 选择下一步,而不是“还能做什么”,优先级可以这样排。
      
      ## Priority 1:Matched Sampling-Budget Control
      
      原因:
      
      > 它可能改变 central benchmark advantage 的解释。
      
      这是最高 information gain。
      
      ---
      
      ## Priority 2:Failure / Boundary Analysis
      
      原因:
      
      > overall success 已经比较充分,再加一个类似 benchmark 信息增益不大。
      
      相比之下:
      
      > **什么时候有效 / 什么时候失效**
      
      会显著增加 scientific understanding。
      
      ---
      
      ## Priority 3:更强的 Representation Control
      
      目标:
      
      > 区分“parallel tempering algorithm 本身”与“discrete representation 改变 move topology”的贡献。
      
      这是对 General Principle:
      
      > **representation + search**
      
      最直接的进一步检验。
      
      ---
      
      ## Priority 4:Ranking Evaluation Reframing
      
      与其继续堆 correlation,可以优先问:
      
      > **在实际 unknown-native scenario 中,SLL 能把 high-quality alternative state 富集到 top-k 到什么程度?**
      
      例如:
      
      - hit rate@k;
      - enrichment factor;
      - probability of recovering both states after top-k selection;
      - calibration。
      
      这些 metric 可能比平均 TM-score 更贴近真正的 selection question。
      
      ---
      
      # 15. 用 WIT 冻结 Research Storyline
      
      按照 WIT,可以把最新版 MSFold 的主线冻结为:
      
      ## Central Question
      
      > **A pretrained protein language model may contain information about multiple conformations; can a better inference-time search strategy access that latent capability?**
      
      ## Central Claim
      
      > **Parallel tempering in ESM3's discrete structure-token space improves access to alternative conformational states while preserving dominant-state accuracy.**
      
      ## Key Findings
      
      1. MSFold achieves a 51.6% multi-conformation success rate on 312 proteins;
      2. improvement mainly comes from Fold 2 recovery;
      3. improvement persists beyond the ESM3 temporal cutoff;
      4. sampling strategy is a critical determinant;
      5. replica exchange couples exploration and refinement to enable basin transitions;
      6. discrete token representation changes accessible structural moves;
      7. broader generation exposes ranking as a remaining bottleneck.
      
      ## General Principle
      
      > **A model's usable predictive capability depends not only on what its representation encodes, but also on how the learned space is searched and how sampled outputs are selected.**
      
      有了这条 frozen storyline 后,对任何新 experiment 都应该问:
      
      > **它会改变 central claim、排除重要 competing hypothesis、明确 boundary,还是只是再增加 supporting example?**
      
      ---
      
      ---
      
      # 16. WIT 是如何应用于 MSFold 的?
      
      MSFold 更适合作为 **WIT 的应用案例**,而不是证明 WIT 有效的独立证据。MSFold 与 WIT 在一定程度上是并行生长的:当时我们已经在使用其中一些 reasoning patterns,只是还没有把它们系统地命名、组织成今天的 WIT。
      
      回顾这个过程,可以看到 WIT 至少在五个方面发挥作用。
      
      ## 16.1 从 Finding 打开 New Questions
      
      MSFold 的 Results 很明显沿着如下问题链生长:
      
      > works
      > → generalizes?
      > → why?
      > → how?
      > → what role does representation play?
      > → how to rank?
      
      这对应 WIT 的核心循环:
      
      > **Finding → New Question → Answerability Check → New Analysis / Experiment → New Finding**
      
      关键不是要求每个项目都把所有问题做一遍,而是避免把一个有价值的 finding 当作 reasoning 的终点。
      
      ## 16.2 区分 Evidence 与 Inference Depth
      
      MSFold 多处需要区分:
      
      > **Fact → restrained 1-hop interpretation → integrated 2-hop interpretation → broader abstraction**
      
      例如:
      
      - benchmark data 支持 Fold 2 recovery 的提高;
      - same-representation sampling comparison 支持 sampling 的重要性;
      - mechanism experiments 支持 replica exchange 的作用;
      - representation analysis 支持关于 accessible moves 的有边界 claim;
      - Discussion 再把这些 findings 综合成更高层的 **representation + search** principle。
      
      这不仅帮助决定 **写什么**,也帮助决定一个 claim **应该放在哪里、应该说多强**。
      
      ## 16.3 把 Alternative Explanations 变成实验
      
      MSFold 中多个问题都可以写成 competing hypotheses:
      
      - memorization vs. generalization;
      - representation vs. sampling;
      - high temperature alone vs. temperature coupling with replica exchange。
      
      WIT 不满足于把它们写进 Discussion,而进一步问:
      
      > **什么 experiment 最能区分这些 plausible explanations?**
      
      temporal holdout、same-representation sampling comparison、exchange ablation 都属于 discrimination power 很高的实验。
      
      ## 16.4 用 Information Gain 选择下一步
      
      当 central benchmark 已经比较充分以后,WIT 不鼓励因为“容易做”或“大家都做”就继续增加类似实验,而是问:剩余哪一个 uncertainty 最重要?
      
      对当前 MSFold,优先级自然变成:
      
      1. matched sampling-budget / ensemble-size control;
      2. systematic failure / boundary-condition analysis;
      3. 更强的 representation-effect isolation;
      4. task-oriented ranking evaluation。
      
      这些实验的价值在于:它们可能改变或显著 sharpen central claim 的解释,而不仅仅再增加一个 supporting example。
      
      ## 16.5 把 Research 与 Writing 连起来
      
      WIT 不把 paper structure 看成研究结束后的包装,而把它看成 research logic 的外显形式。MSFold 的 Results 可以压缩成:
      
      > **Performance → Generalization → Cause → Process → Mechanism → Representation → Selection**
      
      Discussion 再将这些 local findings 综合为 broader interpretation。
      
      这正是 **Writing Is Thinking** 的含义:当你认真决定 Results 和 Discussion 应该怎么写时,会反过来暴露 underlying reasoning chain 是否完整、是否跳跃、是否缺少关键 link。
      
      ---
      
      # 17. MSFold 案例的 Take-Home Message
      
      MSFold 展示的是 WIT 作为 **human–LLM collaborative scientific research and writing** 工具的一种用法。它不是让 LLM 自动把科研做完、研究者喝咖啡等 paper,而是让 LLM 去扩展、挑战、组织和审计研究者的思考,同时让研究者继续承担最关键的 scientific judgments。
      
      在这个案例里,WIT 把四件事连成一条链:
      
      > **打开 scientific question space**
      > → **把 finding 变成有判别力的 next question**
      > → **让 claim 与 evidence 对齐**
      > → **把 reasoning chain 组织成 coherent scientific story**
      
      因此,最值得做的下一步不是简单的“更多实验”,而是最大程度减少 central scientific interpretation 不确定性的实验。
      
      > **Advance the research. Grow the researcher.**
      
      ---
      
      ## Source
      
      本案例由 WIT GitHub 仓库当前版本的 `检验WIT-using-MSFold.md` 重构而来;该文件基于 manuscript:
      
      > *Sampling in structure-token space enables accurate prediction of multiple protein conformations*
      
      审阅文件:`MSFold (3)(1).pdf`。
      
      本文对 manuscript 中仍保留的 `XXX`、待补 figure / supplementary references 等 draft placeholders 不作为 scientific reasoning 缺陷处理。
      
    • Applying-WIT-to-MSFold.md 34.2 KB
      # Applying WIT to MSFold: A Retrospective Case Study
      
      **Paper:** *Sampling in structure-token space enables accurate prediction of multiple protein conformations*  
      **Reviewed version:** latest manuscript provided by the user, `MSFold (3)(1).pdf`  
      **Topic:** multi-conformation protein prediction using ESM3 structure-token space and parallel tempering  
      **Role in WIT:** Application case study. MSFold was developed and written while WIT itself was still taking shape, so this is not an independent test of WIT. Instead, it retrospectively shows how WIT can be used in a real research project to organize scientific questions, interpret findings, choose informative next steps, calibrate claims, and shape the paper's scientific story.
      
      ---
      
      ## 1. How WIT Is Applied in This Case
      
      Applying WIT to MSFold provides a way to reconstruct and continue the project as a sequence of scientific reasoning moves rather than as a sequence of writing edits:
      
      > **WIT helps connect the project's research decisions and writing decisions: the central question, experiments, findings, follow-up questions, claim strength, and paper structure can be treated as one continuous reasoning process.**
      
      Using WIT on the current manuscript highlights several concrete places where the research or writing can still be strengthened:
      
      1. **The Introduction has a strong “missing component = sampling” storyline, but ranking is also presented as a real bottleneck; the hierarchy between the primary and secondary bottlenecks should be made clearer.**
      2. **In Results 2.6, the manuscript moves from a limited representation comparison toward the broader statement that sampling depends on representation; this is already close to a 2-hop / general-principle claim and should be carefully bounded.**
      3. **The benchmark-level gain of SLL is relatively modest, so the wording around an “effective criterion” should remain aligned with evidence strength.**
      4. **Because multi-conformation success is based on the best structures in an ensemble, whether baseline comparisons use matched sampling budgets / ensemble sizes is a high-priority reviewer question; this is not made clear in the main text.**
      5. **Boundary-condition analysis could be strengthened: on which proteins and transition types does MSFold succeed or fail? This may be a relatively low-cost, high-information-gain analysis.**
      
      The purpose of this case is therefore not to show that MSFold “checks every WIT box.” Instead, it illustrates how WIT can be used as a working partner throughout a real project:
      
      > **Finding → Question → Competing Explanations → Discriminating Test → Interpretation → Next Step → Scientific Story**
      
      Even when the story is already strong, WIT keeps asking where claims outrun evidence, which explanations remain viable, and which boundary conditions are still unknown.
      
      
      ---
      
      # 2. Introduction: Final Goal → Missing Component → This Study
      
      ## 2.1 Final Goal
      
      The manuscript establishes a clear final goal:
      
      > **Predict multiple functionally relevant conformations from a single protein sequence rather than only one dominant structure.**
      
      This is not presented merely as an algorithmic task. It is connected to ligand recognition, catalysis, transport, and allosteric regulation, while current experimental structural repositories provide only a partial view of conformational equilibria.
      
      The first Introduction function is therefore clear:
      
      > **Why does multi-conformation prediction matter?**
      
      ---
      
      ## 2.2 Necessary Components
      
      The manuscript then decomposes accurate multi-conformation prediction into two required capabilities:
      
      > **efficient sampling + precise selection**
      
      That is:
      
      1. the conformational space must be explored sufficiently to discover alternative states;
      2. plausible conformations must then be identified from the sampled ensemble.
      
      This is an important decomposition because it prevents “generating many structures” from being equated with “solving multi-conformation prediction.”
      
      ---
      
      ## 2.3 Established Components
      
      Previous work has already established several relevant components:
      
      - MD can explore dynamics but is limited by sampling cost, accessible timescales, and force-field accuracy;
      - AlphaFold / RoseTTAFold-like methods accurately predict dominant conformations;
      - MSA sampling, clustering, diffusion, and flow-matching methods can increase structural diversity;
      - ESM3 provides a discrete structure-token representation and decoder.
      
      ESM3 is particularly important because it provides:
      
      > **3D structure → residue-level structure-token sequence → 3D structure**
      
      Thus, representation and decoding machinery already exist.
      
      ---
      
      ## 2.4 Missing Component
      
      The manuscript identifies the central missing component as:
      
      > **an effective sampling strategy**
      
      Standard ESM3 decoding, including argmax and temperature-controlled randomization, remains fundamentally local and rarely crosses barriers into alternative basins.
      
      MSFold therefore contributes:
      
      > **parallel tempering in structure-token space**
      
      and additionally introduces SLL for ranking.
      
      This closely matches the WIT Introduction logic:
      
      > **Final Goal → Necessary Components → Established Components → Missing Component → This Study**
      
      ### WIT Application
      
      **Strong.**
      
      However, there is one logical hierarchy that could be sharpened.
      
      The preceding paragraph explicitly states that:
      
      > **ranking is a further challenge**
      
      yet the next paragraph says:
      
      > **the missing component is effective sampling**
      
      while MSFold itself addresses both sampling and ranking.
      
      The storyline would become even cleaner if it explicitly distinguished:
      
      > **Primary missing component: effective global sampling**  
      > **Secondary unresolved bottleneck: reliable ranking / selection**
      
      This would preserve a focused central claim while acknowledging the second bottleneck.
      
      ---
      
      # 3. The Six-Dimensional Scientific Question Space
      
      WIT defines “opening up a problem” through six scientific dimensions:
      
      > **Existence → Determinants → Cause → Mechanism → Boundary Conditions → Magnitude**
      
      MSFold addresses nearly all six dimensions, but not equally completely.
      
      ---
      
      ## 3.1 Whether → Existence
      
      Question:
      
      > **Can MSFold actually recover multiple experimentally determined conformations?**
      
      Core benchmark:
      
      - 312 proteins;
      - success requires both Fold 1 and Fold 2 to reach TM-score ≥ 0.75;
      - MSFold: 161 / 312, success rate = **0.516**;
      - AlphaFold3: 0.378;
      - AlphaFold2: 0.360;
      - other methods are lower.
      
      Thus Existence is directly established:
      
      > **MSFold has a real advantage in multi-conformation recovery.**
      
      ---
      
      ## 3.2 What → Determinants
      
      Question:
      
      > **What factors determine success in multi-conformation prediction?**
      
      The manuscript progressively identifies several determinants:
      
      - sampling strategy;
      - replica exchange;
      - temperature regime;
      - representation of the search space;
      - ranking criterion;
      - sampling budget / parameter setting.
      
      Section 2.3 explicitly states:
      
      > **Sampling strategy is critical for multi-conformation prediction.**
      
      Section 2.6 then extends the determinant analysis from algorithm to representation.
      
      Thus the What dimension is well developed.
      
      ---
      
      ## 3.3 Why → Cause
      
      Question:
      
      > **Why do standard prediction / decoding methods tend to miss alternative conformations?**
      
      The primary causal explanation is:
      
      > **standard decoding and local sampling are biased toward the dominant local basin / dominant mode.**
      
      A particularly useful comparison holds the ESM3 structure-token representation fixed while changing the sampling strategy:
      
      - argmax;
      - iterative sampling;
      - MSFold parallel tempering.
      
      Their success rates are approximately 0.330, 0.370, and 0.516, respectively.
      
      This makes:
      
      > **sampling accessibility**
      
      a strong causal explanation for part of the performance gap.
      
      ---
      
      ## 3.4 How → Mechanism
      
      Question:
      
      > **How does parallel tempering make alternative states accessible?**
      
      The manuscript goes beyond “parallel tempering works” and analyzes:
      
      - low-T replicas: local refinement;
      - high-T replicas: broader exploration;
      - replica exchange: coupling the two;
      - trajectories: basin-to-basin transitions;
      - disabling exchange: reduced transition ability;
      - single-temperature low T: trapping in the dominant basin;
      - single-temperature high T: distorted structures.
      
      This forms a clear mechanism chain:
      
      > **local-search trapping**  
      > → **temperature-separated exploration/refinement**  
      > → **replica exchange**  
      > → **barrier crossing / basin transitions**  
      > → **alternative-state recovery**
      
      This is an excellent illustration of the Why/How distinction:
      
      > **Why:** the problem arises from local-search bias.  
      > **How:** parallel tempering resolves it through high-T exploration + low-T refinement + exchange.
      
      ---
      
      ## 3.5 When → Boundary Conditions
      
      The manuscript already provides several forms of boundary evidence:
      
      - diverse conformational changes;
      - temporal holdout;
      - pre-cutoff / post-cutoff structures;
      - multiple case studies;
      - high / intermediate / low temperature;
      - parameter sensitivity.
      
      The temporal evaluation is especially valuable:
      
      > On 32 proteins whose experimentally determined structures were released after the ESM3 training-data cutoff, MSFold achieves a success rate of **0.594**.
      
      This makes simple structural memorization an insufficient explanation of the gain.
      
      However, WIT would continue asking:
      
      > **When does MSFold succeed, and when does it fail?**
      
      A systematic boundary map is still incomplete. Useful analyses could include:
      
      - apo–holo vs. fold switching vs. secondary-structure changes;
      - small vs. large proteins;
      - single-domain vs. multi-domain proteins;
      - performance vs. Fold1–Fold2 structural distance;
      - performance vs. token entropy / diversity;
      - saturation as replica count or sampling steps increase.
      
      These analyses may offer:
      
      > **low experimental cost + high information gain**
      
      and therefore deserve high priority under WIT.
      
      ---
      
      ## 3.6 To What Extent → Magnitude
      
      The manuscript quantifies magnitude at several levels:
      
      - overall success rate: 0.516;
      - Fold 1 average TM-score: 0.821;
      - Fold 2 average TM-score: 0.740;
      - temporal holdout success rate: 0.594;
      - LAOBP endpoint TM-scores: approximately 0.99 / 0.98;
      - SLL vs. pTM / pLDDT correlations and ranking comparisons.
      
      Thus the paper does not stop at:
      
      > **“It works.”**
      
      It also asks:
      
      > **“How much does it work, and where does the gain come from?”**
      
      The Fold 1 / Fold 2 decomposition is particularly informative because it shows that the improvement is mainly due to alternative-state recovery rather than a trade-off that sacrifices dominant-state accuracy.
      
      ---
      
      # 4. Results Titles: Do They Form a “Small Essay”?
      
      The current Results titles are:
      
      1. **MSFold recovers multiple protein conformations across diverse proteins**
      2. **MSFold generalizes beyond known structures**
      3. **Sampling strategy is critical for multi-conformation prediction**
      4. **MSFold traverses conformational landscapes to access alternative states**
      5. **Parallel tempering enables exploration beyond local basins**
      6. **The discrete structure-token space reshapes conformational sampling**
      7. **Sequence log-likelihood improves selection of plausible conformations**
      
      Read in sequence, they reconstruct a clear question chain:
      
      > **Does it work?**  
      > → **Does it generalize beyond memorization?**  
      > → **What causes the gain?**  
      > → **What does successful exploration look like?**  
      > → **How does parallel tempering create transitions?**  
      > → **Why does representation matter?**  
      > → **After exploration, how should useful outputs be selected?**
      
      ### WIT Application
      
      **Very strong.**
      
      This is a canonical example of **Finding-driven Results**.
      
      For MSFold, the practical WIT lesson is:
      
      > **Results should expose the logical progression of the scientific story.**
      
      Here that progression is finding-driven. WIT does not require a particular title syntax; it asks whether the sequence of Results makes the reasoning recoverable.
      
      ---
      
      # 5. Results Subsections: Fact → Restrained 1-hop Opinion
      
      ## 5.1 Section 2.1: Performance
      
      ### Fact
      
      - MSFold success rate = 0.516;
      - Fold 1 accuracy remains comparable to AlphaFold3 / AlphaFold2;
      - Fold 2 accuracy is substantially higher.
      
      ### 1-hop Opinion
      
      The manuscript concludes that:
      
      > MSFold's primary improvement comes from enhanced Fold 2 recovery while maintaining Fold 1 accuracy.
      
      This is a clean:
      
      > **Fact → 1-hop Opinion**
      
      without prematurely jumping to the broader “representation + search” principle.
      
      ### WIT Application
      
      **Very good.**
      
      ---
      
      ## 5.2 Section 2.2: Generalization vs. Memorization
      
      ### Question
      
      > Is the improvement genuine generalization, or a consequence of training-data exposure / memorization?
      
      ### Test
      
      > A post-ESM3-cutoff temporal subset.
      
      ### Fact
      
      > MSFold achieves 0.594 success, comparable to 0.516 on the full benchmark, and continues to outperform most methods.
      
      ### 1-hop Opinion
      
      > The advantage is unlikely to be explained by simple memorization.
      
      The word:
      
      > **unlikely**
      
      is appropriately restrained. The manuscript does not claim that every possible form of memorization has been completely ruled out.
      
      ### A Borderline Inference
      
      In the second temporal analysis, AlphaFold2 remains below MSFold even when both endpoint structures were available before the cutoff. The manuscript then states that the major limitation likely lies in the prediction objective and inference procedure.
      
      This conclusion is one step further from the immediate fact.
      
      The qualifier `likely` helps, but from a WIT perspective this is close to:
      
      > **1.5-hop**
      
      because “not exposure alone” does not fully discriminate among every remaining explanation.
      
      A more precise interpretation would be:
      
      > **Prior exposure alone is insufficient; objective and inference remain plausible major contributors.**
      
      ---
      
      # 6. Section 2.3: Competing Hypotheses and a Discriminating Experiment
      
      This is one of the strongest WIT-style sections in the manuscript.
      
      Possible explanations:
      
      > **H1: The performance gain comes from the ESM3 representation.**  
      > **H2: The performance gain comes from improved sampling.**
      
      Experimental design:
      
      > Hold the ESM3 structure-token representation fixed and vary only the decoding / sampling strategy.
      
      Results:
      
      - argmax: 0.330;
      - iterative sampling: 0.370;
      - MSFold: 0.516.
      
      This provides meaningful discrimination and supports:
      
      > **sampling strategy is a critical determinant**
      
      ### WIT Application
      
      **Scientifically stronger than an ordinary baseline comparison.**
      
      It does not merely ask:
      
      > “Which method scores higher?”
      
      It asks:
      
      > **“What causes the score difference?”**
      
      This is:
      
      > **Competing Hypotheses → Discriminating Experiment**
      
      ---
      
      # 7. Sections 2.4–2.5: From Finding to Mechanism
      
      ## 7.1 Section 2.4: What Happens During Successful Sampling?
      
      The LAOBP trajectory shows:
      
      - periods of local refinement;
      - crossings between Fold 1 / Fold 2 similarity curves;
      - transitions between dominant and alternative basins;
      - two major clusters in the sampled ensemble;
      - similar behavior in additional case studies.
      
      This section first establishes:
      
      > **MSFold does not merely output two endpoints; its sampling process accesses distinct conformational basins.**
      
      Its primary reasoning function is:
      
      > **What does the process look like?**
      
      ---
      
      ## 7.2 Section 2.5: Why / How Do Transitions Occur?
      
      The next section asks a more mechanistic question:
      
      > **What mechanism supports basin transitions?**
      
      Evidence includes:
      
      - greater diversity in high-T replicas;
      - greater local concentration in low-T replicas;
      - reduced transition ability when replica exchange is disabled;
      - dominant-basin trapping under single-temperature low T;
      - structural distortion under single-temperature high T.
      
      This supports:
      
      > **high-T exploration + low-T refinement + replica exchange**
      
      ### WIT Application
      
      Sections 2.4 and 2.5 may look similar superficially, but their reasoning functions differ:
      
      > **2.4 = phenomenon / process observation**  
      > **2.5 = mechanism test**
      
      The separation is therefore justified rather than redundant.
      
      ---
      
      # 8. Section 2.6: Does the Representation Claim Go Too Far?
      
      This section is important because it moves the story from:
      
      > **sampling algorithm matters**
      
      toward:
      
      > **representation + sampling jointly matter**
      
      Evidence includes:
      
      - individual token updates can change a large fraction of residues;
      - large-scale token changes correspond to substantial conformational moves;
      - a structural basin can correspond to multiple token sequences;
      - in 10-ns Cartesian-space MD, most tested proteins do not undergo state transitions.
      
      The manuscript then states that:
      
      > **the success of sampling algorithms depends not only on the algorithm itself, but also on the representation of the search space**
      
      This is an attractive principle.
      
      From the perspective of WIT inference distance, however, it is also one of the places that deserves the most caution.
      
      The local experiment compares:
      
      > **ESM3 discrete token sampling**
      
      with:
      
      > **limited-timescale Cartesian MD**
      
      whereas the conclusion approaches:
      
      > **sampling algorithms in general depend on representation**
      
      This moves beyond a purely local 1-hop interpretation and toward a Discussion-level abstraction.
      
      ### WIT Suggestion
      
      A more restrained Results-level statement would be:
      
      > **Within the tested settings, the discrete structure-token representation changes the scale and connectivity of accessible conformational moves, facilitating broader exploration.**
      
      The Discussion can then abstract to:
      
      > **conformational prediction is a joint problem of representation and search.**
      
      This would produce a cleaner hierarchy:
      
      > **Fact → 1-hop → 2-hop / General Principle**
      
      The latest Discussion already does something important and correct: it explicitly limits the 10-ns MD comparison to sampling accessibility under the tested conditions and does not treat it as evidence for a general limitation of MD.
      
      That is strongly consistent with WIT.
      
      ---
      
      # 9. Section 2.7: Claim–Evidence Alignment for SLL
      
      The scientific motivation for SLL arises naturally:
      
      > **Successful broad exploration creates a new bottleneck: selection.**
      
      This is a good example of:
      
      > **Finding → New Question**
      
      For the LAOBP case:
      
      - SLL Spearman ρ = 0.71;
      - pTM = 0.65;
      - pLDDT = 0.68;
      - two high-quality structures are ranked within the top 1% by SLL while receiving much poorer ranks from pTM / pLDDT.
      
      However, on the 312-protein benchmark, the improvement in the top-ranked average TM-score is relatively modest:
      
      - Fold 1: SLL 0.720 vs. pTM 0.715 vs. pLDDT 0.697;
      - Fold 2: SLL 0.611 vs. pTM 0.607 vs. pLDDT 0.593.
      
      Thus:
      
      > **There is evidence that SLL improves selection.**
      
      But there is not evidence that:
      
      > **SLL solves the ranking problem.**
      
      The Discussion appropriately acknowledges that reliable selection of alternative conformations remains unresolved.
      
      ### WIT Application
      
      The overall claim is reasonably restrained.
      
      However, the Results conclusion:
      
      > **SLL provides an effective criterion**
      
      could be calibrated more tightly, for example:
      
      > **SLL provides a complementary reference-free criterion and modestly improves ranking on the benchmark, with larger gains in selected challenging cases.**
      
      This would align claim strength more closely with evidence strength.
      
      ---
      
      # 10. Discussion: Applying the WIT Core Functions
      
      The revised WIT does not require a fixed six-part Discussion. It focuses on two core functions:
      
      > **Integrated Interpretation → Broader Meaning / justified abstraction**
      
      MSFold performs these functions well.
      
      ---
      
      ## 10.1 Integrated Interpretation
      
      The first Discussion paragraph integrates:
      
      - benchmark improvement;
      - alternative-state recovery;
      - temporal holdout;
      - comparison with standard decoding.
      
      It then concludes that multi-conformation capability depends not only on what a pretrained model has learned, but also on how that information is accessed at inference time.
      
      This is not a repetition of Results.
      
      It is:
      
      > **multiple local findings → integrated 2-hop interpretation**
      
      ### One Important Inference
      
      The manuscript states that:
      
      > the pretrained model already encodes structural information relevant to multiple conformations.
      
      This is not a directly observed Fact.
      
      It is an interpretation jointly supported by:
      
      - fixed-representation sampling comparisons;
      - alternative-state recovery;
      - temporal generalization.
      
      This is exactly the type of statement that belongs in Discussion rather than in a local Results paragraph.
      
      ---
      
      ## 10.2 Broader Meaning / General Principle
      
      The second Discussion paragraph proposes:
      
      > **conformational prediction should be viewed as a joint problem of representation and search.**
      
      This principle is supported by multiple Results:
      
      - 2.3: sampling matters;
      - 2.5: the search mechanism matters;
      - 2.6: representation matters.
      
      It is therefore not an unsupported abstraction.
      
      This is a strong example of:
      
      > **multiple 1-hop findings → 2-hop interpretation → abstraction**
      
      ---
      
      ## 10.3 Optional Extensions: Limitations / Future Studies
      
      For MSFold, explicit limitations are scientifically useful because several unresolved bottlenecks are central to the story.
      
      The current Discussion identifies:
      
      - the upper bound imposed by information encoded in ESM3;
      - quantization / decoding errors;
      - computational cost;
      - the 10-ns MD timescale boundary;
      - unresolved ranking.
      
      Most of these limitations arise naturally from the scientific story rather than from a generic template:
      
      > **Finding → unresolved question → current constraint**
      
      Ranking is especially clear:
      
      > broader exploration → more conformations → harder selection
      
      Thus the Future Study direction:
      
      > better exploration + better selection
      
      is not a wish list.
      
      ### WIT Application
      
      **This is a good use of optional Discussion extensions.**
      
      WIT does not treat Limitations / Future Studies as mandatory surface sections. In MSFold, however:
      
      > **when an unresolved bottleneck is central to the scientific story, it should be stated explicitly.**
      
      ---
      
      # 11. Using WIT for Falsification / Counterexample Checks
      
      ## 11.1 “The MSFold Gain Is Merely Memorization”
      
      Potential falsifier:
      
      > The advantage disappears on post-cutoff proteins.
      
      Observed result:
      
      > It does not.
      
      Thus this competing explanation is substantially weakened.
      
      ---
      
      ## 11.2 “If the ESM3 Representation Is Good, Sampling Does Not Matter”
      
      Potential falsifier:
      
      > Argmax / iterative / MSFold behave similarly under the same representation.
      
      Observed result:
      
      > They do not.
      
      The sampling claim is strengthened.
      
      ---
      
      ## 11.3 “High Temperature Alone Is Enough”
      
      Potential falsifier:
      
      > Single high-T sampling works as well as MSFold.
      
      Observed behavior:
      
      > High T produces many distorted structures, whereas low T remains trapped in the dominant basin.
      
      Thus:
      
      > **high temperature alone is insufficient.**
      
      This supports the replica-exchange mechanism.
      
      ---
      
      ## 11.4 “Replica Exchange Is Not Important”
      
      Potential falsifier:
      
      > Disabling exchange leaves behavior unchanged.
      
      Observed result:
      
      > Transition and exploration ability decrease.
      
      Thus the exchange mechanism is strengthened.
      
      ---
      
      ## 11.5 “SLL Is a Reliable Ranking Solution”
      
      Here the data themselves provide a useful soft counterexample:
      
      > Benchmark-level gains are modest, and the Discussion acknowledges that reliable ranking remains unresolved.
      
      The WIT response is not to discard SLL, but to:
      
      > **Counterexample / weak gain → narrow the claim → better claim**
      
      That is:
      
      > SLL is a useful complementary reference-free ranking signal, not a complete solution to the ranking problem.
      
      This nicely illustrates:
      
      > **A successful falsification sharpens the claim.**
      
      ---
      
      # 12. Claim–Evidence Mapping
      
      | Major Claim | Main Evidence | WIT Judgment |
      |---|---|---|
      | MSFold improves multi-conformation recovery | 312-protein benchmark; 0.516 success rate | Strong support |
      | Gain mainly comes from Fold 2 recovery | Fold 1 / Fold 2 decomposition | Strong support |
      | Improvement is not simple memorization | post-ESM3-cutoff temporal subset | Strong evidence, but not an absolute exclusion of every memorization mechanism |
      | Sampling is a critical determinant | same ESM3 representation + different sampling strategies | Strong discriminating evidence |
      | Parallel tempering enables basin transitions | trajectories + exchange ablation + single-T comparison | Strong mechanism support |
      | Representation reshapes accessible conformational moves | token-update analysis + MD comparison | Supported, but generalization must remain bounded by tested conditions |
      | SLL improves ranking | case study + 312-protein benchmark | Supported, but overall gain is modest |
      | The pretrained PLM encodes more conformational information than standard decoding reveals | synthesis of sampling and temporal evidence | Reasonable 2-hop interpretation, not direct observation |
      | Prediction depends jointly on representation and search | synthesis of 2.3 / 2.5 / 2.6 | Reasonable General Principle |
      
      ---
      
      # 13. Reviewer Stress Test
      
      WIT asks:
      
      > **What are the strongest questions a demanding reviewer could raise?**
      
      For the current manuscript, at least four deserve attention.
      
      ## 13.1 Is the Sampling Budget / Ensemble Size Matched Fairly?
      
      MSFold generates:
      
      > **40 replicas × 500 steps = 20,000 conformations**
      
      while the success metric uses the best matching structures in the sampled ensemble.
      
      A reviewer can therefore reasonably ask:
      
      > **Do baselines use comparable sample counts / computational budgets?**
      
      If some baselines generate far fewer structures, best-of-N evaluation itself can confer an advantage.
      
      The main text does not make this control clear.
      
      If the Supplementary Material already controls this carefully:
      
      > **surface that fact more explicitly in the main paper.**
      
      If it does not:
      
      > **this is a high-priority discriminating control.**
      
      It could materially change the interpretation of the central performance claim.
      
      ---
      
      ## 13.2 Is the Representation-vs-MD Comparison Fair Enough?
      
      The 10-ns MD comparison naturally invites the objection:
      
      > **Is the lack of transitions simply due to insufficient timescale?**
      
      The Discussion already narrows the claim appropriately.
      
      However, if the manuscript wants stronger evidence for:
      
      > **representation reshapes sampling**
      
      a more matched representation-level control would be preferable to relying heavily on short MD.
      
      ---
      
      ## 13.3 Where Are the Boundary Conditions?
      
      The 312-protein benchmark is strong, but a reviewer may ask:
      
      > **Which proteins succeed, which fail, and why?**
      
      If the existing data permit it, valuable analyses include:
      
      - performance vs. conformational-change type;
      - performance vs. protein size;
      - performance vs. Fold1–Fold2 structural distance;
      - performance vs. token entropy / diversity;
      - performance vs. domain architecture.
      
      This is a classic:
      
      > **When → Boundary Conditions**
      
      question and likely offers high information gain at relatively low cost.
      
      ---
      
      ## 13.4 Is the Ranking Claim Too Strong?
      
      SLL shows an advantage, but the benchmark-average gain is modest.
      
      Thus:
      
      > **“improves” is well supported; “solves,” “reliable,” or “substantially better” would require more caution.**
      
      This is not a fatal flaw. It is a claim-calibration issue.
      
      ---
      
      # 14. Information Gain: Choosing the Next Step with WIT
      
      If WIT is used to choose the next experiment—not simply to list everything that could be done—the priorities might be:
      
      ## Priority 1: Matched Sampling-Budget Control
      
      Reason:
      
      > It could change the interpretation of the central benchmark advantage.
      
      This has the highest information gain.
      
      ---
      
      ## Priority 2: Failure / Boundary Analysis
      
      Reason:
      
      > Overall performance is already well established; one more similar benchmark may add little information.
      
      By contrast:
      
      > **when the method succeeds or fails**
      
      would substantially deepen scientific understanding.
      
      ---
      
      ## Priority 3: A Stronger Representation Control
      
      Goal:
      
      > Separate the contribution of the parallel-tempering algorithm from the contribution of a discrete representation that changes move topology.
      
      This would directly strengthen the General Principle:
      
      > **representation + search**
      
      ---
      
      ## Priority 4: Reframe Ranking Evaluation Around the Actual Task
      
      Rather than accumulating more correlation statistics, ask:
      
      > **In an unknown-native scenario, how effectively does SLL enrich high-quality alternative states into the top-k?**
      
      Useful metrics might include:
      
      - hit rate@k;
      - enrichment factor;
      - probability of recovering both states after top-k selection;
      - calibration.
      
      These may align more directly with the actual selection problem than average TM-score.
      
      ---
      
      # 15. Freezing the Research Storyline with WIT
      
      Under WIT, the latest MSFold storyline can be frozen as:
      
      ## Central Question
      
      > **A pretrained protein language model may contain information about multiple conformations; can a better inference-time search strategy access that latent capability?**
      
      ## Central Claim
      
      > **Parallel tempering in ESM3's discrete structure-token space improves access to alternative conformational states while preserving dominant-state accuracy.**
      
      ## Key Findings
      
      1. MSFold achieves a 51.6% multi-conformation success rate on 312 proteins;
      2. the gain mainly comes from Fold 2 recovery;
      3. the gain persists beyond the ESM3 temporal cutoff;
      4. sampling strategy is a critical determinant;
      5. replica exchange couples exploration and refinement to enable basin transitions;
      6. discrete token representation changes accessible structural moves;
      7. broader generation exposes ranking as a remaining bottleneck.
      
      ## General Principle
      
      > **A model's usable predictive capability depends not only on what its representation encodes, but also on how the learned space is searched and how sampled outputs are selected.**
      
      Once this storyline is frozen, every proposed new experiment should be evaluated by asking:
      
      > **Will it change the central claim, eliminate an important competing hypothesis, clarify a boundary condition, or merely add another supporting example?**
      
      ---
      
      ---
      
      # 16. How WIT Was Applied to MSFold
      
      MSFold is best treated as an **application case**, not as evidence that independently validates WIT. The project and WIT developed partly in parallel: some of the reasoning patterns were already being used before they were given their current WIT names.
      
      A retrospective reconstruction shows five especially important uses.
      
      ## 16.1 Opening Findings into New Questions
      
      The Results story grows through a sequence such as:
      
      > works
      > → generalizes?
      > → why?
      > → how?
      > → what role does representation play?
      > → how should outputs be ranked?
      
      This is the WIT loop in practice:
      
      > **Finding → New Question → Answerability Check → New Analysis / Experiment → New Finding**
      
      The point is not to force every project through the same list of questions. The point is to prevent a useful finding from being treated as the end of the reasoning process.
      
      ## 16.2 Separating Evidence from Inference Depth
      
      MSFold repeatedly benefits from distinguishing:
      
      > **Fact → restrained 1-hop interpretation → integrated 2-hop interpretation → broader abstraction**
      
      For example:
      
      - benchmark data support improved Fold 2 recovery;
      - same-representation sampling comparisons support the importance of sampling;
      - mechanism experiments support the role of replica exchange;
      - representation analyses support a bounded claim about accessible moves;
      - the Discussion can then integrate these findings into the broader **representation + search** principle.
      
      This helps decide not only **what to say**, but also **where a claim belongs** and **how strongly it should be stated**.
      
      ## 16.3 Turning Alternative Explanations into Experiments
      
      Several parts of MSFold can be formulated as competing hypotheses:
      
      - memorization vs. generalization;
      - representation vs. sampling;
      - high temperature alone vs. temperature coupling with replica exchange.
      
      WIT turns these from discussion points into a design question:
      
      > **What experiment most strongly distinguishes the plausible explanations?**
      
      Temporal holdout, same-representation sampling comparisons, and exchange ablation are examples of experiments with high discriminating value.
      
      ## 16.4 Choosing High-Information Next Steps
      
      Once the main benchmark result is established, WIT discourages adding experiments merely because they are easy or conventional. It asks which remaining uncertainty matters most.
      
      For the current MSFold manuscript, this leads to priorities such as:
      
      1. matched sampling-budget / ensemble-size controls;
      2. systematic failure and boundary-condition analysis;
      3. stronger isolation of representation effects;
      4. task-oriented evaluation of ranking.
      
      These steps are valuable because they can change or sharpen the interpretation of the central claim, not merely add more supporting examples.
      
      ## 16.5 Connecting Research and Writing
      
      WIT treats the paper structure as a consequence of the research logic rather than as a separate polishing stage. In MSFold:
      
      > **Performance → Generalization → Cause → Process → Mechanism → Representation → Selection**
      
      provides the Results progression, while the Discussion integrates these local findings into a broader interpretation.
      
      This is the sense in which **Writing Is Thinking**: deciding how to write the Results and Discussion exposes whether the underlying reasoning chain is complete, overextended, or missing an important link.
      
      ---
      
      # 17. Take-Home Message from the MSFold Case
      
      MSFold illustrates WIT as a tool for **human–LLM collaborative scientific research and writing**. It does not ask the LLM to autonomously produce a paper while the researcher waits. Instead, it uses the LLM to extend, challenge, organize, and audit the researcher's reasoning while keeping the researcher responsible for the scientific judgments that matter most.
      
      In this case, WIT helps with four connected tasks:
      
      > **open the scientific question space**
      > → **turn findings into discriminating next questions**
      > → **keep claims aligned with evidence**
      > → **convert the reasoning chain into a coherent scientific story**
      
      The most valuable next work is therefore not simply “more experiments.” It is the work that most reduces uncertainty about the central scientific interpretation. 
      
      > **Advance the research. Grow the researcher.**
      
      ---
      
      ## Source
      
      This case study is derived from the current WIT repository version of `Assess-WIT-using-MSFold.md`, which in turn is based on the manuscript:
      
      > *Sampling in structure-token space enables accurate prediction of multiple protein conformations*
      
      Reviewed manuscript file: `MSFold (3)(1).pdf`.
      
      Draft placeholders such as `XXX` and incomplete figure / supplementary references are not treated here as scientific-reasoning flaws.
      
    • README.md 85 B
      Case studies illustrate how WIT can be used in real scientific research and writing.
      
  • references
    • WIT-Scientific-thinking-and-writing-skill.md 67 KB
      # WIT: A Scientific Thinking and Writing Skill
      
      By Dongbo Bu  
      Institute of Computing Technology,  
      Chinese Academy of Sciences  
      Email: dbu@ict.ac.cn  
      2026/08/27
      
      > **Running examples:** This document mainly uses **MSFold** and **AlphaGo** as recurring examples. MSFold represents protein-structure / bioinformatics research, whereas AlphaGo represents a classic AI method/system study. The two examples are used to explain, stress-test, and refine WIT across different research styles, not to imply that all papers should follow the same surface form.
      
      ## 1. What WIT Means
      
      **WIT = Writing Is Thinking.**
      
      WIT is a workflow for **using writing to drive scientific thinking**. It connects problem formulation, experimental results, Discussion, Limitations, and Future Studies into one continuous process of scientific reasoning.
      
      Its core idea is simple:
      
      > **Writing is not merely the expression of completed thinking; it is part of scientific thinking itself.**
      
      WIT uses the **REWRITE loop** as its execution cycle:
      
      > **Research Question → Examine Literature → Work → Read Finding → Interrogate → Test Answerability → Extend / Exit**
      
      WIT is the overarching framework; REWRITE is its operational mechanism.
      
      > **Usage principle:** WIT is not merely a checklist. For important rules, it should explain *why*, provide a representative example, and give an actionable decision criterion.
      
      > **Meta-principle: WIT specifies the logical functions of scientific reasoning, not a rigid surface template for scientific prose.**
      >
      > The same reasoning function can be expressed through different prose forms. WIT should require the author to know why a section exists, what its evidence supports, and how it advances the story, but it should not require every paper to use the same subsection titles, paragraph order, or ritualized Limitations / Future Studies structure.
      >
      > **Reasoning structure ≠ Surface prose structure.**
      
      > **WIT is a question generator, not a checklist completer.**
      
      Its role is to expose overlooked scientific dimensions, competing explanations, and unresolved questions. It does not require every project to answer every generated question, nor every paper to explicitly contain every WIT module.
      
      > **Human–LLM collaboration principle: Advance the research. Grow the researcher.**
      
      WIT is designed not only to improve the scientific work, but also to strengthen the researcher's ability to ask questions, interpret evidence, compare explanations, design experiments, calibrate claims, and make scientific judgments. A human–LLM collaboration is not successful if the paper improves while the researcher becomes a passive supervisor of AI output.
      
      > **Automate labor; augment judgment.**
      
      The LLM should freely reduce low-learning-value labor when useful, but it should not automate away the reasoning that the researcher should learn to perform. The researcher should remain actively involved at the high-value judgment points: choosing what question matters, interpreting an important finding, proposing and comparing competing hypotheses, selecting discriminating experiments, deciding how strongly the evidence supports a claim, identifying boundaries, and deciding when the study is sufficient.
      
      The LLM therefore acts as a **scaffold, challenger, generator, and auditor of reasoning**: it can ask for an initial human judgment, expose overlooked alternatives, stress-test that judgment, and help refine it. This does not mean withholding useful answers or forcing a Socratic dialogue every time; when the user asks for a direct answer or faces a deadline, WIT should answer directly while still making the key assumptions, alternatives, and decision logic visible.
      
      ## 2. What Problems Does WIT Address?
      
      WIT is designed to address several common difficulties in scientific research and paper writing:
      
      1. **At the beginning of a project, there is only a vague idea, and it is unclear how to truly “open up” the problem.**  
         How can a single initial question be expanded into a researchable **question space** through **Whether / What / Why / How / When / To what extent**?
      
      2. **Results and Discussion are often mixed together.**  
         How far should a Results subsection go? When should it remain a Fact, and when can it include a 1-hop Opinion? Why should Discussion not simply repeat the Results?
      
      3. **A finding generates many new questions, but it is unclear which ones should be answered now and which should be left for later.**  
         Which questions can be answered with the current data or additional experiments and should become new Results? Which questions truly belong in Discussion, Limitations, and Future Studies?
      
      4. **Discussion is often either too shallow or overly speculative.**  
         How can multiple 1-hop Opinions be integrated into a 2-hop Interpretation, and then abstracted into a General Principle beyond this study while remaining within the evidence boundary?
      
      5. **Introduction often becomes a literature list of “many previous studies exist, but a gap remains.”**  
         How can we start from the final scientific goal, identify the necessary components already established by previous studies, and clarify which **missing component** is provided by the present study?
      
      6. **Limitations and Future Studies often become ritualized sections.**  
         How can limitations arise from important questions that the current study cannot answer, and how can future studies naturally follow from these unresolved questions?
      
      7. **Real research includes unexpected results, reviewer challenges, and deadlines.**  
         How should unexpected findings be handled? How can the work be stress-tested from a reviewer’s perspective? How can a minimal but complete, credible, and defensible scientific story be formed under limited time and resources?
      
      8. **AI can improve research output while weakening the researcher's own scientific ability if too much reasoning is outsourced.**  
         How can human–LLM collaboration advance the project while also training the researcher's ability to formulate questions, interpret findings, compare hypotheses, choose experiments, and make independent scientific judgments?
      
      ## 3. How to Use WIT
      
      WIT has two main ways of use:
      
      (1) **Agent Skill mode**: if the agent / IDE supports Agent Skills or can read skill files in a project, it is recommended to use [`SKILL.md`](../SKILL.md) as the entry point. The agent can then follow its rules to select the appropriate WIT mode, load the full workflow, and arrange appropriate **control transfer** between the researcher and the LLM.
      
      (2) **Directly load the WIT workflow**: if the current environment does not support skill discovery, or if WIT is only needed in a single conversation, provide the full WIT workflow file directly to the AI and explicitly ask it to work according to WIT.
      
      The relationship between the two is:
      
      > **The full WIT file defines the method; `SKILL.md` defines how an agent invokes and executes that method.**
      
      In other words, the full WIT file serves as a **reference / specification**, whereas [`SKILL.md`](../SKILL.md) serves as an **agent-facing executable collaboration protocol**. It specifies when WIT should be used, which mode should be invoked, what supporting material should be loaded, which steps may be automated by the AI, which key judgments should keep the researcher actively involved, and when the process should stop.
      
      ### 3.1 File Entry Points and Downloads
      
      WIT GitHub repository:
      
      > [https://github.com/deltadbu/WIT-skill](https://github.com/deltadbu/WIT-skill)
      
      Main files:
      
      - [`SKILL.md`](../SKILL.md) — agent execution entry point and human–LLM collaboration protocol; [direct download](https://github.com/deltadbu/WIT-skill/raw/refs/heads/main/wit/SKILL.md)
      - [`WIT-Scientific-thinking-and-writing-skill.md`](WIT-Scientific-thinking-and-writing-skill.md) — English full workflow; [direct download](https://github.com/deltadbu/WIT-skill/raw/refs/heads/main/wit/references/WIT-Scientific-thinking-and-writing-skill.md)
      - [`WIT-科学思考及写作skill.md`](WIT-科学思考及写作skill.md) — Chinese full workflow; [direct download](https://github.com/deltadbu/WIT-skill/raw/refs/heads/main/wit/references/WIT-%E7%A7%91%E5%AD%A6%E6%80%9D%E8%80%83%E5%8F%8A%E5%86%99%E4%BD%9Cskill.md)
      
      For long-term use, it is recommended to clone the entire repository rather than downloading only one file:
      
      ```bash
      git clone https://github.com/deltadbu/WIT-skill.git
      ```
      
      The installable skill package is the repository's [`wit/`](../) directory. Keeping that directory intact preserves the relative paths among `SKILL.md`, `references/`, `tests/`, and `case-studies/`.
      
      ---
      
      ### 3.2 Using `SKILL.md`: Let an Agent Coordinate WIT Automatically
      
      If the agent supports Agent Skills, place or install the repository's [`wit/`](../) directory—the installable WIT skill package—in a skill directory that the agent can discover. Different platforms may use different skill locations or installation mechanisms; WIT does not assume one fixed directory.
      
      Once the agent discovers [`SKILL.md`](../SKILL.md), it first reads its metadata and execution rules. It should treat `SKILL.md` as a **dispatcher + workflow controller + collaboration protocol**, mainly responsible for:
      
      (1) deciding whether the current task should use WIT;
      
      (2) identifying the appropriate mode, such as **Open a question, Advance from a finding, Choose the next experiment, Review Results, Review Discussion, Stress-test a study, or Deadline Mode**;
      
      (3) loading either the full [`WIT-Scientific-thinking-and-writing-skill.md`](WIT-Scientific-thinking-and-writing-skill.md) or [`WIT-科学思考及写作skill.md`](WIT-科学思考及写作skill.md), depending on the language and task;
      
      (4) loading supporting material from `tests/` or `case-studies/` inside the skill package (repository paths: [`wit/tests/`](../tests/) and [`wit/case-studies/`](../case-studies/)) only when needed, rather than placing all material into context by default;
      
      (5) executing WIT's REWRITE decision loop, reasoning invariants, and stop rule;
      
      (6) arranging **control transfer** in human–LLM collaboration: low-learning-value labor may be automated by the AI, while high-learning-value judgment nodes should keep the researcher substantively involved.
      
      Therefore, when using [`SKILL.md`](../SKILL.md), the user does not need to manually specify the entire WIT procedure at every step. For example, one can simply say:
      
      > Use WIT to analyze this finding and decide the next experiment.
      
      or:
      
      > Use WIT to review the Results and Discussion of this paper.
      
      If the agent has correctly discovered and loaded WIT, it should use [`SKILL.md`](../SKILL.md) to select the appropriate mode automatically, rather than asking the user to restate the full WIT workflow.
      
      If the current tool **cannot automatically discover skills** but can read files in the repository, the user can explicitly instruct:
      
      > Read [SKILL.md](../SKILL.md) first, then use WIT according to its routing, human–LLM collaboration rules, and stop rule. When the full English workflow is needed, read [WIT-Scientific-thinking-and-writing-skill.md](WIT-Scientific-thinking-and-writing-skill.md).
      
      Core principle:
      
      > **`SKILL.md` is not a replacement for the full WIT framework; it is the entry point and control program for an agent to execute WIT.**
      
      ---
      
      ### 3.3 When Skill Discovery Is Not Supported: Load the Full WIT Workflow Directly
      
      If WIT is being used in an ordinary ChatGPT conversation, or the current agent does not support automatic discovery of [`SKILL.md`](../SKILL.md), the simplest approach is to load the full WIT workflow directly.
      
      #### In ChatGPT
      
      Download and upload the English [`WIT-Scientific-thinking-and-writing-skill.md`](WIT-Scientific-thinking-and-writing-skill.md), or the Chinese [`WIT-科学思考及写作skill.md`](WIT-科学思考及写作skill.md), and then say:
      
      > Please read this WIT workflow and follow it throughout this conversation.
      
      After that, you can directly say:
      
      > Use WIT to expand this finding: ...
      
      or:
      
      > Use WIT to review the Results and Discussion of this paper.
      
      If you also want agent-level routing, human–LLM control transfer, and the stop rule, provide [`SKILL.md`](../SKILL.md) together with the full workflow.
      
      If you start a new conversation and these files are not automatically available, provide the files or GitHub links again.
      
      #### In ChatGPT Work / Project Spaces
      
      Place [`SKILL.md`](../SKILL.md) together with the English [`WIT-Scientific-thinking-and-writing-skill.md`](WIT-Scientific-thinking-and-writing-skill.md) or Chinese [`WIT-科学思考及写作skill.md`](WIT-科学思考及写作skill.md) in the relevant project materials, then say at the beginning of the project:
      
      > Please read WIT's `SKILL.md` and full workflow, and use them as the human–LLM collaborative rules for scientific thinking and writing in this project.
      
      This allows WIT to be used together with manuscript drafts, experimental results, and project documents over the long term.
      
      #### In VSCode / Codex / Copilot
      
      It is recommended to clone the [WIT GitHub repository](https://github.com/deltadbu/WIT-skill). If the tool supports Agent Skills, make the repository's [`wit/`](../) directory discoverable as the skill package and let it execute WIT from [`wit/SKILL.md`](../SKILL.md). If it does not, explicitly instruct it in the chat or project instructions:
      
      > Read [SKILL.md](../SKILL.md) first, then load the appropriate full WIT workflow and follow its routing, reasoning, human–LLM collaboration, and stop rules.
      
      Core principle:
      
      > **If Skill Discovery is supported: enter WIT through `SKILL.md`.**  
      > **If Skill Discovery is not supported: load the full WIT workflow directly; if more stable routing and collaboration control are needed, load `SKILL.md` as well.**
      
      ---
      
      ### 3.4 How to Invoke WIT After Loading
      
      #### Mode A: Open Up a Research Question
      
      Input:
      
      > Use WIT to open up this question: XXX.
      
      Expand it mainly from:
      
      > **Whether / What / Why / How / When / To what extent**
      
      The goal is to turn a vague question into a researchable **question space**.
      
      ---
      
      #### Mode B: Advance Research from a Finding
      
      Input:
      
      > This is a finding: XXX. Use WIT to expand it.
      
      Expected output:
      
      1. Fact;
      2. 1-hop Opinion;
      3. Literature positioning;
      4. New Why / How / What / When / Whether / To what extent questions;
      5. Which questions are answerable now;
      6. The most valuable additional experiments or analyses;
      7. Which questions are not answerable now;
      8. 2-hop Interpretation;
      9. General Principle;
      10. Potential Limitation, if relevant to the current story;
      11. Potential Future Study, if worth making explicit.
      
      This is the most important day-to-day use of WIT:
      
      > **Every time an important finding appears, open up the problem again.**
      
      ---
      
      #### Mode C: Review Results
      
      Input:
      
      > Use WIT to review the Results.
      
      Check whether:
      
      - the subsection titles form a clear and progressive scientific story;
      - each subsection has a logical function rather than merely adding another experiment;
      - the evidence supports a clear **Fact → restrained 1-hop Opinion**;
      - a **Question / Finding-driven** organization is appropriate, or whether a **Component / Pipeline-driven** organization is more natural for a method or system paper;
      - even when the surface structure follows a pipeline, the underlying **Question → Test → Finding → Next Step** reasoning remains recoverable;
      - questions answerable now have been prematurely pushed into Discussion;
      - key controls, competing explanations, counterexamples, or unexpected results have been overlooked;
      - **Early figures and their captions:** do not require Figure 1 itself to summarize the entire paper. In some method papers, Figure 1 is more useful for introducing the **problem, representation, or motivation** before presenting the solution; AlphaDev is a representative example. What matters is that an **early overview figure—often, but not necessarily, Figure 1—** allows the reader to understand the paper's central idea, major components, and high-level operation before reading the detailed Results. Audit the function of the early figures, not a fixed figure number.
      - **Figure captions should communicate both evidence and the immediate takeaway:** for a Results figure, the caption should not merely list panels or restate axes. It should primarily describe the **Fact** shown by the figure and may add a **restrained 1-hop Opinion**—the immediate take-home message directly supported by that figure. Avoid 2-hop interpretation, broad abstraction, or general principles that require integrating multiple findings; those usually belong in the Results text or Discussion. A useful test is: **What is shown? → What does this figure immediately tell us?** For overview or conceptual figures, the caption may instead explain the idea, workflow, and components without forcing a result-level opinion.
      - **Natural-language explanation and concrete walkthrough for mathematical or algorithmic methods:** if the core method is expressed mainly through mathematical formulas or computer algorithms, do not rely on formal notation or pseudocode alone. First explain the **basic idea in natural language**: what problem the formula or algorithm is trying to solve, what the key terms or steps mean, and why the procedure is designed this way. Then provide a minimal concrete example that instantiates the key terms in the formulas with actual values, or walks through the important steps of the algorithm on a small input. Prefer the **smallest example that exposes the essential mechanism**. The goal is to let the reader first grasp the idea conceptually and then **mentally execute the method**. A representative example is AlphaDev, where sorting only three numbers provides a compact and intuitive illustration of how the discovered algorithm operates.
      - **Worked case study for a method paper:** does the Results section contain a compact, representative case that lets the reader trace the method from input through the key intermediate steps to output? The case should explain **how the proposed method operates, why it succeeds, why existing methods fail or become unreliable on the same case, and what mechanism or design choice creates the difference.** Prefer the **simplest case that isolates the essential difference** between methods. The purpose is to explain mechanism, not to replace benchmark-level evidence.
      - **Difference-focused benchmark analysis:** aggregate metrics such as accuracy, precision, recall, AUC, correlation, or success rate establish whether and by how much a method performs better overall, but they rarely explain why. Check whether the authors identify the cases or subsets on which the proposed and existing methods exhibit the largest or most systematic differences, then analyze **where the performance gap comes from, why existing methods fail there, and why the proposed method succeeds.** Prefer a systematic pattern of differential performance over isolated cherry-picked examples.
      
      For method papers, these checks serve complementary functions:
      
      > **Early figures orient the reader: Figure 1 may introduce the problem, while an early overview figure explains the central idea and high-level operation; natural-language explanation plus a minimal concrete walkthrough makes a mathematical or algorithmic method operationally understandable; benchmark statistics establish whether and by how much the method works; a worked case explains why it succeeds and why existing methods fail; difference-focused analysis explains where the gain comes from and whether the mechanism is systematic.**
      
      Principle:
      
      > **Results should expose the logical progression of the scientific story, not obey a single title format.**
      
      ---
      
      #### Mode D: Review Discussion
      
      Input:
      
      > Use WIT to review the Discussion.
      
      First check the **core functions**:
      
      - does it integrate local findings / 1-hop Opinions into an overall interpretation?
      - does it go beyond repeating the Results?
      - does it explain the broader meaning and, when justified, abstract toward a General Principle?
      
      Then check the **optional extensions**, when scientifically useful:
      
      - should the findings reopen a new question space?
      - is there a meaningful boundary or limitation that should be stated?
      - do Future Studies address genuine unresolved questions rather than appear as ritual prose?
      
      Principle:
      
      > **The core function of Discussion is interpretation and abstraction. New Questions, Limitations, and Future Studies are valuable extensions when needed, not mandatory surface sections in every paper.**
      
      ---
      
      #### Mode E: Identify the Next Experiment
      
      Input:
      
      > Use WIT to determine the most valuable next experiment.
      
      Prioritize:
      
      1. questions that could change the central claim;
      2. major challenges a reviewer is likely to raise;
      3. competing explanations;
      4. boundary conditions;
      5. low-cost, high-information experiments.
      
      ---
      
      #### Mode F: Deadline Mode
      
      Input:
      
      > The deadline is close. Use WIT to help me close the project.
      
      Output:
      
      - Must do;
      - Should do;
      - Can omit;
      - Should be written as a Limitation;
      - Should be left for Future Study;
      - Whether the central claim should be narrowed.
      
      WIT's day-to-day use can be summarized in one sentence:
      
      > **If Skill Discovery is supported, let the agent enter WIT through `SKILL.md`; otherwise, load the full workflow directly. Then start from a question; every important finding should generate new questions; answerable questions can extend the current Results, while unanswerable ones should first be judged for importance before being developed into Discussion, Limitations, or Future Studies.**
      
      ## 4. REWRITE: The Core Research Loop
      
      WIT advances research through the **REWRITE** loop:
      
      > **Research Question → Examine Literature → Work / Experiment → Read Finding → Interrogate Finding → Test Answerability → Extend / Exit**
      
      These are not seven independent modules, but one continuous cycle.
      
      ### 4.1 R — Research Question
      
      Problem formulation is not merely choosing a topic. It means:
      
      > **expanding a single question into a scientific question space.**
      
      The goal is not to mechanically list interrogative words, but to map the problem into six relatively orthogonal scientific dimensions.
      
      #### 4.1.1 Whether → Existence
      
      First ask:
      
      > **Does the phenomenon really exist?**
      
      This is the basic existence question.
      
      For example:
      
      > Does MSFold truly recover alternative conformations more effectively than standard decoding?
      
      This dimension asks whether the phenomenon is stable, statistically supported, reproducible, and more than an accidental observation.
      
      #### 4.1.2 What → Determinants
      
      Once the phenomenon is established, ask:
      
      > **What variables or factors determine whether it occurs and how strongly?**
      
      For example:
      
      > What protein properties determine whether MSFold successfully recovers alternative conformations?
      
      Possible determinants may include protein size, conformational-change type, sequence identity, token-space diversity, and sampling budget.
      
      This dimension asks:
      
      > **What determines the outcome?**
      
      #### 4.1.3 Why → Cause
      
      Why does not ask for the detailed process. It asks:
      
      > **What causes the phenomenon?**
      
      That is, identify the causal driver.
      
      For example:
      
      > Why does standard decoding fail to recover alternative conformations?
      
      Possible causes include:
      
      - the states are not encoded in the representation;
      - search is trapped in a local mode;
      - the decoding objective favors the dominant state.
      
      Thus:
      
      > **Why asks what causes the phenomenon.**
      
      #### 4.1.4 How → Mechanism
      
      How is different from Why.
      
      Why asks:
      
      > **What causes the phenomenon?**
      
      How asks:
      
      > **Through what process or mechanism does that cause produce the outcome?**
      
      For example:
      
      > If standard decoding fails because it becomes trapped in local modes, how does parallel tempering help cross token-space barriers and reach alternative states?
      
      Thus:
      
      > **Cause → Mechanism → Outcome**
      
      and:
      
      > **Why asks what causes it; How asks through what mechanism the cause produces the effect.**
      
      #### 4.1.5 When → Boundary Conditions
      
      When should not be interpreted only as time.
      
      It represents the more general question:
      
      > **Under what conditions does the conclusion hold or fail?**
      
      For example:
      
      - For which protein classes does it hold?
      - For which conformational-change types does it fail?
      - Is the effect stronger in low-data or high-data regimes?
      - Over what sequence-identity range does it generalize?
      - In which tissues, cell types, or spatial regions does it hold?
      
      Many questions that might previously have been phrased as “Where” are actually boundary-condition questions, so Where does not need to be a separate dimension.
      
      This dimension asks:
      
      > **What are the boundary conditions?**
      
      #### 4.1.6 To what extent → Magnitude
      
      Finally ask:
      
      > **How large is the effect, and over what range does it hold?**
      
      For example:
      
      - How much does the success rate improve?
      - Is the gain concentrated in only a few samples?
      - How large a conformational change can be recovered?
      - Is the effect practically meaningful, not merely statistically significant?
      
      This dimension concerns:
      
      > **Magnitude / effect size / range**
      
      Thus, opening up a research question can be represented as six scientific dimensions:
      
      > **Existence → Determinants → Cause → Mechanism → Boundary Conditions → Magnitude**
      
      corresponding to:
      
      > **Whether → What → Why → How → When → To what extent**
      
      Not every project must answer all six. Their purpose is to systematically ask:
      
      > **Which important dimensions of the question space have not yet been opened?**
      
      
      #### 4.1.7 A More Familiar Example: Opening the Six-Dimensional Question Space with AlphaGo
      
      The MSFold example is useful for protein-structure research, but it may be unfamiliar to readers outside the field. AlphaGo provides a more widely understood example.
      
      The classic Nature paper *Mastering the game of Go with deep neural networks and tree search* begins from a clear difficulty: Go has an enormous search space. The paper describes an approximate branching factor of 250 and a typical game depth of 150, making exhaustive search infeasible. AlphaGo addresses this by reducing both the effective **breadth** and **depth** of search: a policy network prioritizes promising moves, a value network evaluates positions, and both are integrated with Monte Carlo tree search (MCTS).
      
      The purpose of this example is not to mechanically attach six interrogative words to AlphaGo. It is to show that:
      
      > **the same scientific result can be opened along six distinct scientific dimensions.**
      
      ##### Whether → Existence: Can AlphaGo Actually Reach Professional-Level Go?
      
      The first question is simply:
      
      > **Can deep neural networks combined with tree search actually solve the long-standing computer-Go problem at professional level?**
      
      The paper provides direct evidence:
      
      - AlphaGo achieved a **99.8% win rate** against other Go programs;
      - it defeated European Go champion Fan Hui **5–0** in the formal match.
      
      This establishes:
      
      > **The phenomenon exists: the approach reaches a level previously unattained by computer Go.**
      
      Whether does not ask *why* the system succeeds. It first asks:
      
      > **Does it succeed at all?**
      
      ##### What → Determinants: What Determines AlphaGo's Playing Strength?
      
      Once the phenomenon is established, the next question is:
      
      > **What variables or components determine how strong AlphaGo becomes?**
      
      The paper analyzes several determinants.
      
      For example:
      
      - the supervised policy network predicted expert moves with **57.0% accuracy**, compared with **44.4%** for the previous state of the art reported by the authors;
      - small improvements in policy prediction accuracy produced substantial improvements in playing strength;
      - after reinforcement learning, the RL policy network won **more than 80%** of games against the supervised policy network;
      - without search, the RL policy network won **85%** of games against Pachi;
      - policy quality, value estimation, rollout policy, and search all affect final playing strength.
      
      Thus What asks:
      
      > **Which components or variables determine performance?**
      
      It does not yet ask why those components work.
      
      ##### Why → Cause: Why Could AlphaGo Break Through Where Traditional Go Programs Struggled?
      
      Why asks for the causal explanation.
      
      The paper identifies two central obstacles:
      
      - enormous **search breadth**;
      - enormous **search depth**.
      
      Exhaustive search is therefore infeasible.
      
      A central causal explanation for AlphaGo's success is:
      
      > **it does not search all possibilities equally; learned policy and value functions drastically reduce the effective search space.**
      
      More specifically:
      
      - the policy network reduces effective **breadth**;
      - the value network reduces effective **depth**.
      
      Thus:
      
      > **Why = What caused the breakthrough?**
      
      This is deeper than simply saying “because it used deep learning.”
      
      A more informative statement is:
      
      > **AlphaGo succeeds because learning makes an otherwise intractable search problem tractable enough to search.**
      
      ##### How → Mechanism: How Do the Policy and Value Networks Actually Change MCTS?
      
      Once the cause has been identified, How asks:
      
      > **Through what computational process does that cause produce stronger play?**
      
      The mechanism can be decomposed as follows.
      
      1. **Policy network guides selection and expansion**
      
      The search does not explore all legal moves equally. The policy network gives higher priority to promising moves.
      
      2. **Value network evaluates leaf positions**
      
      At a leaf node, the value network can directly estimate the probability of winning rather than always playing the game to completion.
      
      3. **Rollout provides an additional evaluation**
      
      A fast rollout policy simulates the game to the end and provides another estimate.
      
      4. **MCTS backs the information up**
      
      The value-network estimate and rollout outcome are combined and propagated back through the search path to update action values and visit counts.
      
      Thus:
      
      > **Cause:** learning reduces the effective search space.  
      > **Mechanism:** policy-guided search + value evaluation + rollout + MCTS backup implement that reduction.
      
      This illustrates the distinction:
      
      > **Why asks what causes the success; How asks through what computational mechanism that cause produces stronger play.**
      
      ##### To What Extent → Magnitude: How Much Stronger Is AlphaGo?
      
      A magnitude question is not satisfied with:
      
      > “AlphaGo is strong.”
      
      It asks:
      
      > **How large is the improvement, and how large are the component effects?**
      
      The paper provides several quantitative scales:
      
      - **99.8% win rate** against other Go programs;
      - **5–0** against Fan Hui;
      - RL policy versus SL policy: **>80% win rate**;
      - RL policy without search versus Pachi: **85% win rate**;
      - a single value-network evaluation approached the accuracy of Monte Carlo rollouts using the RL policy while using about **15,000 times less computation**.
      
      These answer:
      
      > **How large is the effect?**
      
      The difference between Whether and To what extent is therefore clear:
      
      > **Whether: Is there an effect?**  
      > **To what extent: How large is it?**
      
      ##### When → Boundary Conditions: Under What Conditions Does AlphaGo's Advantage Hold or Fail?
      
      This dimension is especially instructive.
      
      The AlphaGo paper already shows that the approach works:
      
      - on full-sized Go;
      - against multiple computer Go programs;
      - against a professional-level human opponent.
      
      These provide partial evidence about the boundary.
      
      However, the paper does not systematically map all boundary conditions.
      
      WIT would therefore continue by asking:
      
      - Does the advantage persist when the search budget is greatly reduced?
      - Can MCTS compensate when the policy network is weak?
      - When does systematic error in the value estimate cause search to fail?
      - Does the advantage hold across opponents with very different styles and strengths?
      - Can the broader **learning + search** principle transfer beyond Go to other large decision problems?
      
      These are not all claims already answered by the original paper. They are new questions naturally generated from its findings:
      
      > **Under what conditions does the conclusion hold or fail?**
      
      That is the essence of Boundary Conditions.
      
      ---
      
      AlphaGo makes the six dimensions easy to distinguish:
      
      | Dimension | Question in the AlphaGo example |
      |---|---|
      | **Whether → Existence** | Can deep networks + search really achieve professional-level Go? |
      | **What → Determinants** | Which factors—policy accuracy, RL, value estimation, search—determine playing strength? |
      | **Why → Cause** | Why can this approach break through the traditional computer-Go bottleneck? |
      | **How → Mechanism** | How do policy/value networks interact with MCTS to change the search process? |
      | **When → Boundary Conditions** | Under what search budgets, opponent types, model qualities, and task conditions does the advantage hold or fail? |
      | **To what extent → Magnitude** | How large are the gains in win rate, playing strength, and computational efficiency? |
      
      Thus a finding such as:
      
      > **AlphaGo defeated a professional Go player.**
      
      is not the end of the research process. Once opened along the six dimensions, it becomes:
      
      > **Does it exist? → What determines it? → Why does it happen? → How does it happen? → Under what conditions? → How large is the effect?**
      
      That is what WIT means by:
      
      > **opening up the problem.**
      
      **Reference:** Silver, D. et al. *Mastering the game of Go with deep neural networks and tree search*. Nature 529, 484–489 (2016), doi:10.1038/nature16961.
      
      
      ### 4.2 E — Examine the Literature
      
      The literature is not decoration for the Introduction; it is the coordinate system for scientific reasoning.
      
      WIT uses two literature checkpoints.
      
      #### 4.2.1 Literature Checkpoint 1: Before the Study
      
      After defining the core scientific question, ask:
      
      1. Has this question already been answered?
      2. What explanations have previous studies proposed?
      3. What competing hypotheses exist?
      4. What boundary conditions are already known?
      5. What exactly does the present study add?
      
      The purpose is not to “collect enough references,” but to determine:
      
      > **Where is the novelty?**
      
      and:
      
      > **Which existing explanations must the current study distinguish among?**
      
      #### 4.2.2 Literature Checkpoint 2: After an Important Finding
      
      After every important finding, return to the literature and ask:
      
      > **How does this finding relate to existing knowledge?**
      
      Possible relationships include:
      
      - **Confirm**: supports an existing conclusion;
      - **Contradict**: conflicts with an existing conclusion;
      - **Refine**: adds boundaries, conditions, or a more precise interpretation;
      - **Extend**: generalizes an existing conclusion to new tasks, systems, or settings;
      - **Reframe**: changes how the original problem should be understood.
      
      This checkpoint directly determines the depth of the Discussion.
      
      What matters is usually not:
      
      > “Our result is consistent with a previous study.”
      
      but:
      
      > **What does our finding change, refine, or extend in the existing understanding?**
      
      ---
      
      ### 4.3 W — Work / Experiment
      
      Every experiment should correspond to a clear question.
      
      Do not perform an ablation simply because “papers usually need ablation.” Do not draw a t-SNE plot simply because others do.
      
      First ask:
      
      > **What question is this experiment intended to answer?**
      
      The ideal structure is:
      
      > **Question → Experiment → Data → Finding**
      
      Experiments are tools for answering questions, not the organizing units of Results.
      
      ---
      
      ### 4.4 R — Read the Finding
      
      After obtaining a result, do not immediately move on to the next experiment. First distinguish three levels.
      
      #### 4.4.1 Data
      
      Raw observations or quantitative results.
      
      #### 4.4.2 Finding
      
      A qualitative statement directly supported by the data.
      
      #### 4.4.3 1-hop Opinion
      
      A one-step interpretation that remains close to the data.
      
      A Results subsection can therefore be compressed into a practical **Fact–Opinion** structure:
      
      > **Question → Experiment → Fact → 1-hop Opinion**
      
      #### 4.4.4 Fact
      
      A fact directly obtained from experiments or analyses, including data, comparisons, observed phenomena, and statistical results.
      
      Example:
      
      > Method A significantly outperforms Method B under distribution shift.
      
      #### 4.4.5 1-hop Opinion
      
      A **one-step interpretation** of the Fact. It may contain author judgment, but it must remain close to the current result and should not jump directly to a broader theoretical claim.
      
      Example:
      
      > These results suggest that X may improve robustness under distribution shift.
      
      Therefore:
      
      > **Results = Fact + 1-hop Opinion**
      
      The end of a Results subsection should answer:
      
      > **What does this specific Fact mean?**
      
      But the Opinion should remain only one step away from the Fact.
      
      Core rule:
      
      > **Results may contain opinion, but only opinion that is one hop away from the fact.**
      
      Do not jump directly to a field-level general principle.
      
      ---
      
      #### 4.4.6 Unexpected Results and Serendipitous Findings: The Anomaly Branch
      
      Real research is not perfectly linear. Experiments often produce results opposite to the hypothesis, anomalous samples, unexpected subgroups, apparently failed but reproducible phenomena, or observations inconsistent with existing theory.
      
      After every important experiment, ask:
      
      > **Did the result match the original expectation?**
      
      If yes, continue with the normal WIT / REWRITE process.
      
      If no, enter the **Anomaly Branch**.
      
      ##### (1) Is it a technical error?
      
      Check for:
      
      - data-processing errors;
      - implementation bugs;
      - measurement errors;
      - data leakage;
      - batch effects;
      - sample contamination;
      - statistical artifacts.
      
      If yes:
      
      > Fix the problem and rerun the experiment.
      
      ##### (2) Is it random noise?
      
      Ask:
      
      > Is it reproducible?
      
      If not:
      
      > Do not treat it as a major finding yet.
      
      ##### (3) Is it a stable, reproducible anomaly?
      
      If yes, do not force it back into the original hypothesis.
      
      Ask:
      
      > **Does this anomaly imply that the original scientific question was incomplete—or even wrong?**
      
      At this point, allow:
      
      > **Unexpected Finding → Rewrite the Question**
      
      REWRITE is therefore not merely a way to organize completed results; it also allows new findings to redefine the research question itself.
      
      ---
      
      ### 4.5 I — Interrogate the Finding
      
      Every important finding should generate new questions.
      
      Systematically ask:
      
      - **Why?**
      - **How?**
      - **What?**
      - **Whether?**
      - **When?**
      - **Where?**
      - **To what extent?**
      
      A good finding should open up a new question space.
      
      ---
      
      
      #### 4.5.1 Competing Hypotheses: Do Not Settle on a Single Explanation Too Early
      
      After a finding appears, do not ask only:
      
      > **Why did this happen?**
      
      Force yourself to construct multiple plausible explanations:
      
      > **Finding → Hypothesis A / Hypothesis B / Hypothesis C**
      
      For example:
      
      > Finding: MSFold still outperforms baselines on unseen proteins.
      
      Possible explanations include:
      
      - **H1:** the ESM3 token space contains generalizable alternative conformations;
      - **H2:** the gain mainly comes from a larger sampling budget;
      - **H3:** some structural-transition classes in the benchmark are intrinsically easier for parallel tempering.
      
      The next step should not be to choose the most appealing explanation, but to ask:
      
      > **What experiment can distinguish these competing hypotheses?**
      
      Mechanistic research should prioritize **discriminating experiments**, not merely accumulate supportive evidence.
      
      ---
      
      #### 4.5.2 Falsification / Counterexample Check: Attack Your Own Conclusion
      
      For every important conclusion, ask:
      
      > **What result would falsify this conclusion?**
      
      and:
      
      > **What is the most plausible counterexample?**
      
      For example:
      
      > Claim: structural information improves OOD robustness.
      
      Do not only search for more supporting datasets. Also ask:
      
      - Is there a dataset where adding structural information hurts performance?
      - Does the advantage disappear when sequence diversity is already high?
      - Is the gain actually caused by additional model capacity rather than structural information itself?
      
      The purpose is to determine:
      
      > **Where is the evidence boundary of the claim?**
      
      If a conclusion has no clear potential falsifier, it is often not yet scientifically defined tightly enough.
      
      ---
      
      
      ##### What if a counterexample is found?
      
      Do not immediately conclude that the original conclusion is simply “wrong.” First classify the counterexample.
      
      - **Technical or data problem**: bug, measurement error, data leakage, sample contamination.  
        → Fix the issue and rerun the experiment.
      - **Random noise**: the result is not reproducible.  
        → Do not overturn the conclusion yet, but reduce confidence.
      - **Stable, reproducible counterexample**:  
        → The claim should usually be revised rather than discarded outright.
      
      A stable counterexample can have three major scientific consequences:
      
      1. **Narrow the claim**
      
      For example:
      
      > “X always improves Y”
      
      may become:
      
      > “X improves Y under conditions A and B.”
      
      2. **Reveal a boundary condition**
      
      The counterexample may answer:
      
      > **When does this conclusion fail?**
      
      This can be more informative than accumulating additional supportive examples.
      
      3. **Rewrite the hypothesis or principle**
      
      If the counterexample strikes at the proposed mechanism, revisit the central claim and possibly the entire research storyline.
      
      Thus:
      
      > **Counterexample → Boundary Condition → Better Claim**
      
      A good counterexample does not necessarily weaken a paper; it can make the conclusion more precise and credible.
      
      ##### What if no counterexample is found?
      
      Do not conclude that the claim has been “proven.”
      
      Instead, ask two further questions.
      
      1. **Was the falsification attempt strong enough?**
      
      Did we actively test:
      
      - adverse conditions;
      - extreme cases;
      - OOD data;
      - negative controls;
      - alternative explanations;
      - subgroups where failure is plausible?
      
      2. **Is the claim genuinely falsifiable?**
      
      If every possible result can be accommodated by rephrasing the claim, the claim may be too vague.
      
      State explicitly:
      
      > **What observation would make me accept that this claim is false?**
      
      The full logic is:
      
      > **Claim → Potential Falsifier → Test → Counterexample?**
      
      If a stable counterexample is found:
      
      > **Revise Claim / Identify Boundary / Rewrite Hypothesis**
      
      If none is found:
      
      > **The claim gains support, but is not “proven.”**
      
      If the potential falsifier cannot yet be tested:
      
      > **Record it as an unresolved question → Limitation / Future Study**
      
      A useful summary is:
      
      > **A failed falsification strengthens a claim; a successful falsification sharpens it.**
      
      
      ### 4.6 T — Test Answerability
      
      This is the key decision point in the REWRITE loop.
      
      For every new question, ask:
      
      > **Can the current study answer this question through additional analysis or experiments?**
      
      #### 4.6.1 If YES
      
      Do not move it prematurely into Discussion.
      
      Instead:
      
      > **New Question → New Experiment / Analysis → New Finding → New Results**
      
      For example, if A outperforms B on OOD data and the next question is whether the advantage is consistent across protein families, and the existing dataset already contains multiple families, analyze it now and turn it into a new Results subsection.
      
      Likewise, if “Which component drives the improvement?” can be answered through ablation, it should not be left as Future Work.
      
      #### 4.6.2 If NO
      
      Only then should the question move into Discussion:
      
      > **New Question → Why the Current Study Cannot Answer It → Limitation → Future Study**
      
      ---
      
      
      #### 4.6.3 Information Gain: Choose the Experiment That Most Changes Your Belief
      
      When several questions are answerable, do not choose only by convenience.
      
      A better rule is:
      
      > **Prioritize experiments with the highest information gain and discrimination power.**
      
      Ask:
      
      - If the result is A, how much would my interpretation change?
      - If the result is B, would I revise the central claim?
      - Can the experiment distinguish two currently plausible explanations?
      - Does it merely add another supporting example, or substantially reduce uncertainty?
      
      For example:
      
      - one more similar benchmark may increase confidence only slightly;
      - a control experiment that separates two competing mechanisms may change the interpretation entirely.
      
      Thus:
      
      > **Next experiment ≠ easiest experiment**  
      > **Next experiment = highest-value information test**
      
      ---
      
      ### 4.7 E — Extend / Exit
      
      If a new question is answerable now:
      
      > **New Question → New Experiment / Analysis → New Results**
      
      If it is not answerable now or lies beyond scope:
      
      > **New Question → Discussion → Limitation → Future Study**
      
      ## 5. Results: Expose the Logical Progression of the Scientific Story
      
      The strong rule is not:
      
      > **Every Results subsection must be titled as a scientific question or answer.**
      
      The deeper requirement is:
      
      > **Results should expose the logical progression of the scientific story.**
      
      The reader should be able to recover:
      
      > **why this part is needed → what evidence was obtained → what local conclusion the evidence supports → why the next part follows.**
      
      ### 5.1 Two Valid Modes of Results Organization
      
      #### (1) Question / Finding-driven
      
      This is often natural for discovery, mechanism, and hypothesis-testing studies.
      
      Subsection titles may state the answer directly, for example:
      
      - “Module X is the primary contributor to the performance gain”
      - “The performance advantage persists under distribution shift”
      - “The learned representation better separates functional states”
      
      The advantage is:
      
      > **the title directly tells the reader what was learned.**
      
      #### (2) Component / Pipeline-driven
      
      For method, system, and engineering papers, component- or pipeline-driven organization can be equally effective.
      
      The AlphaGo Nature paper is an important counterexample to an overly rigid WIT rule. Its major research sections proceed through:
      
      - Supervised learning of policy networks
      - Reinforcement learning of policy networks
      - Reinforcement learning of value networks
      - Searching with policy and value networks
      - Evaluating the playing strength of AlphaGo
      
      These headings are clearly **component / pipeline-driven**.
      
      Yet together they form a coherent storyline:
      
      > **learn a policy → improve it by self-play → learn a value function → integrate policy and value into search → evaluate the complete system**
      
      Thus:
      
      > **Pipeline-driven ≠ a mere list of techniques.**
      
      The key question is whether the components form a meaningful logical progression.
      
      ### 5.2 Separate Reasoning Structure from Surface Prose
      
      The reasoning behind a Results subsection can often be reconstructed as:
      
      > **Motivation / Question → Experiment / Analysis → Fact → 1-hop Opinion → Next Question / Next Step**
      
      But the prose does not need to mechanically state each element.
      
      AlphaGo sometimes advances with pipeline language such as “the second stage of the training pipeline ...” rather than repeatedly writing “we next asked whether ...”.
      
      That is perfectly acceptable when the underlying reasoning remains clear.
      
      Therefore:
      
      > **WIT requires recoverable reasoning, not formulaic prose.**
      
      ### 5.3 Fact + Restrained 1-hop Opinion Remains a Strong Rule
      
      Results may interpret evidence, but the interpretation should stay close to that evidence.
      
      Principle:
      
      > **Results: Fact → restrained 1-hop Opinion**
      
      For example, AlphaGo's tournament results provide strong facts about playing strength. The paper draws local conclusions about system strength rather than immediately jumping to a universal theory of intelligence.
      
      Likewise, when combined value-network and rollout evaluation outperformed either mechanism alone, the authors inferred that the two position-evaluation mechanisms were complementary.
      
      That is a clean example of:
      
      > **Fact → 1-hop Opinion**
      
      A useful test is:
      
      > **If the data in this subsection were removed, would the opinion still stand?**
      
      If not, it is probably a local Results-level interpretation.  
      If the claim requires multiple findings across the paper, it more likely belongs in Discussion.
      
      ### 5.4 The Real Test for Results Titles: Can They Form a Small Essay?
      
      Extract all Results subsection titles and read them in order.
      
      Ask:
      
      > **Do they independently reveal how the scientific story progresses?**
      
      The story may be:
      
      > **Question → Finding → Mechanism → Boundary**
      
      or, as in AlphaGo:
      
      > **Component A → Component B → Integration → System Evaluation**
      
      The small-essay test therefore evaluates:
      
      > **logical progression**
      
      not a single required title style.
      
      ### 5.5 Guided Template: Results Subsection
      
      For less experienced researchers, use the reasoning template first, then choose the final prose form.
      
      #### (1) Motivation / Scientific Question
      
      > Why is this part needed? We want to know: ________________________.
      
      #### (2) Experiment / Analysis
      
      > To answer or advance this question, we: ________________________.
      
      #### (3) Fact
      
      > The data directly show: ________________________.
      
      #### (4) 1-hop Opinion
      
      > These results locally suggest: ________________________.
      
      #### (5) Next Question / Next Step
      
      > Therefore, the next logical step is: ________________________.
      
      For a method or system paper, also ask:
      
      > **What indispensable logical function does this subsection serve in the overall pipeline?**
      
      The template is meant to expose reasoning, not force the final paper into five formulaic sentences.
      
      ---
      
      ## 6. Discussion: Core Functions + Optional Extensions
      
      AlphaGo provides an important correction to WIT:
      
      > **A strong Discussion does not have to explicitly contain New Questions, Limitations, and Future Studies.**
      
      Discussion should therefore distinguish between:
      
      > **Core Functions (strong rules)**  
      > and  
      > **Optional Extensions (used when scientifically helpful)**
      
      ### 6.1 Core Function 1: Integrated Interpretation
      
      Discussion should first move beyond line-by-line repetition of Results.
      
      It should integrate local findings / 1-hop Opinions and ask:
      
      > **Taken together, what do these results mean?**
      
      This can be represented as:
      
      > **Multiple local findings → Integrated Interpretation**
      
      When the synthesis moves one level beyond individual Results subsections, it can also be described as:
      
      > **multiple 1-hop Opinions → 2-hop Interpretation**
      
      But “2-hop” describes reasoning depth; it is not a required sentence pattern.
      
      ### 6.2 Core Function 2: Broader Meaning / General Principle
      
      After integration, Discussion often asks:
      
      > **Why do these results matter beyond the immediate experiments?**
      
      Possible forms include:
      
      - deeper causal interpretation;
      - conceptual comparison with an existing paradigm or baseline;
      - transferable mechanism;
      - broader implication;
      - a General Principle.
      
      For MSFold, one possible abstraction is:
      
      > **Model capability is jointly determined by representation and search.**
      
      But not every paper needs a grand General Principle. Abstraction should stop where the evidence stops.
      
      ### 6.3 AlphaGo as a Reverse Validation: No Fixed Six-Part Discussion Is Required
      
      AlphaGo's Discussion is short, yet logically strong.
      
      It performs roughly three functions:
      
      #### (1) Integrated Interpretation
      
      It brings policy networks, value networks, reinforcement learning, and tree search together as one system rather than repeating individual performance numbers.
      
      #### (2) Deeper Meaning
      
      Through comparison with conventional high-intensity search systems, it highlights a deeper computational idea: the system does not merely examine more positions; it learns **where to search and how to evaluate**.
      
      This moves beyond component-level Results toward a deeper interpretation.
      
      #### (3) Beyond This Study
      
      The final part treats Go as an instance of a broader class of difficult decision/search problems and points toward implications of the learning + search idea beyond Go.
      
      Thus AlphaGo completes:
      
      > **this study → broader meaning / abstraction**
      
      without explicitly writing:
      
      > New Questions → Limitations → Future Studies
      
      AlphaGo is therefore an important counterexample:
      
      > **The reasoning functions of Discussion can be complete even when the surface structure does not contain ritualized limitations or future-work paragraphs.**
      
      ### 6.4 Optional Extension 1: Raise New Questions
      
      When new questions are important for defining boundaries or opening the next research stage, continue with:
      
      > **Integrated Interpretation / General Principle → New Question Space**
      
      Use the six scientific dimensions:
      
      > **Existence → Determinants → Cause → Mechanism → Boundary Conditions → Magnitude**
      
      corresponding to:
      
      > **Whether → What → Why → How → When → To what extent**
      
      These questions are tools for continuing research. They do not all have to appear in the current paper.
      
      ### 6.5 Optional Extension 2: Limitations
      
      A Limitation is valuable when an important unresolved question is genuinely constrained by the current study design, data, or scope.
      
      A meaningful limitation remains:
      
      > **A constraint that prevents the current study from answering an important question raised by its own findings.**
      
      It is not:
      
      > **a generic list of things the authors did not do.**
      
      But if the claim boundaries are already clear and an explicit limitation paragraph adds little, WIT should not require one merely for template completeness.
      
      ### 6.6 Optional Extension 3: Future Studies
      
      Future Study is most useful when it directly addresses an unresolved question generated by the current work:
      
      > **Finding → New Question → Limitation → Future Study**
      
      This remains a powerful:
      
      > **research-planning logic**
      
      but it does not always need to become a:
      
      > **paper-surface paragraph**
      
      ### 6.7 How to Judge a Discussion
      
      More important than asking whether it contains a Limitations or Future Work section is asking:
      
      - does it move beyond repeating Results?
      - does it provide an integrated interpretation?
      - does it explain broader meaning at a level justified by the evidence?
      - is the abstraction too strong?
      - if it proposes a General Principle, do multiple findings support it?
      - if important boundaries or unresolved questions exist, are they handled honestly?
      - does the ending leave a clear take-home message?
      
      Thus:
      
      > **Core Discussion: Interpretation → Broader Meaning**
      
      When useful, extend to:
      
      > **→ New Questions → Boundary / Limitations → Future Studies**
      
      ### 6.8 Guided Template: Discussion
      
      Complete the core first, then decide whether optional extensions are needed.
      
      #### Core A: Integrated Interpretation
      
      > Taken together, the findings show: ________________________.
      
      #### Core B: Broader Meaning
      
      > At a deeper level, these findings mean: ________________________.
      
      > Within the evidence boundary, a broader interpretation is: ________________________.
      
      #### Optional C: New Question Space
      
      > The most important unresolved question is: ________________________.
      
      #### Optional D: Limitation
      
      > The current study cannot answer it because: ________________________.
      
      #### Optional E: Future Study
      
      > Answering it would require: ________________________.
      
      Principle:
      
      > **Do not write Optional C–E merely to complete a template.**
      
      ---
      
      ## 7. Claim–Evidence Mapping: What Exactly Supports Each Major Claim?
      
      Before the reviewer stress test, extract all major claims and map each one to its evidence:
      
      > **Claim → Figure / Table → Evidence → Strength → Remaining Uncertainty**
      
      Example:
      
      | Claim | Supporting evidence | Remaining uncertainty |
      |---|---|---|
      | MSFold improves alternative-conformation recovery | Fig. 2 benchmark | Could sampling budget explain part of the gain? |
      | The improvement is not due to memorization | unseen-protein test | Does this hold across all protein classes? |
      | Parallel tempering enables broader exploration | sampling analysis / ablation | Does it correspond to physical energy barriers? |
      
      Check:
      
      - Is there any major claim without direct evidence?
      - Is one figure being asked to support too many conclusions?
      - Does Discussion contain claims never supported in Results?
      - Is the strength of the claim greater than the strength of the evidence?
      - Should “demonstrates” be narrowed to “suggests”?
      
      This step establishes:
      
      > **Claim–Evidence alignment**
      
      If a claim cannot be mapped to explicit evidence:
      
      > **add evidence, narrow the claim, or remove it.**
      
      ---
      
      ## 8. Reviewer Stress Test
      
      Before submission, actively switch to a reviewer’s perspective.
      
      Ask:
      
      > **If I were the most demanding and knowledgeable reviewer, what would be the top three challenges to this paper?**
      
      Generate:
      
      > **Top 3 Reviewer Challenges**
      
      Then classify them.
      
      ### 8.1 Can be resolved with additional experiments now
      
      → Return to Results.
      
      ### 8.2 Can be resolved through existing data analysis or interpretation
      
      → Strengthen Results / Discussion.
      
      ### 8.3 Cannot genuinely be resolved within the current study
      
      → Write as a Limitation and design a corresponding Future Study.
      
      ### 8.4 Fatal Flaw
      
      Examples include an unfair central comparison, mismatch between claim and experimental design, severe confounding, or a central conclusion unsupported by the available evidence.
      
      Do not simply place these in Limitations.
      
      Instead:
      
      > **Redesign the study or narrow the central claim.**
      
      Therefore:
      
      > **Reviewer challenge ≠ limitation**
      
      The reviewer perspective is a **stress test**; a limitation is only one possible outcome.
      
      ---
      
      ## 9. Deadline Mode: How to Close a Study Under Limited Time
      
      WIT / REWRITE can create a problem: every finding can generate more questions, so research can continue indefinitely.
      
      Real research is constrained by submission deadlines, graduation timelines, computational resources, wet-lab costs, and project duration.
      
      Therefore, WIT includes a:
      
      > **Stop Rule / Minimum Sufficient Story**
      
      ### 9.1 Near a Deadline: prioritize three types of questions
      
      #### (1) Questions that could change the central claim
      
      If a different answer would invalidate or substantially narrow the main conclusion, address it first.
      
      #### (2) Fatal questions a reviewer is highly likely to raise
      
      Examples include:
      
      - data leakage;
      - memorization;
      - unfair baselines;
      - missing key controls;
      - alternative explanations.
      
      Address these first.
      
      #### (3) Low-cost questions that substantially improve interpretation
      
      Examples:
      
      - a simple ablation;
      - subgroup analysis;
      - error analysis;
      - a key negative control.
      
      If the cost is low and the information gain is high, prioritize them.
      
      ### 9.2 Questions that can be deferred
      
      If a question:
      
      - does not change the central claim;
      - does not affect the main evidence chain;
      - requires substantial new experiments;
      - is better suited to a separate study;
      
      then move it to:
      
      > **Discussion → Limitation → Future Study**
      
      
      ### 9.3 Research Storyline Freeze: Prevent the Project from Expanding Without Focus
      
      At a mature stage of the project, temporarily freeze the storyline by writing down:
      
      > **Central Question**  
      > **Central Claim**  
      > **3–5 Key Findings**  
      > **General Principle**
      
      Then ask for every proposed new experiment:
      
      > **Will it change or substantially strengthen this storyline?**
      
      If it:
      
      - changes the central claim → high priority;
      - rules out an important competing hypothesis → high priority;
      - establishes an important boundary condition → potentially valuable;
      - merely adds another similar result → proceed cautiously.
      
      Storyline Freeze does not mean the story can never change. A strong unexpected finding may justify rewriting it.
      
      Its purpose is to prevent:
      
      > **the project from losing its center because too many “also possible” experiments are added.**
      
      ---
      
      ### 9.4 Stop Rule
      
      
      Stop expanding Results when:
      
      1. the central scientific question has been answered credibly;
      2. major competing explanations have been ruled out to a reasonable extent;
      3. major reviewer challenges have been addressed;
      
    • WIT-科学思考及写作skill.md 65.9 KB
      # WIT:科学思考及写作skill
      
      By Dongbo Bu  
      Institute of Computing Technology,  
      Chinese Academy of Sciences  
      Email: dbu@ict.ac.cn  
      2026/08/28
      
      
      > **贯穿示例:** 本文主要以 **MSFold** 和 **AlphaGo** 两篇研究作为贯穿示例。MSFold 代表蛋白质结构 / 生物信息学研究,AlphaGo 代表经典 AI / method-system research。两者用于从不同研究范式解释、检验和修正 WIT,而不是要求所有论文都采用相同的表面写法。
      
      ## 1. WIT 的释义
      
      **WIT = Writing Is Thinking。**
      
      WIT 是一套“**用写作推动科学思考**”的工作流:把开题、实验结果、Discussion、Limitations 和 Future Studies 串成一个连续的科研推理过程。
      
      它的核心观点只有一句:
      
      > **写作不是思考完成后的表达,而是科学思考本身的一部分。**
      
      WIT 内部使用 **REWRITE loop** 作为执行循环:
      
      > **Research Question → Examine Literature → Work → Read Finding → Interrogate → Test Answerability → Extend / Exit**
      
      其中,WIT 是总框架,REWRITE 是具体运行机制。
      
      > **使用原则:** WIT 不是只给结论的 checklist。对于关键规则,应同时说明“为什么”、给出典型例子,并提供可执行的判断标准。
      
      > **元原则:WIT 约束的是科学思考需要完成的逻辑功能,而不是论文表面的固定格式。**
      >
      > 同一种 reasoning function,可以用不同的 prose form 表达。WIT 应要求作者能够回答“为什么这一部分存在、证据支持什么、下一步逻辑是什么”,但不要求每篇论文都机械地使用同一种 subsection title、同一种段落顺序或固定的 Limitations / Future Studies 模板。
      >
      > **Reasoning structure ≠ Surface prose structure.**
      
      > **WIT 用来生成问题,而不是用来完成清单。**
      
      WIT 的作用是帮助研究者暴露可能遗漏的 scientific dimensions、competing explanations 和 unresolved questions;它不要求每个 project 把所有问题逐项回答,也不要求每篇论文把所有模块逐项写完。
      
      > **人机协同原则:推进研究,成长研究者。**  
      > **Advance the research. Grow the researcher.**
      
      WIT 不仅要帮助研究者把研究做得更好,也要让研究者在协同过程中提升提问、解释证据、比较不同解释、设计实验、校准 claim 和进行 scientific judgment 的能力。如果 paper 变好了,而研究者只是被动等待 AI 产出,那么这种 human–LLM collaboration 并没有达到 WIT 的目标。
      
      > **Automate labor; augment judgment.**  
      > **自动化劳动,增强判断。**
      
      对于低 learning-value 的劳动,例如检索整理、格式处理、重复性 coding、机械分析和文字整理,可以充分利用 AI 自动化;但不应把研究者本应掌握的 reasoning 自动化掉。研究者应持续参与高价值的 judgment nodes:什么问题值得研究、一个重要 finding 应如何解释、有哪些 competing hypotheses、哪个 experiment 最有 discrimination power、证据能够支持多强的 claim、结论的 boundary 在哪里,以及什么时候应该停止扩展研究。
      
      因此,LLM 在 WIT 中应主要扮演 **scaffold、challenger、generator 和 auditor of reasoning**:必要时先让研究者给出初步判断,再补充遗漏、提出反例、比较 alternatives、帮助修正判断。这并不意味着每次都要强行采用苏格拉底式问答;如果用户明确要求直接答案,或者处于 Deadline Mode,WIT 应直接帮助完成任务,同时仍把关键假设、备选解释和决策逻辑显性化。
      
      ## 2. WIT 解决哪些痛点问题
      
      WIT 主要解决科学研究与论文写作中以下常见困惑:
      
      (1)**开题时只有一个模糊想法,不知道怎样把问题真正“打开”。**  
         如何从 **Whether / What / Why / How / When / To what extent** 等角度,把一个单点问题展开成可以研究的 question space?
      
      (2)**Results 与 Discussion 经常混在一起。**  
         Results 的 subsection 应该写到什么程度?什么时候只是 Fact,什么时候可以加入 1-hop Opinion?Discussion 为什么不能只是把 Results 再重复一遍?
      
      (3)**一个 finding 会产生很多新问题,但不知道哪些应该现在做,哪些应该留到以后。**  
         哪些问题能够通过当前数据或补充实验回答,并应继续形成新的 Results?哪些问题才真正属于 Discussion、Limitations 和 Future Studies?
      
      (4)**Discussion 容易要么太浅,要么过度拔高。**  
         如何从多个 1-hop Opinions 综合成 2-hop Interpretation,再进一步抽象成 beyond this study 的 General Principle,同时保持证据边界?
      
      (5)**Introduction 常常只是“已有工作很多,但仍有 gap”的文献罗列。**  
         如何从最终科学目标出发,识别 previous studies 已经建立的必要组件,并说明本研究补上的真正是哪个 **missing component**?
      
      (6)**Limitations 和 Future Studies 容易变成例行公事。**  
         如何让 limitation 来自“当前研究无法回答的重要问题”,并让 future study 成为针对这些 unresolved questions 的自然下一步?
      
      (7)**真实研究会出现异常结果、审稿人挑战和 deadline。**  
         如何处理 unexpected findings?如何从 reviewer perspective 做压力测试?如何在有限时间和资源下形成一个最小但完整、可信、可辩护的 scientific story?
      
      (8)**AI 可以提高研究产出,但如果把 reasoning 过度外包,也可能削弱研究者自身的科研能力。**  
         如何让 human–LLM collaboration 在推进 project 的同时,也训练研究者提出问题、解释 finding、比较 hypotheses、选择实验和独立进行 scientific judgment 的能力?
      
      ## 3. 如何使用 WIT
      
      WIT 有两种主要使用方式:
      
      (1)**Agent Skill 模式**:如果所使用的 agent / IDE 支持 Agent Skills 或能够读取项目中的 skill 文件,推荐让 [`SKILL.md`](../SKILL.md) 作为入口,由 agent 按其中的规则自动选择 WIT 的工作模式、加载完整 workflow,并在人与 LLM 之间安排合适的 control transfer。
      
      (2)**直接加载 WIT workflow**:如果当前环境不支持 skill discovery,或者只是想在一次对话中使用 WIT,可以直接把完整的 WIT workflow 文件提供给 AI,再明确要求它按照 WIT 工作。
      
      二者的关系是:
      
      > **完整 WIT 文件定义方法;`SKILL.md` 定义 agent 如何调用和执行这个方法。**
      
      也就是说,完整 WIT 文件更像 **reference / specification**;[`SKILL.md`](../SKILL.md) 更像 **agent-facing executable collaboration protocol**:它规定什么时候调用 WIT、调用哪个 mode、需要加载哪些材料、哪些步骤可以由 AI 自动完成、哪些关键 judgment 应让研究者参与,以及什么时候停止。
      
      ### 3.1 文件入口与下载
      
      WIT 的 GitHub 仓库:
      
      > [https://github.com/deltadbu/WIT-skill](https://github.com/deltadbu/WIT-skill)
      
      主要文件:
      
      - [`SKILL.md`](../SKILL.md) — agent 的执行入口与人机协同协议;[直接下载](https://github.com/deltadbu/WIT-skill/raw/refs/heads/main/wit/SKILL.md)
      - [`WIT-科学思考及写作skill.md`](WIT-科学思考及写作skill.md) — 中文完整 workflow;[直接下载](https://github.com/deltadbu/WIT-skill/raw/refs/heads/main/wit/references/WIT-%E7%A7%91%E5%AD%A6%E6%80%9D%E8%80%83%E5%8F%8A%E5%86%99%E4%BD%9Cskill.md)
      - [`WIT-Scientific-thinking-and-writing-skill.md`](WIT-Scientific-thinking-and-writing-skill.md) — English full workflow;[直接下载](https://github.com/deltadbu/WIT-skill/raw/refs/heads/main/wit/references/WIT-Scientific-thinking-and-writing-skill.md)
      
      如果希望长期使用 WIT,推荐直接 clone 整个仓库,而不是只下载一个文件:
      
      ```bash
      git clone https://github.com/deltadbu/WIT-skill.git
      ```
      
      真正可安装的 skill package 是仓库中的 [`wit/`](../) 子目录。保持这个目录整体不变,就能保证 `SKILL.md`、`references/`、`tests/` 和 `case-studies/` 之间的相对路径正确。
      
      ---
      
      ### 3.2 使用 `SKILL.md`:让 Agent 按 WIT 自动协同
      
      如果 agent 支持 Agent Skills,推荐把仓库中的 [`wit/`](../) 子目录——也就是真正可安装的 WIT skill package——安装或放入该 agent 能够发现的 skill 目录。不同平台的 skill 安装位置可能不同,应按照相应平台的规则配置;WIT 本身不假定某一个固定目录。
      
      Agent 发现 [`SKILL.md`](../SKILL.md) 后,首先读取其中的 metadata 和执行规则。它会把 `SKILL.md` 作为一个 **dispatcher + workflow controller + collaboration protocol**,主要负责:
      
      (1)判断当前任务是否应该使用 WIT;
      
      (2)识别当前需要的 mode,例如 **Open a question、Advance from a finding、Choose the next experiment、Review Results、Review Discussion、Stress-test a study、Deadline Mode**;
      
      (3)根据语言与任务,加载完整的 [`WIT-科学思考及写作skill.md`](WIT-科学思考及写作skill.md) 或 [`WIT-Scientific-thinking-and-writing-skill.md`](WIT-Scientific-thinking-and-writing-skill.md);
      
      (4)需要时再加载 skill package 内的 `tests/` 或 `case-studies/`(仓库路径分别为 [`wit/tests/`](../tests/) 和 [`wit/case-studies/`](../case-studies/))中的 supporting materials,而不是默认把所有材料一次性塞进 context;
      
      (5)执行 WIT 的 REWRITE decision loop、reasoning invariants 和 stop rule;
      
      (6)在人机协同中安排 **control transfer**:低 learning-value 的劳动可以由 AI 自动完成;高 learning-value 的 judgment nodes,应让研究者保持实质参与。
      
      因此,使用 [`SKILL.md`](../SKILL.md) 时,用户不需要每一步都手工指定 WIT 的全部流程。可以直接说:
      
      > Use WIT to analyze this finding and decide the next experiment.
      
      或者:
      
      > 按 WIT 检查这篇论文的 Results 和 Discussion。
      
      如果 agent 已经正确发现并加载 WIT,它应根据 [`SKILL.md`](../SKILL.md) 自动选择相应的 mode,而不是要求用户重新描述整套 WIT 流程。
      
      如果当前工具**不能自动发现 skill**,但能够读取 repo 中的文件,也可以显式要求:
      
      > 请先读取 [SKILL.md](../SKILL.md),再按照其中的 routing、human–LLM collaboration rules 和 stop rule 使用 WIT。需要中文完整 workflow 时,读取 [WIT-科学思考及写作skill.md](WIT-科学思考及写作skill.md)。
      
      核心原则是:
      
      > **`SKILL.md` 不是 WIT 理论正文的替代品,而是 agent 执行 WIT 的入口和控制程序。**
      
      ---
      
      ### 3.3 不支持 Skill Discovery 时:直接加载完整 WIT workflow
      
      如果只是使用普通 ChatGPT 对话,或者当前 agent 不支持自动发现 [`SKILL.md`](../SKILL.md),最简单的方法是直接加载完整的 WIT workflow。
      
      #### 在 ChatGPT 中
      
      下载并上传中文 [`WIT-科学思考及写作skill.md`](WIT-科学思考及写作skill.md),或英文 [`WIT-Scientific-thinking-and-writing-skill.md`](WIT-Scientific-thinking-and-writing-skill.md),然后说:
      
      > 请读取这个 WIT workflow,并在本次对话中按它工作。
      
      之后可以直接说:
      
      > 按 WIT 展开这个 finding:……
      
      或:
      
      > 按 WIT 检查这篇论文的 Results 和 Discussion。
      
      如果希望同时采用 agent-level 的 routing、human–LLM control transfer 和 stop rule,也可以把 [`SKILL.md`](../SKILL.md) 一并提供给 AI。
      
      如果开启新的对话,而这些文件没有自动带入,就需要重新提供文件或 GitHub 链接。
      
      #### 在 ChatGPT Work / 项目空间中
      
      可以把 [`SKILL.md`](../SKILL.md) 与中文 [`WIT-科学思考及写作skill.md`](WIT-科学思考及写作skill.md) 或英文 [`WIT-Scientific-thinking-and-writing-skill.md`](WIT-Scientific-thinking-and-writing-skill.md) 放进对应项目资料中,然后在项目开始时说:
      
      > 请读取 WIT 的 `SKILL.md` 和完整 workflow,并把它们作为本项目的人机协同科学思考与写作规则。
      
      这样可以让 WIT 与论文草稿、实验结果、项目文档一起长期使用。
      
      #### 在 VSCode / Codex / Copilot 中
      
      推荐直接 clone [WIT GitHub repo](https://github.com/deltadbu/WIT-skill)。若工具支持 Agent Skills,则把仓库中的 [`wit/`](../) 子目录作为可发现的 skill package,并让其从 [`wit/SKILL.md`](../SKILL.md) 自动发现和执行 WIT;若不支持,则在对话或项目 instruction 中明确要求:
      
      > Read [SKILL.md](../SKILL.md) first, then load the appropriate full WIT workflow and follow its routing, reasoning, human–LLM collaboration, and stop rules.
      
      核心原则是:
      
      > **支持 Skill Discovery:以 `SKILL.md` 为入口。**  
      > **不支持 Skill Discovery:直接加载完整 WIT workflow;需要更稳定的 routing 与协同控制时,再同时加载 `SKILL.md`。**
      
      ---
      
      ### 3.4 加载后如何调用
      
      #### 模式 A:开题——把问题打开
      
      输入:
      
      > 按 WIT 把这个问题打开:XXX。
      
      重点从以下方向展开:
      
      > **Whether / What / Why / How / When / To what extent**
      
      目标是把一个模糊问题展开成可研究的 **question space**。
      
      ---
      
      #### 模式 B:从一个 Finding 继续推进研究
      
      输入:
      
      > 这是一个 finding:XXX。按 WIT 展开。
      
      重点输出:
      
      (1)Fact;
      (2)1-hop Opinion;
      (3)Literature positioning;
      (4)新的 Why / How / What / When / Whether / To what extent 问题;
      (5)哪些当前可回答;
      (6)最值得补的实验或分析;
      (7)哪些当前不可回答;
      (8)2-hop Interpretation;
      (9)General Principle;
      (10)Potential Limitation(若与当前 story 相关);
      (11)Potential Future Study(若值得显式展开)。
      
      这是 WIT 最核心的日常使用方式:
      
      > **每得到一个重要 finding,就重新把问题打开一次。**
      
      ---
      
      #### 模式 C:检查 Results
      
      输入:
      
      > 按 WIT 检查 Results。
      
      重点检查:
      
      - subsection titles 连起来,是否能形成清楚、递进的 scientific story;
      - 每个 subsection 的存在是否有明确的逻辑功能,而不只是“又做了一个实验”;
      - 是否形成清楚的 **Fact → restrained 1-hop Opinion**;
      - **Question / Finding-driven** 的组织是否合适;如果是 method / system paper,**Component / Pipeline-driven** 的组织是否更自然;
      - 即使表面按 pipeline 组织,背后的 **Question → Test → Finding → Next Step** reasoning 是否仍然可恢复;
      - 是否有本来可以回答的问题被过早扔进 Discussion;
      - 是否遗漏关键 control、competing explanation、counterexample 或异常结果;
      - **Early figures 及其 captions:** 不应强制要求 Figure 1 本身概括全文。有些 method paper 中,Figure 1 更适合先介绍**问题、representation 或 motivation**,再在后续的 early figure 中给出方法 overview;AlphaDev 就是一个典型例子。真正需要检查的是:是否存在一个**较早出现的 overview figure——通常可以是 Figure 1,但不必强制是 Figure 1——**使 reader 在尚未阅读详细 Results 之前,就能掌握全文的 central idea、主要 components 和方法的 high-level operation。应检查 early figures 所承担的 reasoning function,而不是固定 figure number。
      - **Figure caption 既要说明 evidence,也可以给出 immediate takeaway:** 对于 Results figure,caption 不应只是机械地列出各 panel 或重复坐标轴信息。应以图中直接展示的 **Fact** 为主体,并可以加入一个 **restrained 1-hop Opinion**——即由该图直接支持的 immediate take-home message。不要在 caption 中写需要综合多个 findings 才成立的 2-hop interpretation、broad abstraction 或 general principle;这些通常应留在 Results 正文或 Discussion。一个简单检查是:**图里展示了什么?→ 这张图立即告诉我们什么?** 对于 overview / conceptual figure,caption 则可以重点解释 idea、workflow 和 components,不必强行加入结果性的 opinion。
      - **数学公式 / 计算机算法方法的自然语言解释与 concrete walkthrough:** 如果论文的核心方法主要由数学公式或计算机算法描述,不应只依赖 formal notation 或 pseudocode。应首先用**自然语言解释其 basic idea**:该公式或算法试图解决什么问题,关键项或关键步骤分别意味着什么,以及为什么要这样设计。然后再提供一个尽可能简单的 concrete example:给出公式中关键变量或关键项的具体数值,逐步展示计算过程;或者在一个很小的输入上,逐步展示算法中的重要步骤。优先选择**能够暴露核心 mechanism 的最小例子**。其目标是让 reader 先在概念上理解方法,再能够在头脑中**“执行”一次这个方法**。AlphaDev 是一个很典型的例子:用三个数的排序,就能够非常直观地展示算法的运行过程。
      - **方法类 paper 的 worked case study:** Results 中是否提供了一个简洁、representative 的 case,使 reader 可以从 input 出发,沿着关键 intermediate steps 一直看到 output?这个 case 应说明:**本文方法如何运行、为什么能够 succeed、为什么 existing methods 在同一个 case 上 failed 或变得 unreliable,以及究竟是哪一个 mechanism / design choice 造成了这种差异。** 优先选择**能够隔离方法本质差异的最简单 case**。Case study 的作用是解释 mechanism,而不是替代 benchmark-level evidence。
      - **Difference-focused benchmark analysis:** accuracy、precision、recall、AUC、correlation、success rate 等 aggregate metrics 可以说明方法总体上是否更好、好多少,却通常不能解释 why。应检查作者是否从 benchmark 中找出 proposed method 与 existing methods 表现差异最大或具有系统性差异的 cases / subsets,并进一步分析:**performance gap 来自哪里?existing methods 为什么在这些 cases 上失败?本文方法为什么能够成功?** 应优先寻找 systematic pattern of differential performance,而不是只挑选少数漂亮的 cherry-picked examples。
      
      对于 method paper,这几类 evidence / explanation 的功能彼此互补:
      
      > **Early figures 负责给 reader 定向:Figure 1 可以先介绍问题,而后续较早出现的 overview figure 再解释 central idea 和 high-level operation;自然语言解释加 minimal concrete walkthrough 让数学或算法方法真正变得 operationally understandable;benchmark statistics 说明方法是否有效、提高多少;worked case 解释为什么本文方法 succeed、为什么 existing methods fail;difference-focused analysis 解释 improvement 从哪里来,以及这个 mechanism 是否具有系统性。**
      
      原则:
      
      > **Results should expose the logical progression of the scientific story, not obey a single title format.**
      
      ---
      
      #### 模式 D:检查 Discussion
      
      输入:
      
      > 按 WIT 检查 Discussion。
      
      先检查 **core functions**:
      
      - 是否把多个局部 findings / 1-hop Opinions 综合成 integrated interpretation;
      - 是否超越逐条重复 Results;
      - 是否进一步说明 broader meaning,必要时上升到 General Principle。
      
      再检查 **optional extensions**(仅在研究需要时):
      
      - 是否值得从 findings 再打开新的 question space;
      - 是否存在真正需要明确写出的 boundary / limitation;
      - Future Studies 是否能够针对 unresolved questions,而不是例行公事。
      
      原则:
      
      > **Discussion 的核心功能是 interpretation and abstraction;New Questions、Limitations 和 Future Studies 是有价值的扩展,但不是每篇论文都必须显式出现的固定段落。**
      
      ---
      
      #### 模式 E:寻找下一步实验
      
      输入:
      
      > 按 WIT 判断下一步最值得做什么实验。
      
      优先考虑:
      
      (1)会改变 central claim 的问题;
      (2)reviewer 最可能提出的关键挑战;
      (3)competing explanations;
      (4)boundary conditions;
      (5)低成本、高信息量的实验。
      
      ---
      
      #### 模式 F:Deadline Mode
      
      输入:
      
      > Deadline 很近,按 WIT 帮我收束。
      
      输出:
      
      - 必须做;
      - 最好做;
      - 可以不做;
      - 应写入 Limitation;
      - 应留给 Future Study;
      - central claim 是否需要收缩。
      
      WIT 的日常使用可以压缩成一句话:
      
      > **支持 Skill Discovery 时,让 agent 从 `SKILL.md` 进入 WIT;不支持时,直接加载完整 workflow。然后从问题开始,每得到一个 finding 再产生问题;能回答的继续做,不能回答的先判断是否重要,再决定是否进入 Discussion、Limitations 或 Future Studies。**
      
      ## 4. REWRITE:核心研究循环
      
      WIT 内部通过 **REWRITE** 循环推进研究:
      
      > **Research Question → Examine Literature → Work / Experiment → Read Finding → Interrogate Finding → Test Answerability → Extend / Exit**
      
      七个步骤不是彼此独立的模块,而是一个连续循环。
      
      ### 4.1 R — Research Question:研究问题
      
      “开题”不只是确定一个题目,而是:
      
      > **把一个单点问题展开成一个 scientific question space。**
      
      这里不应只是机械地罗列疑问词。更本质的做法,是把问题展开成六类彼此尽量正交的科学维度。
      
      #### 4.1.1 Whether → Existence:现象是否存在?
      
      首先问:
      
      > **这个现象真的存在吗?**
      
      这是最基础的 existence question。
      
      例如:
      
      > MSFold 是否真的比标准 decoding 更容易恢复 alternative conformations?
      
      这一层的目标是确认:
      
      - 现象是否稳定存在;
      - 是否具有统计显著性;
      - 是否可重复;
      - 是否只是偶然结果。
      
      #### 4.1.2 What → Determinants:哪些因素决定它?
      
      当现象存在后,进一步问:
      
      > **哪些变量、因素或属性决定这个现象的强弱与出现?**
      
      例如:
      
      > 哪些 protein properties 决定 MSFold 是否能够恢复 alternative conformations?
      
      可能的 determinants 包括 protein size、conformational change type、sequence identity、token-space diversity、sampling budget 等。
      
      这一层关注:
      
      > **What determines the outcome?**
      
      #### 4.1.3 Why → Cause:为什么会发生?
      
      Why 问的不是“具体过程怎么发生”,而是:
      
      > **什么原因导致了这个现象?**
      
      即寻找 causal driver。
      
      例如:
      
      > 为什么标准 decoding 难以恢复 alternative conformations?
      
      可能原因包括:
      
      - representation 中根本没有编码这些 states;
      - search 被限制在局部 mode;
      - decoding objective 偏好 dominant state。
      
      因此:
      
      > **Why asks what causes the phenomenon.**
      
      #### 4.1.4 How → Mechanism:原因通过什么机制产生结果?
      
      How 与 Why 不同。
      
      Why 已经回答:
      
      > **是什么原因导致了现象?**
      
      How 则继续问:
      
      > **这个原因通过什么过程、路径或机制产生结果?**
      
      例如:
      
      > 如果原因是标准 decoding 容易陷入局部 mode,那么 parallel tempering 是如何帮助跨越 token-space barrier、进入 alternative states 的?
      
      因此:
      
      > **Cause → Mechanism → Outcome**
      
      以及:
      
      > **Why asks what causes it; How asks through what mechanism the cause produces the effect.**
      
      #### 4.1.5 When → Boundary Conditions:什么条件下成立或失效?
      
      这里的 When 不应只理解为时间。
      
      它代表更一般的:
      
      > **Under what conditions does the conclusion hold or fail?**
      
      例如:
      
      - 在哪些 protein classes 上成立?
      - 在哪些 conformational change types 上失效?
      - 在 low-data regime 还是 high-data regime 更明显?
      - 在什么 sequence identity 范围内成立?
      - 在哪些组织、细胞类型、空间区域中成立?
      
      原先可以归入 “Where” 的很多问题,本质上也属于这一维度,因此不再单独设置 Where。
      
      这一层关注:
      
      > **Boundary conditions**
      
      #### 4.1.6 To what extent → Magnitude:效应有多大?
      
      最后问:
      
      > **这个效应到底有多强?范围有多宽?**
      
      例如:
      
      - success rate 提高多少?
      - improvement 是否只在少数 samples 中出现?
      - 能恢复多大幅度的 conformational change?
      - effect size 是否具有实际意义,而不只是统计显著?
      
      这一层关注:
      
      > **Magnitude / effect size / range**
      
      因此,“把问题打开”可以压缩为六个科学维度:
      
      > **Existence → Determinants → Cause → Mechanism → Boundary Conditions → Magnitude**
      
      对应:
      
      > **Whether → What → Why → How → When → To what extent**
      
      这六维并不要求每个 project 都全部回答,而是用于系统检查:
      
      > **当前研究的问题空间中,还有哪些重要维度没有被打开?**
      
      
      #### 4.1.7 一个更大众化的例子:用 AlphaGo 把六维 question space 打开
      
      MSFold 的例子适合蛋白质结构研究,但对非本领域读者不够直观。下面用 AlphaGo 的经典 Nature 论文 *Mastering the game of Go with deep neural networks and tree search* 说明六个维度。
      
      AlphaGo 面对的核心困难很清楚:围棋的搜索空间极大,论文估计围棋平均 branching factor 约为 250、典型 game depth 约为 150,因此 exhaustive search 不可行。AlphaGo 的核心思想是同时压缩搜索树的**宽度**和**深度**:policy network 用来优先选择有希望的落子,value network 用来评估局面,再与 Monte Carlo tree search (MCTS) 结合。
      
      这里最重要的不是把六个疑问词硬套在 AlphaGo 上,而是看:
      
      > **同一个研究成果,可以沿六种不同的 scientific dimensions 被继续打开。**
      
      ##### Whether → Existence:AlphaGo 真的能达到专业围棋水平吗?
      
      这是最直接的 existence question:
      
      > **深度神经网络 + tree search 的组合,是否真的能够解决此前计算机围棋长期无法突破的问题?**
      
      论文给出了非常直接的证据:
      
      - AlphaGo 对其他 Go programs 的胜率达到 **99.8%**;
      - 对欧洲围棋冠军 Fan Hui 的正式比赛结果为 **5:0**。
      
      因此这一层回答的是:
      
      > **The phenomenon exists:这种方法确实能够达到此前计算机围棋没有达到的水平。**
      
      注意,Whether 不是问“为什么成功”,只是先确认:
      
      > **它到底成功了没有?**
      
      ##### What → Determinants:哪些因素决定 AlphaGo 的棋力?
      
      确认 AlphaGo 很强之后,下一步自然不是立刻问机制,而是先问:
      
      > **什么因素决定它有多强?**
      
      论文实际上分析了多个 determinants。
      
      例如:
      
      - supervised policy network 对 expert moves 的预测准确率达到 **57.0%**,高于当时其他研究的 **44.4%**;
      - policy prediction accuracy 的小幅提高会带来明显的 playing-strength 提高;
      - reinforcement learning 后的 policy network 在 head-to-head 中对 supervised policy network 的胜率超过 **80%**;
      - 不使用 search 时,RL policy network 对 Pachi 的胜率达到 **85%**;
      - value network、rollout policy、policy network 以及 search budget 都会影响最终棋力。
      
      因此 What 问的是:
      
      > **Which components or variables determine performance?**
      
      而不是:
      
      > **这些因素为什么有效?**
      
      ##### Why → Cause:为什么传统方法难以解决围棋,而 AlphaGo 能突破?
      
      Why 关注 causal explanation。
      
      论文一开始就指出两个核心困难:
      
      - **search breadth 太大**:每个局面有大量可能落子;
      - **search depth 太深**:一盘棋需要搜索很长的未来序列。
      
      因此 exhaustive search 在围棋上不可行。
      
      AlphaGo 能够突破,一个核心 causal explanation 是:
      
      > **它不再平等地搜索所有可能性,而是利用学习到的 policy 和 value,把有效搜索空间大幅压缩。**
      
      更具体地说:
      
      - policy network 降低有效 **breadth**;
      - value network 降低有效 **depth**。
      
      所以:
      
      > **Why = Cause:成功的原因是什么?**
      
      这里得到的是一个比“用了 deep learning”更本质的解释:
      
      > **AlphaGo succeeds because learning makes an otherwise intractable search problem tractable enough to search.**
      
      ##### How → Mechanism:policy/value network 如何真正改变 MCTS?
      
      知道“原因是缩小有效搜索空间”之后,How 才继续问:
      
      > **这个原因具体通过什么算法过程产生效果?**
      
      AlphaGo 的 mechanism 可以进一步拆开:
      
      (1)**Policy network guides selection and expansion**
      
      在搜索树中,不再平均探索所有合法动作,而是优先探索 policy network 认为更可能的 moves。
      
      (2)**Value network evaluates leaf positions**
      
      搜索到 leaf node 后,不必每次都完整模拟到终局;value network 可以直接估计该局面的 winning probability。
      
      (3)**Rollout provides another evaluation**
      
      fast rollout policy 继续模拟到游戏结束,得到另一种局面估值。
      
      (4)**MCTS backs the information up**
      
      value-network evaluation 和 rollout result 被组合,并沿搜索路径向上 backup,更新 action values 和 visit counts。
      
      因此:
      
      > **Cause:学习缩小了有效搜索空间。**  
      > **Mechanism:policy-guided search + value evaluation + rollout + MCTS backup 具体实现了这种压缩。**
      
      这正好体现:
      
      > **Why asks what causes the success; How asks through what computational mechanism that cause produces stronger play.**
      
      ##### To what extent → Magnitude:AlphaGo 到底强了多少?
      
      Magnitude question 不满足于:
      
      > “AlphaGo 很强。”
      
      而要问:
      
      > **强多少?计算效率提高多少?各个组件带来多大提升?**
      
      AlphaGo 论文给出了多个量化尺度:
      
      - 对其他 Go programs:**99.8% win rate**;
      - 对 Fan Hui:**5–0**;
      - RL policy 对 SL policy:**>80% win rate**;
      - RL policy 在没有 search 的情况下对 Pachi:**85% win rate**;
      - 单次 value-network evaluation 的准确度接近使用 RL policy 的 Monte Carlo rollouts,但计算量约少 **15,000 倍**。
      
      这些数字回答的不是“有没有作用”,而是:
      
      > **How large is the effect?**
      
      因此,Whether 与 To what extent 的区别也非常清楚:
      
      > **Whether:有没有?**  
      > **To what extent:有多大?**
      
      ##### When → Boundary Conditions:AlphaGo 在什么条件下成立、什么时候会失效?
      
      这是一个特别值得注意的例子。
      
      AlphaGo 论文已经证明:
      
      - 方法在 full-sized Go 上有效;
      - 对多种计算机程序有效;
      - 对一位职业级人类棋手有效。
      
      这些给出了部分 boundary evidence。
      
      但论文并没有系统画出完整的 boundary map。
      
      沿 WIT 继续追问,可以产生:
      
      - 当 search budget 大幅降低时,优势是否仍然存在?
      - policy network 较弱时,MCTS 还能否补偿?
      - value estimate 出现系统性偏差时,search 会在什么时候失效?
      - 对不同风格、不同水平的人类棋手,效果是否一致?
      - 这种 **learning + search** 的原则能否迁移到围棋之外的其他大规模决策问题?
      
      这些问题不是原论文都已经回答了,而是由 AlphaGo 的 finding 自然生成的:
      
      > **Under what conditions does the conclusion hold or fail?**
      
      这正是 Boundary Conditions 的意义。
      
      ---
      
      用 AlphaGo 可以非常直观地看到六维之间的区别:
      
      | Dimension | AlphaGo 中的问题 |
      |---|---|
      | **Whether → Existence** | 深度网络 + search 是否真的能达到专业围棋水平? |
      | **What → Determinants** | policy accuracy、RL、value network、search 等哪些因素决定棋力? |
      | **Why → Cause** | 为什么这种方法能够突破传统计算机围棋? |
      | **How → Mechanism** | policy/value network 如何嵌入 MCTS 并改变搜索过程? |
      | **When → Boundary Conditions** | 在什么 search budget、opponent、模型质量和任务条件下仍成立或失效? |
      | **To what extent → Magnitude** | 胜率、棋力和计算效率到底提高了多少? |
      
      因此,一个 finding:
      
      > **AlphaGo defeated a professional Go player.**
      
      并不是研究的终点。  
      沿六维 scientific question space 展开后,它立刻变成:
      
      > **有没有 → 由什么决定 → 为什么 → 如何实现 → 在什么条件下成立 → 到底强多少**
      
      这就是 WIT 所说的:
      
      > **把问题打开。**
      
      **Reference:** Silver, D. et al. *Mastering the game of Go with deep neural networks and tree search*. Nature 529, 484–489 (2016), doi:10.1038/nature16961.
      
      
      ### 4.2 E — Examine the Literature:审视文献
      
      文献不是 Introduction 的装饰,而是科研推理的坐标系。
      
      WIT 设置两个文献检查点。
      
      #### 4.2.1 Literature Checkpoint 1:研究开始之前
      
      在确定核心 scientific question 后,检查:
      
      (1)这个问题是否已经被回答?
      (2)已有工作提供了哪些解释?
      (3)存在哪些 competing hypotheses?
      (4)已知的 boundary conditions 是什么?
      (5)当前工作相对于已有研究究竟新增什么?
      
      这里的目标不是“找够引用”,而是判断:
      
      > **Novelty 在哪里?**
      
      以及:
      
      > **当前研究真正需要区分哪些已有解释?**
      
      #### 4.2.2 Literature Checkpoint 2:得到重要 Finding 之后
      
      每得到一个重要 finding,再回到文献,问:
      
      > **这个 finding 与已有认识是什么关系?**
      
      可以分成:
      
      - **Confirm**:验证已有结论;
      - **Contradict**:与已有结论冲突;
      - **Refine**:对已有结论增加边界、条件或更精细解释;
      - **Extend**:把已有结论推广到新的任务、体系或场景;
      - **Reframe**:改变问题本身的理解方式。
      
      这一检查点直接决定 Discussion 的深度。
      
      真正值得讨论的通常不是:
      
      > “我们的结果与某文一致。”
      
      而是:
      
      > **我们的发现改变、修正或扩展了什么认识?**
      
      ---
      
      ### 4.3 W — Work / Experiment:开展研究与实验
      
      每个实验都应当对应一个明确问题。
      
      不要因为“论文通常需要 ablation”就做 ablation;也不要因为“别人画 t-SNE”就画 t-SNE。
      
      先问:
      
      > **这个实验究竟要回答什么问题?**
      
      理想结构是:
      
      > **Question → Experiment → Data → Finding**
      
      实验是回答问题的工具,不是论文 Results 的组织单位。
      
      ---
      
      ### 4.4 R — Read the Finding:读懂 Finding
      
      得到结果以后,不要立刻跳到下一项实验。首先区分三个层级。
      
      #### 4.4.1 Data:数据
      
      原始观察或定量结果。
      
      #### 4.4.2 Finding:发现
      
      数据直接支持的定性陈述。
      
      #### 4.4.3 1-hop Opinion:一步推论
      
      在不远离数据的前提下,向前走一步。
      
      因此,一个 Results subsection 可以压缩成非常实用的 **Fact–Opinion** 结构:
      
      > **Question → Experiment → Fact → 1-hop Opinion**
      
      其中:
      
      #### 4.4.4 Fact
      
      实验或分析直接得到的事实,包括数据、比较、观察到的现象和统计结果。
      
      例如:
      
      > Method A 在 distribution shift 下显著优于 Method B。
      
      #### 4.4.5 1-hop Opinion
      
      对 Fact 做出的**一步解释**。它可以包含作者判断,但必须紧贴当前结果,不能直接跨越到更一般的理论层面。
      
      例如:
      
      > 这些结果提示,X 可能提高了模型在 distribution shift 下的稳健性。
      
      因此:
      
      > **Results = Fact + 1-hop Opinion**
      
      Results subsection 的结尾应当回答:
      
      > **这个具体 Fact 意味着什么?**
      
      但这个 Opinion 只能离 Fact 一步远。
      
      核心原则:
      
      > **Results 可以有 opinion,但只能是离 fact 一步远的 opinion。**
      
      不要在这里直接跳到领域级 general principle。
      
      ---
      
      #### 4.4.6 异常结果与偶然发现:Anomaly Branch
      
      真实科研不是完全线性的。实验经常会出现与假设相反的结果、异常样本、unexpected subgroup、看似失败但可重复的现象,或与已有理论不一致的结果。
      
      因此每个重要实验之后,都应检查:
      
      > **结果是否符合原先预期?**
      
      如果符合,进入正常 WIT / REWRITE 流程。
      
      如果不符合,进入异常结果分支:
      
      ##### (1)是技术错误吗?
      
      检查数据处理错误、实现 bug、测量误差、数据泄漏、batch effect、样本污染、统计假象等。
      
      如果是:
      
      > 修正后重新实验。
      
      ##### (2)是随机噪声吗?
      
      问:
      
      > 是否可重复?
      
      如果不可重复:
      
      > 暂不作为主要 finding。
      
      ##### (3)是稳定、可重复的异常现象吗?
      
      如果是,不要强行把它塞回原 hypothesis。
      
      应当问:
      
      > **这个异常是否意味着原来的 scientific question 不完整,甚至问错了?**
      
      此时允许:
      
      > **Unexpected Finding → Rewrite the Question**
      
      也就是说,REWRITE 不仅“整理已有成果”,还允许研究问题被新发现重新定义。
      
      ---
      
      ### 4.5 I — Interrogate the Finding:追问 Finding
      
      每个重要 finding 都应该继续产生问题。
      
      系统追问:
      
      - **Why?**
      - **How?**
      - **What?**
      - **Whether?**
      - **When?**
      - **Where?**
      - **To what extent?**
      
      一个好的 finding 应该能够打开新的 question space。
      
      ---
      
      
      #### 4.5.1 Competing Hypotheses:不要过早接受单一解释
      
      当一个 finding 出现后,不要只问:
      
      > **“为什么会这样?”**
      
      还应强制提出多个可能解释:
      
      > **Finding → Hypothesis A / Hypothesis B / Hypothesis C**
      
      例如:
      
      > Finding:MSFold 在未见蛋白上仍然优于 baseline。
      
      可能解释包括:
      
      - **H1:** ESM3 的 token space 确实编码了可泛化的 alternative conformations;
      - **H2:** improvement 主要来自更大的 sampling budget;
      - **H3:** benchmark 中某些结构类型更容易被 parallel tempering 恢复。
      
      真正有价值的下一步不是选择“最顺眼”的解释,而是问:
      
      > **什么实验能够区分这些 competing hypotheses?**
      
      因此,机制研究应优先寻找 **discriminating experiment**,而不是仅仅继续积累支持性证据。
      
      ---
      
      #### 4.5.2 Falsification / Counterexample Check:主动攻击自己的结论
      
      对每个重要 conclusion,再反过来问:
      
      > **什么结果会推翻这个 conclusion?**
      
      以及:
      
      > **最可能的 counterexample 是什么?**
      
      例如:
      
      > Claim:结构信息提高了模型的 OOD robustness。
      
      不要只继续找支持这个 claim 的数据,还应问:
      
      - 是否存在加入结构信息后性能下降的数据集?
      - 在 sequence diversity 已经很高时,这个优势是否消失?
      - improvement 是否其实来自额外参数量,而不是结构信息本身?
      
      这一检查的目的不是“证明自己错”,而是确定:
      
      > **这个 claim 的证据边界在哪里?**
      
      如果一个 conclusion 没有明确的潜在反证条件,它往往还没有被定义得足够科学。
      
      ---
      
      
      ##### 如果找到了 counterexample,怎么办?
      
      不要立刻宣布“结论错了”。先判断反例属于哪一类:
      
      - **技术或数据问题**:如 bug、measurement error、data leakage、sample contamination。  
        → 修正后重新实验。
      - **偶然噪声**:结果不可重复。  
        → 暂不推翻 conclusion,但应降低 confidence。
      - **稳定、可重复的 counterexample**:  
        → 这通常意味着原 claim 需要被修正,而不是简单丢弃。
      
      稳定反例通常有三种价值:
      
      (1)**收缩 claim**
      
      例如:
      
      > “X always improves Y”
      
      收缩为:
      
      > “X improves Y under conditions A and B.”
      
      (2)**发现 boundary condition**
      
      反例可能告诉我们:
      
      > **这个结论什么时候不成立?**
      
      这往往比继续堆支持性结果更有科学价值。
      
      (3)**重写 hypothesis / principle**
      
      如果反例击中了核心机制,就需要重新检查 central claim,甚至重写 research storyline。
      
      因此:
      
      > **Counterexample → Boundary Condition → Better Claim**
      
      一个好的反例未必削弱论文,反而可能让结论更精确、更可信。
      
      ##### 如果找不到 counterexample,怎么办?
      
      也不能因此说:
      
      > **“Conclusion 已被证明。”**
      
      应继续检查两个问题:
      
      (1)**我们是否真的进行了足够强的 falsification?**
      
      是否主动测试了:
      
      - 最不利条件;
      - extreme cases;
      - OOD data;
      - negative controls;
      - alternative explanations;
      - 可能失败的 subgroups。
      
      (2)**这个 claim 是否具有可证伪性?**
      
      如果无论出现什么结果,都可以通过改写措辞让 claim 保持成立,那么这个 claim 可能定义得过于含糊。
      
      应明确写出:
      
      > **什么结果出现时,我会承认这个 claim 不成立?**
      
      因此完整流程是:
      
      > **Claim → Potential Falsifier → Test → Counterexample?**
      
      如果找到稳定反例:
      
      > **Revise Claim / Identify Boundary / Rewrite Hypothesis**
      
      如果没有找到:
      
      > **Claim gains support, but is not “proven.”**
      
      如果当前无法检验潜在 falsifier:
      
      > **记录为 unresolved question → Limitation / Future Study**
      
      可以把这一原则压缩成:
      
      > **A failed falsification strengthens a claim; a successful falsification sharpens it.**
      
      
      ### 4.6 T — Test Answerability:检验可回答性
      
      这是 WIT 中 REWRITE loop 最关键的决策点。
      
      对于每个新问题,问:
      
      > **当前研究能不能通过额外分析或实验回答这个问题?**
      
      #### 4.6.1 如果答案是 YES
      
      不要把它过早写进 Discussion。
      
      应该继续做:
      
      > **New Question → New Experiment / Analysis → New Finding → New Results**
      
      例如,若发现 A 在 OOD 上优于 B,接着问“这种优势是否在不同 protein family 中一致?”而现有数据已包含多个 family,就应该立即分析,并把结果做成新的 Results subsection。
      
      同样,如果“哪个模块带来主要提升?”可以通过 ablation 回答,就不应该写成 Future Work。
      
      #### 4.6.2 如果答案是 NO
      
      问题才进入 Discussion:
      
      > **New Question → Why current study cannot answer → Limitation → Future Study**
      
      ---
      
      
      #### 4.6.3 Information Gain:下一实验不是“最容易做”,而是“最能改变判断”
      
      当多个问题都可以继续实验时,不应只按方便程度选择。
      
      更好的原则是:
      
      > **优先选择 information gain 最大、最能区分 competing hypotheses 的实验。**
      
      可以问:
      
      - 如果实验结果为 A,我的判断会改变多少?
      - 如果结果为 B,我是否会修改 central claim?
      - 这个实验能否区分两个目前都合理的解释?
      - 它只是增加更多 supporting data,还是会真正减少 uncertainty?
      
      例如:
      
      - 再增加一个相似 benchmark,可能只让 confidence 从 0.85 变成 0.88;
      - 一个针对 competing hypotheses 的 control experiment,可能直接决定机制解释 A 还是 B。
      
      通常后者更值得优先做。
      
      因此:
      
      > **Next experiment ≠ easiest experiment**  
      > **Next experiment = highest-value information test**
      
      ---
      
      ### 4.7 E — Extend / Exit:拓展或收束
      
      如果新问题当前可回答:
      
      > **New Question → New Experiment / Analysis → New Results**
      
      如果当前不可回答或超出研究范围:
      
      > **New Question → Discussion → Limitation → Future Study**
      
      ## 5. Results:呈现 scientific story 的逻辑推进
      
      Results 的强规则不是:
      
      > **每个 subsection 必须按 scientific question 命名。**
      
      更本质的要求是:
      
      > **Results should expose the logical progression of the scientific story.**
      
      也就是说,读者应能够看出:
      
      > **为什么做这一部分 → 得到了什么 evidence → evidence 支持什么局部结论 → 为什么下一部分自然出现。**
      
      ### 5.1 两种都合理的 Results 组织方式
      
      #### (1)Question / Finding-driven
      
      更适合 discovery、mechanism、hypothesis-testing 型研究。
      
      例如标题可以直接写主要 answer:
      
      - “模块 X 是性能提升的主要来源”
      - “性能优势在 distribution shift 下仍然保持”
      - “学到的 representation 更好地区分不同功能状态”
      
      这种写法的优点是:
      
      > **标题直接告诉读者学到了什么。**
      
      #### (2)Component / Pipeline-driven
      
      对于 method、system、engineering 型论文,按组件或 pipeline stage 组织完全可以成立。
      
      AlphaGo 的 Nature 论文就是一个重要 counterexample。其主要研究段落依次是:
      
      - Supervised learning of policy networks
      - Reinforcement learning of policy networks
      - Reinforcement learning of value networks
      - Searching with policy and value networks
      - Evaluating the playing strength of AlphaGo
      
      这些标题明显是 **component / pipeline-driven**,而不是把每个标题写成一个 scientific answer。
      
      但它们连起来形成了很清楚的 storyline:
      
      > **learn a policy → improve the policy by self-play → learn a value function → integrate policy and value into search → evaluate the complete system**
      
      因此:
      
      > **Pipeline-driven ≠ technique list。**
      
      真正的问题是:
      
      > **这些组件是否形成一个有逻辑因果关系的 progression?**
      
      而不是标题里有没有出现 question / finding。
      
      ### 5.2 Reasoning structure 与 surface prose 要分开
      
      一个 Results subsection 背后的 reasoning,通常可以还原为:
      
      > **Motivation / Question → Experiment / Analysis → Fact → 1-hop Opinion → Next Question / Next Step**
      
      但这不意味着正文必须机械地逐项写出这些句子。
      
      AlphaGo 的一些段落直接用类似:
      
      > “The second stage of the training pipeline aims at ...”
      
      来推进,而没有显式写:
      
      > “We next asked whether ...”
      
      这种 surface prose 完全没有问题,只要背后的 reasoning chain 是清楚的。
      
      因此:
      
      > **WIT 要求 reasoning 可恢复,不要求 prose 模板化。**
      
      ### 5.3 Fact + restrained 1-hop Opinion:这条仍然是强规则
      
      Results 可以解释,但解释要紧贴 evidence。
      
      原则:
      
      > **Results:Fact → restrained 1-hop Opinion**
      
      例如 AlphaGo 的 tournament results 给出非常强的 Fact:系统几乎击败了所有比较的 Go programs,并在正式比赛中击败职业棋手。
      
      作者由此做的是局部、直接的能力解释,而没有立刻跳到:
      
      > “This reveals a universal principle of intelligence.”
      
      又例如,在 position evaluation 的比较中,value network 与 rollout 的组合效果优于单独使用任一者,于是作者得到一个很近的解释:
      
      > **the two position-evaluation mechanisms are complementary.**
      
      这正是:
      
      > **Fact → 1-hop Opinion**
      
      一个判断标准是:
      
      > **如果删掉当前 subsection 的数据,这个 opinion 是否还站得住?**
      
      如果不能,通常说明它仍然是 Results 允许的局部解释;  
      如果需要整篇论文多个 findings 才能支持,就更可能属于 Discussion。
      
      ### 5.4 Results titles 的真正检查方法:能否组成一篇“小 essay”?
      
      把 Results subsection titles 全部抽取出来,按顺序读一遍。
      
      问:
      
      > **它们能否独立讲出 scientific story 是怎样推进的?**
      
      这个 story 可以是:
      
      > **Question → Finding → Mechanism → Boundary**
      
      也可以像 AlphaGo:
      
      > **Component A → Component B → Integration → System Evaluation**
      
      因此,小 essay test 检查的是:
      
      > **logical progression**
      
      而不是:
      
      > **标题形式是否统一。**
      
      ### 5.5 初学者模板:Results Subsection
      
      对于经验较少的研究者,可以先用 reasoning template 思考,再决定最终 prose 怎么写。
      
      #### (1)Motivation / Scientific Question
      
      > 为什么需要这一部分?我们想知道:________________________。
      
      #### (2)Experiment / Analysis
      
      > 为了回答或推进这个问题,我们:________________________。
      
      #### (3)Fact
      
      > 数据直接显示:________________________。
      
      #### (4)1-hop Opinion
      
      > 这些结果局部地提示:________________________。
      
      #### (5)Next Question / Next Step
      
      > 因此下一步自然需要:________________________。
      
      如果是 method / system paper,还可以问:
      
      > **这个 subsection 在整个 pipeline 中承担什么不可替代的逻辑功能?**
      
      模板是用于暴露 reasoning,不是要求最终文章照着五句话写。
      
      ---
      
      ## 6. Discussion:Core Functions + Optional Extensions
      
      AlphaGo 对 WIT 的一个重要修正是:
      
      > **好的 Discussion 不一定显式包含 New Questions、Limitations 和 Future Studies。**
      
      因此,Discussion 应区分:
      
      > **Core Functions(强规则)**  
      > 与  
      > **Optional Extensions(按研究需要展开)**
      
      ### 6.1 Core Function 1:Integrated Interpretation
      
      Discussion 首先不能只是逐条重复 Results。
      
      它需要把多个局部 findings / 1-hop Opinions 综合起来,回答:
      
      > **Taken together,这些结果整体意味着什么?**
      
      可以写成:
      
      > **Multiple local findings → Integrated Interpretation**
      
      如果多个 findings 的综合需要比单个 subsection 多走一步,也可以理解为:
      
      > **multiple 1-hop Opinions → 2-hop Interpretation**
      
      但这里的“2-hop”是 reasoning depth 的描述,不是要求第一段必须机械使用某种句式。
      
      ### 6.2 Core Function 2:Broader Meaning / General Principle
      
      在 integrated interpretation 之后,Discussion 通常还需要回答:
      
      > **为什么这些结果值得超越当前实验本身去理解?**
      
      可能的上升方式包括:
      
      - deeper causal interpretation;
      - 与已有 paradigm / baseline 的概念性比较;
      - transferable mechanism;
      - broader implication;
      - General Principle。
      
      例如 MSFold 中可以上升为:
      
      > **Model capability is jointly determined by representation and search.**
      
      但并非每篇论文都必须提出一个宏大的“general principle”。证据只支持到哪里,就上升到哪里。
      
      ### 6.3 AlphaGo 的反向验证:Discussion 不需要固定六段式
      
      AlphaGo 的 Discussion 很短,但逻辑非常完整。
      
      它大致完成了三件事:
      
      #### (1)Integrated Interpretation
      
      把 policy network、value network、reinforcement learning 和 tree search 合起来解释为一个完整 system,而不是再逐条汇报胜率。
      
      #### (2)Deeper Meaning
      
      通过与传统高强度 search system 的比较,强调 AlphaGo 的关键不是“看更多 positions”,而是:
      
      > **learn where to search and how to evaluate.**
      
      这已经从组件结果上升到更深的 computational interpretation。
      
      #### (3)Beyond This Study
      
      最后把 Go 看成大规模 decision/search problem 的代表,并讨论这种 learning + search 思路对其他难解 AI problems 的启发。
      
      这已经完成:
      
      > **this study → broader meaning / abstraction**
      
      但 AlphaGo 并没有显式写:
      
      > New Questions → Limitations → Future Studies
      
      因此,AlphaGo 是一个很重要的 counterexample:
      
      > **Discussion 的 reasoning functions 可以完整,而 surface structure 不必包含固定的 limitations/future-work 段落。**
      
      ### 6.4 Optional Extension 1:Raise New Questions
      
      当新的 scientific questions 对理解研究边界或开启下一研究阶段有价值时,可以继续:
      
      > **Integrated Interpretation / General Principle → New Question Space**
      
      此时优先使用六个 scientific dimensions:
      
      > **Existence → Determinants → Cause → Mechanism → Boundary Conditions → Magnitude**
      
      也就是:
      
      > **Whether → What → Why → How → When → To what extent**
      
      但这一步是为了**继续研究**,不意味着这些问题都必须写进当前论文 Discussion。
      
      ### 6.5 Optional Extension 2:Limitations
      
      当一个重要 unresolved question 的确受到当前 study design / data / scope 限制时,Limitation 很有价值。
      
      真正的 limitation 仍然是:
      
      > **一个限制,使当前研究无法回答由自身 findings 引出的重要问题。**
      
      不是:
      
      > **随便列出一些“我们没做的事情”。**
      
      但如果论文并不需要一个独立 limitation 段落才能诚实界定 claim,也不应为了模板完整而强行添加。
      
      ### 6.6 Optional Extension 3:Future Studies
      
      Future Study 最有价值的情况是:
      
      > **它明确回答一个当前无法回答、但由本研究自然产生的 unresolved question。**
      
      因此:
      
      > **Finding → New Question → Limitation → Future Study**
      
      仍然是一条很好的科研推理链。
      
      但它是:
      
      > **research-planning logic**
      
      不一定必须变成:
      
      > **paper-surface paragraph**
      
      ### 6.7 Discussion 的判断标准
      
      比“有没有写 limitations/future work”更重要的是:
      
      - 是否超越 Results repetition?
      - 是否形成 integrated interpretation?
      - 是否给出了与 evidence 强度匹配的 broader meaning?
      - abstraction 是否过度?
      - 如果提出 general principle,是否有多个 findings 共同支撑?
      - 如果存在重要 boundary / unresolved question,是否诚实处理?
      - Discussion 的结束是否留下清楚的 take-home message?
      
      因此:
      
      > **Discussion 的核心:Interpretation → Broader Meaning**
      
      必要时再扩展:
      
      > **→ New Questions → Boundary / Limitations → Future Studies**
      
      ### 6.8 初学者模板:Discussion
      
      先完成 core,再决定是否需要 optional extension。
      
      #### Core A:Integrated Interpretation
      
      > 本研究多个 findings 共同表明:________________________。
      
      #### Core B:Broader Meaning
      
      > 更深一层,这些 findings 意味着:________________________。
      
      > 在 evidence 允许的范围内,更一般地可以理解为:________________________。
      
      #### Optional C:New Question Space
      
      > 如果继续打开问题,最重要的 unresolved question 是:________________________。
      
      #### Optional D:Limitation
      
      > 当前研究不能回答它,是因为:________________________。
      
      #### Optional E:Future Study
      
      > 若要回答它,需要:________________________。
      
      原则:
      
      > **不需要为了填满模板而写 Optional C–E。**
      
      ---
      
      ## 7. Claim–Evidence Mapping:投稿前检查“每个结论凭什么成立”
      
      在进入 reviewer stress test 之前,先把论文中的 major claims 全部抽取出来,并为每个 claim 建立 evidence map:
      
      > **Claim → Figure / Table → Evidence → Strength → Remaining Uncertainty**
      
      例如:
      
      | Claim | Supporting evidence | 仍存在的问题 |
      |---|---|---|
      | MSFold improves alternative-conformation recovery | Fig. 2 benchmark | 是否受 sampling budget 影响? |
      | Improvement is not due to memorization | unseen-protein test | 是否对所有 protein classes 都成立? |
      | Parallel tempering enables broader exploration | sampling analysis / ablation | 是否真正对应物理 energy barrier? |
      
      重点检查:
      
      - 有没有 **claim 没有直接 evidence**?
      - 有没有一个 figure 被用来支持过多不同结论?
      - Discussion 中有没有出现 Results 从未支持过的 claim?
      - claim 的强度是否超过 evidence 的强度?
      - 是否应该把 “demonstrates” 收缩成 “suggests”?
      
      这一步本质上是在建立:
      
      > **Claim–Evidence alignment**
      
      如果一个 claim 找不到明确 supporting evidence,应当:
      
      > **补证据、收缩 claim,或删除 claim。**
      
      ---
      
      ## 8. Reviewer Stress Test
      
      在准备投稿前,主动切换到 reviewer perspective。
      
      问:
      
      > **如果我是最挑剔、最专业的 reviewer,这篇文章最可能受到哪三个挑战?**
      
      生成:
      
      > **Top 3 Reviewer Challenges**
      
      然后逐一分类:
      
      ### 8.1 当前可以补实验解决
      
      → 回到 Results。
      
      ### 8.2 可以通过已有数据分析或解释解决
      
      → 补充 Results / Discussion。
      
      ### 8.3 当前研究确实无法解决
      
      → 写入 Limitation,并设计对应 Future Study。
      
      ### 8.4 Fatal Flaw
      
      例如 central comparison 不公平、claim 与实验设计不匹配、关键变量严重混杂、核心结论无法由现有证据支持。
      
      此时不能简单放进 limitations。
      
      应当:
      
      > **重新设计研究或收缩 central claim。**
      
      因此:
      
      > **Reviewer challenge ≠ limitation**
      
      Reviewer perspective 是一个 **stress test**,而 limitation 只是其可能输出之一。
      
      ---
      
      ## 9. Deadline Mode:有限时间下如何收束研究
      
      WIT / REWRITE 容易产生一个问题:每个 finding 都能继续生成问题,因此研究可以无限延伸。
      
      真实科研存在投稿 deadline、学生毕业、计算资源限制、湿实验成本和项目周期限制。
      
      因此必须加入:
      
      > **Stop Rule / Minimum Sufficient Story**
      
      ### 9.1 Deadline 临近时,优先回答三类问题
      
      #### (1)会改变 central claim 的问题
      
      如果这个问题答案不同,会导致论文主结论不成立或需要明显收缩,必须优先处理。
      
      #### (2)Reviewer 极可能提出的致命问题
      
      例如 data leakage、memorization、unfair baseline、缺少关键 control、alternative explanation。
      
      优先处理。
      
      #### (3)低成本但能显著提高解释力的问题
      
      例如简单 ablation、subgroup analysis、error analysis、关键 negative control。
      
      如果成本低、收益高,应优先补。
      
      ### 9.2 可以暂时停止的问题
      
      如果一个问题:
      
      - 不改变 central claim;
      - 不影响主要证据链;
      - 需要大量新实验;
      - 更适合作为独立研究;
      
      则进入:
      
      > Discussion → Limitation → Future Study
      
      
      ### 9.3 Research Storyline Freeze:防止 project 越做越散
      
      研究进行到一定阶段后,先暂时“冻结”一次主线,写下:
      
      > **Central Question**  
      > **Central Claim**  
      > **3–5 Key Findings**  
      > **General Principle**
      
      然后对每一个新实验问:
      
      > **它会改变或显著加强这条主线吗?**
      
      如果答案是:
      
      - **会改变 central claim** → 值得优先做;
      - **能排除重要 competing hypothesis** → 值得做;
      - **能明确重要 boundary condition** → 可能值得做;
      - **只是再增加一个相似结果** → 谨慎继续。
      
      Storyline Freeze 不是不允许研究变化。  
      如果出现强的 unexpected finding,当然可以重新打开并重写 storyline。
      
      它的作用是防止:
      
      > **项目因为不断追加“也可以做”的实验而失去中心。**
      
      ---
      
      ### 9.4 Stop Rule
      
      
      当满足以下条件时,可以停止继续扩展 Results:
      
      (1)Central scientific question 已被可信地回答;
      (2)关键 competing explanations 已排除到合理程度;
      (3)主要 reviewer challenge 已处理;
      (4)最重要的 boundary conditions 已有基本证据;
      (5)剩余问题需要明显超出当前研究范围的新实验。
      
      原则:
      
      > **目标不是回答所有问题,而是形成一个最小但完整、可信、可辩护的 scientific story。**
      
      ---
      
      ## 10. 句子应该放在哪里:决策图
      
      ```text
      这句话在说什么?
      │
      ├─ 直接报告数据、比较或观察?
      │      └─ Results
      │
      ├─ 解释某一个具体结果的一步意义?
      │      └─ Results subsection ending:1-hop opinion
      │
      ├─ 综合多个 findings,解释本研究整体说明什么?
      │      └─ Discussion opening:2-hop interpretation
      │
      ├─ 提出超越本研究的一般机制或原则?
      │      └─ Discussion middle:general principle
      │
      ├─ 由 findings / principle 引出新的未解问题,而且对界定当前工作或开启下一阶段很重要?
      │      └─ 可放入 Discussion;也可以只作为后续 research question
      │
      ├─ 说明一个重要 unresolved question 为什么当前无法回答?
      │      └─ 如有必要,可写成 Limitation
      │
      └─ 说明下一步用什么实验或分析回答该 unresolved question?
             └─ 如有必要,可写成 Future Study
      ```
      
      ---
      
      ## 11. REWRITE loop 的两个循环
      
      ### 11.1 Inner Loop:研究推进循环
      
      > **Finding → Question → Answerability → Experiment → Finding**
      
      作用:
      
      > 推动当前 project 继续生长。
      
      ### 11.2 Outer Loop:理解与抽象循环
      
      > **Finding → Literature → Interpretation → Principle → New Question**
      
      作用:
      
      > 把一个具体结果提升为更深的 scientific understanding。
      
      两个循环共同决定:
      
      > 做哪些新实验,以及 Results / Discussion 应该如何组�
  • tests
    • Assess-WIT-using-AlphaDev-cn.md 8.3 KB
      # 使用 AlphaDev 检验 WIT
      
      ## 检验范围
      
      本文使用:
      
      - Mankowitz, D. J. et al. *Faster sorting algorithms discovered using deep reinforcement learning*. Nature 618, 257–263 (2023). DOI: 10.1038/s41586-023-06004-9。
      - 当前 WIT Agent Skill:[SKILL.md](../SKILL.md)
      - 当前完整 WIT workflow:[WIT-Scientific-thinking-and-writing-skill.md](../references/WIT-Scientific-thinking-and-writing-skill.md)
      
      AlphaDev 发表早于 WIT,因此对于 WIT 的大部分原则,它可以作为 external stress test。但是,WIT 最近新增的“先用自然语言解释 basic idea,再用最小 concrete example 让 reader mentally execute the method”这一条,本身就是受 AlphaDev 启发后加入的。因此 AlphaDev 可以**说明**这条规则,但不能作为这条规则的独立验证。
      
      ## 总体结论
      
      AlphaDev 对 WIT 的主体架构提供了很强的支持:Results 可以灵活组织、Fact 与 interpretation 应控制推理距离、方法论文需要 mechanism-revealing case、需要比较 competing alternatives、需要分析 boundary,而且 Discussion 的抽象应与 evidence 相称。
      
      但 AlphaDev 也暴露了当前 WIT 中一条过强的 surface rule:
      
      > 不应要求 Figure 1 本身一定承担全文 overview。
      
      AlphaDev 的 Figure 1 主要解释 C++ 与 assembly 的关系;真正把 AssemblyGame 的核心思想讲清楚的是 Figure 2。因此,更好的 WIT 规则应是:
      
      > **方法论文应尽早提供一个 overview figure——通常可以是 Figure 1,但不必强制是 Figure 1——使 reader 能掌握 central idea 和 high-level operation。**
      
      这正好符合 WIT 自己的 meta-principle:
      
      > reasoning function 应被要求,但 surface form 不应被强制。
      
      ## 1. Introduction / missing component
      
      AlphaDev 的 Introduction 清楚地指出一个 missing capability:已有 program synthesis 可以生成或优化程序,但要在巨大的 program space 中高效找到同时 correct 且 fast 的程序仍然困难,尤其是直接优化 CPU-level measured latency。AlphaDev 把这一问题改写成 AssemblyGame,并用 RL + search 来解决。
      
      这与 WIT 的 Introduction logic 相容:
      
      `Final Goal → Existing Components → Missing Capability → Why It Matters → This Study`
      
      论文并没有机械地按这个模板写,反而进一步说明 WIT 的原则是对的:检查的是逻辑功能,而不是固定 prose form。
      
      ## 2. Results 的组织方式
      
      AlphaDev 的 Results 并不是 question-driven titles,而是沿着科学故事推进:
      
      - fixed sorting algorithms;
      - variable sorting algorithms;
      - new algorithm discoveries;
      - swap / copy moves;
      - variable-sort mechanism;
      - 与 stochastic search 对比;
      - additional domains;
      - libc++ patch。
      
      这说明一个优秀的方法论文完全可以采用 component / finding-driven 的表面结构,只要背后的 scientific progression 清楚。
      
      因此 AlphaDev 支持 WIT 当前的修正:
      
      > **Results should expose the logical progression of the scientific story, not obey a single title format.**
      
      ## 3. 自然语言解释 + 最小 concrete walkthrough
      
      AlphaDev 先用自然语言解释 AssemblyGame 的 basic idea:agent 逐步选择 low-level CPU instructions 来构造程序,reward 同时考虑 correctness 与 latency。随后再用很小的 sorting example 把状态、action 和 correctness 具体化。
      
      论文还使用三个数的 sorting network 来解释 AlphaDev swap move。reader 能直接看到传统步骤中为什么存在冗余,以及 AlphaDev 为什么可以少掉一条 instruction。
      
      这非常好地说明了 WIT 新增原则:
      
      > **先用自然语言讲清 idea,再用最小例子让方法“跑起来”。**
      
      但这里必须强调:这是 **positive example,而不是 independent validation**,因为这条 WIT 规则就是受 AlphaDev 启发后加入的。
      
      ## 4. Worked comparative case:why ours succeeds / why existing methods fail
      
      AlphaDev 的 comparative case 做得非常好。
      
      对于 swap / copy move,论文把 classic sorting-network logic 与 AlphaDev 改进后的 logic 并排展示,因此读者不仅知道“少了一条 instruction”,还可以理解少掉的原因。
      
      对于 VarSort4,human benchmark 根据 sequence length 分别调用相应 sorting network;AlphaDev 则复用已经排序的前缀,再调用一个 simplified routine。这样就直接解释了 latency gain 从哪里来。
      
      这支持 WIT 的区分:
      
      > **Benchmark evidence 说明方法有效;worked comparative case 解释方法 how and why it works。**
      
      ## 5. Difference-focused benchmark analysis
      
      AlphaDev 不只报告总体 performance,而是进一步区分不同 regime:
      
      - fixed vs variable sorts;
      - branchless vs branching;
      - algorithm length 能否作为 latency proxy;
      - cold start vs warm start stochastic search。
      
      这正是 WIT 所强调的:不能只问“平均提高多少”,还要问:
      
      > **performance gap 出现在哪里?为什么在那里出现?**
      
      AlphaDev 同时提示 WIT 的措辞可以更宽一点。除了 cases / subsets,还应包括:
      
      > **cases, subsets, conditions, or regimes**
      
      因为方法差异有时来自条件或 regime,而不是某几个具体样本。
      
      ## 6. Competing alternatives、falsification 与 boundary
      
      AlphaDev 认真比较了 stochastic superoptimization baseline,并设置 cold-start 与 warm-start variants,也控制 computational resources / wall-clock。论文并没有宣称 AlphaDev 在任何情况下都优越;例如在某些 branchless、warm-start 场景下 stochastic search 更具 computational efficiency。
      
      这符合 WIT 的核心思想:
      
      > 要设计 discriminating comparison,而不是只找 supporting evidence;claim 要有 boundary。
      
      Nature 文章页面的 post-publication discussion 还提供了一个很好的 boundary lesson。论文写到 brute force 可证明 sort 3 不存在短于 17 instructions 的程序;作者后来澄清,这个 lower bound 是针对本文限定的 branchless instruction set。这个例子非常直接地说明:
      
      > **Claim strength and scope must match the tested evidence boundary.**
      
      ## 7. Discussion
      
      AlphaDev 的 Discussion 很短:总结主要 achievement,指出 RL 与 stochastic search 的互补性,并讨论可能的 generalization。它没有机械出现一套“Limitations → Future Work → Conclusion”的固定结构。
      
      这支持 WIT 当前的非模板化原则:
      
      > **Discussion core = Integrated Interpretation → Broader Meaning / justified abstraction.**
      
      New Questions、Boundary、Limitations、Future Studies 有科学需要时才展开,不应成为 ritualized sections。
      
      ## 8. 六维 question space
      
      AlphaDev 很自然地可以用 WIT 的六维 question space 展开:
      
      - **Existence / Whether:** RL 能否发现优于高度优化 human baseline 的 sorting routines?
      - **Determinants / What:** representation、reward、search strategy、branching structure、initialization regime。
      - **Cause / Why:** learned search 为什么能优于某些 stochastic search。
      - **Mechanism / How:** AssemblyGame + neural representation + MCTS-guided RL + correctness/latency reward。
      - **Boundary / When:** branchless / branching、cold / warm start、instruction-set restriction、hardware regime。
      - **Magnitude / To what extent:** instruction savings、latency gains、libc++ downstream impact。
      
      因此六维框架作为 attention map 是合理的;它不需要变成论文的六个 subsection。
      
      ## 最终判断
      
      AlphaDev 没有推翻 WIT 的核心 reasoning architecture。相反,它:
      
      1. 支持 WIT 对 Results / Discussion 非模板化的理解;
      2. 强烈说明 mechanism-revealing worked example 的价值;
      3. 支持对 performance gap 来源进行系统分析;
      4. 支持 competing-method comparison 与 boundary-aware claim;
      5. **反驳了“Figure 1 本身通常必须压缩全文 idea”这一过强 surface rule。**
      
      因此最准确的结论不是:
      
      > “AlphaDev 验证了 WIT。”
      
      而是:
      
      > **AlphaDev 对 WIT 的大部分原则构成了一个很强的 external stress test,并帮助 WIT 把一条过强规则修得更准确。**
      
      最值得立即修改的一条是:
      
      > **方法论文应尽早提供一个 overview figure——often, but not necessarily, Figure 1——使 reader 能够掌握 central idea 和 high-level operation。**
      
      这本身就是 WIT 的运行方式:
      
      `Rule → Counterexample → Boundary → Better Rule`
      
    • Assess-WIT-using-AlphaDev.md 8.3 KB
      # Assessing WIT using AlphaDev
      
      ## Scope
      
      This assessment uses:
      
      - Mankowitz, D. J. et al. *Faster sorting algorithms discovered using deep reinforcement learning*. Nature 618, 257–263 (2023). DOI: 10.1038/s41586-023-06004-9.
      - Current WIT Agent Skill: [SKILL.md](../SKILL.md)
      - Current full WIT workflow: [WIT-Scientific-thinking-and-writing-skill.md](../references/WIT-Scientific-thinking-and-writing-skill.md)
      
      AlphaDev predates WIT and is therefore useful as an external stress test for most of WIT. However, one recently added WIT rule—natural-language explanation followed by a minimal concrete walkthrough—was explicitly motivated by AlphaDev. AlphaDev therefore illustrates that rule but cannot independently validate it.
      
      ## Overall conclusion
      
      AlphaDev strongly supports the main architecture of WIT: flexible Results organization, evidence-to-interpretation distance, mechanism-revealing cases, competing alternatives, boundary analysis, and evidence-proportional Discussion.
      
      At the same time, AlphaDev exposes one important over-specification in the current WIT skill:
      
      > WIT should not require Figure 1 itself to summarize the entire paper.
      
      In AlphaDev, Figure 1 explains the relationship between C++ and assembly, whereas Figure 2 provides the central conceptual view of AssemblyGame. A better WIT rule is:
      
      > **An early overview figure—often, but not necessarily, Figure 1—should allow the reader to grasp the central idea and high-level operation of a method paper.**
      
      This correction is consistent with WIT's own meta-principle: reasoning function should be required, not a fixed surface form.
      
      ## 1. Introduction / missing component
      
      AlphaDev clearly motivates a missing capability. Earlier program-synthesis approaches can generate or optimize programs, but the paper emphasizes the difficulty of efficiently searching for programs that are both correct and fast, especially when optimizing actual CPU-level latency. AlphaDev formulates this search as AssemblyGame and uses reinforcement learning plus search to address it.
      
      This is compatible with WIT's Introduction logic:
      
      `Final Goal → Existing Components → Missing Capability → Why It Matters → This Study`
      
      The paper does not literally follow that template, which supports WIT's principle that the logical function matters more than prose form.
      
      ## 2. Results organization
      
      AlphaDev's Results are not organized as explicit questions. They progress through:
      
      - fixed sorting algorithms;
      - variable sorting algorithms;
      - new algorithmic discoveries;
      - swap and copy moves;
      - variable-sort mechanisms;
      - comparison with stochastic search;
      - additional domains;
      - the libc++ patch.
      
      This is a strong external example of a component/finding-driven Results structure whose underlying scientific progression remains recoverable.
      
      Therefore AlphaDev supports the WIT revision:
      
      > **Results should expose the logical progression of the scientific story, not obey a single title format.**
      
      ## 3. Natural-language explanation and minimal walkthrough
      
      AlphaDev first explains the basic idea of AssemblyGame in ordinary language: the agent constructs a program by selecting low-level instructions, and it is rewarded for correctness and latency. It then makes the process concrete with a small sorting example.
      
      The paper also uses a three-element sorting network to explain the AlphaDev swap move. The reader can see why a previous comparison becomes redundant and how one instruction can be removed.
      
      This is an excellent illustration of the WIT rule:
      
      > Explain the idea in words, then make it executable with a minimal example.
      
      But this is a **positive example, not independent validation**, because AlphaDev explicitly motivated the addition of this rule to WIT.
      
      ## 4. Worked comparative cases: why ours succeeds and existing methods fail
      
      AlphaDev contains unusually strong comparative cases.
      
      For the swap/copy moves, the paper shows the conventional sorting-network logic and the shortened logic discovered by AlphaDev, making the source of the improvement understandable.
      
      For VarSort4, the human benchmark dispatches to a separate sorting network according to sequence length, whereas AlphaDev reuses a partially sorted prefix and then applies a simplified routine. This explains not merely that AlphaDev is faster, but how the algorithmic structure produces the latency gain.
      
      This strongly supports WIT's distinction:
      
      > **Benchmark evidence establishes that a method works; a worked comparative case explains how and why it works.**
      
      ## 5. Difference-focused benchmark analysis
      
      AlphaDev does more than report aggregate superiority. It separates regimes in which the methods behave differently:
      
      - fixed versus variable sorts;
      - branchless versus branching programs;
      - algorithm length as a useful latency proxy versus regimes in which length and latency decouple;
      - cold-start versus warm-start stochastic search.
      
      This is exactly the type of analysis WIT is trying to encourage. It shows that "where the gain comes from" may be expressed not only through individual cases or subsets, but also through **conditions or regimes**.
      
      Suggested refinement to WIT:
      
      > replace "cases/subsets" with **"cases, subsets, conditions, or regimes"**.
      
      ## 6. Competing alternatives, falsification, and boundaries
      
      The paper directly compares AlphaDev with a stochastic superoptimization baseline under matched or favorable conditions, including cold-start and warm-start variants. It also notes cases in which warm-start stochastic search is computationally more efficient, rather than claiming universal superiority.
      
      This is consistent with WIT's preference for discriminating comparisons and boundary-aware claims.
      
      A post-publication discussion on the Nature article page also illustrates why claim boundaries matter. The paper states that brute force established a 17-instruction lower bound for sort 3; an author later clarified that this lower bound is within the restricted branchless instruction set considered in the work. This is a concrete example of the WIT principle:
      
      > **Claim strength and scope must match the evidence and tested boundary.**
      
      ## 7. Discussion
      
      AlphaDev's Discussion is short. It summarizes the main achievement, notes complementary strengths of RL and stochastic search, and discusses possible generalization. It does not contain a ritualized sequence of "limitations → future work → conclusion."
      
      This supports WIT's current formulation:
      
      > **Discussion core = Integrated Interpretation → Broader Meaning / justified abstraction.**
      
      New questions, boundaries, limitations, and future studies are useful when scientifically needed, not mandatory surface sections.
      
      ## 8. Six-dimensional question space
      
      AlphaDev can be naturally interrogated through the WIT question space:
      
      - **Existence / Whether:** Can an RL-based system discover sorting routines better than highly optimized human baselines?
      - **Determinants / What:** Representation, reward definition, search method, branching structure, and initialization regime affect performance.
      - **Cause / Why:** Learned search can explore program space differently from stochastic optimization.
      - **Mechanism / How:** AssemblyGame + neural representation + MCTS-guided RL + correctness/latency reward.
      - **Boundary / When:** Branchless versus branching routines, cold versus warm start, instruction-set restrictions, and hardware regimes.
      - **Magnitude / To what extent:** Instruction savings, latency improvements, and downstream libc++ impact.
      
      The dimensions are useful as an attention map without needing to appear as six explicit paper sections.
      
      ## Verdict
      
      AlphaDev does not falsify WIT's core reasoning architecture. Instead, it:
      
      1. supports WIT's flexible Results and Discussion principles;
      2. strongly illustrates the value of mechanism-revealing worked examples;
      3. supports systematic analysis of where performance differences arise;
      4. supports boundary-aware and competing-method comparisons;
      5. **falsifies an overly rigid surface rule about Figure 1.**
      
      Therefore the appropriate conclusion is not "AlphaDev validates WIT." It is:
      
      > **AlphaDev provides a strong external stress test for most of WIT and sharpens one of its rules.**
      
      The most important revision is:
      
      > **An early overview figure—often, but not necessarily, Figure 1—should convey the central idea and high-level operation of a method paper.**
      
      This outcome is itself consistent with WIT:
      
      `Rule → Counterexample → Boundary → Better Rule`
      
    • Assess-WIT-using-AlphaGo-cn.md 15.4 KB
      # WIT 案例:用AlphaGo论文检验WIT
      
      **论文:** Silver, D. et al. *Mastering the game of Go with deep neural networks and tree search*. Nature 529, 484–489 (2016).  
      **DOI:** 10.1038/nature16961  
      **WIT 角色:** 独立外部检验案例。AlphaGo 论文发表于 WIT 形成之前,因此它特别适合用于 falsify / refine WIT,而不是“证明 WIT”。
      
      ---
      
      ## 1. 为什么选择 AlphaGo?
      
      AlphaGo 是一个很好的 WIT stress test,原因有三:
      
      (1)文章质量高,scientific story 极强;  
      (2)它属于 AI method / system paper,而不是生物学 discovery paper;  
      (3)它的 Results 明显采用 pipeline/component-driven 组织,因此可以检验 WIT 是否把“question-driven Results”规定得过死。
      
      因此,这个案例的目标不是证明:
      
      > **AlphaGo 符合 WIT。**
      
      而是问:
      
      > **WIT 的哪些 reasoning principles 经得住 AlphaGo 检验?哪些必须修正?**
      
      这体现 WIT 自己的原则:
      
      > **A successful falsification sharpens the claim.**
      
      ---
      
      ## 2. Introduction:WIT 的 missing-component 逻辑成立吗?
      
      ### 2.1 Final Goal
      
      AlphaGo 的 final goal 很清楚:
      
      > **在 full-sized Go 上达到甚至超过人类专业棋手水平。**
      
      困难来自围棋极大的搜索空间。论文指出,Go 的典型 branching factor 约为 250、game depth 约为 150,exhaustive search 不可行。
      
      ### 2.2 Necessary Components
      
      论文把大规模搜索问题拆成两个核心需求:
      
      - **Move selection / policy**:减少 search breadth;
      - **Position evaluation / value**:减少 search depth。
      
      换句话说,要解决 Go,不只是“搜索更多”,而是要:
      
      > **更聪明地决定搜哪里,以及更准确地判断一个局面好不好。**
      
      ### 2.3 What Had Been Established?
      
      已有工作已经提供了:
      
      - Monte Carlo tree search;
      - rollout;
      - shallow policy;
      - handcrafted / relatively simple value functions。
      
      这些组件已经把 computer Go 推到 strong amateur level,但还未达到 professional level。
      
      ### 2.4 Missing Component
      
      AlphaGo 真正补上的不是单一一个算法,而是一组此前缺失、且可以被学习出来的高质量组件:
      
      > **deep policy network + deep value network + learning from expert games and self-play + integration with MCTS**
      
      因此,如果用 WIT 表达:
      
      > **Final Goal → Established Components → Missing Capability → This Study**
      
      逻辑非常清楚。
      
      ### 2.5 WIT verdict
      
      **强支持。**
      
      但这里也提示 WIT:
      
      > “Missing component” 不一定是一个单独模块,也可能是一组必须协同工作的 missing capabilities。
      
      ---
      
      ## 3. 用六维 scientific question space 打开 AlphaGo
      
      WIT 将“开题”从疑问词升级为六类 scientific dimensions:
      
      > **Existence → Determinants → Cause → Mechanism → Boundary Conditions → Magnitude**
      
      ### 3.1 Whether → Existence
      
      问题:
      
      > **深度网络 + tree search 是否真的能够达到专业围棋水平?**
      
      论文给出直接 evidence:
      
      - 对其他 Go programs 的胜率达到 99.8%;
      - 对欧洲冠军 Fan Hui 的正式比赛为 5–0。
      
      因此,Existence 得到强回答。
      
      ---
      
      ### 3.2 What → Determinants
      
      问题:
      
      > **哪些因素决定 AlphaGo 的棋力?**
      
      论文逐步分析了:
      
      - supervised policy quality;
      - reinforcement learning;
      - value network;
      - rollout;
      - MCTS;
      - search computation / scaling。
      
      例如,SL policy 的 expert-move prediction accuracy 达到 57.0%;RL policy 对 SL policy 的胜率超过 80%;即使不使用 search,RL policy 对 Pachi 的胜率达到 85%。
      
      因此,论文不只回答“AlphaGo 强”,还分析:
      
      > **What determines how strong it becomes?**
      
      ---
      
      ### 3.3 Why → Cause
      
      问题:
      
      > **为什么传统方法难以解决 Go,而 AlphaGo 能突破?**
      
      核心 causal explanation 是:
      
      > **Go 的原始搜索空间不可处理;AlphaGo 利用 learned policy 和 value function,大幅降低 effective search breadth 和 depth。**
      
      因此,Why 不是简单说:
      
      > “因为用了 deep learning。”
      
      而是:
      
      > **learning makes an otherwise intractable search problem tractable enough to search effectively.**
      
      ---
      
      ### 3.4 How → Mechanism
      
      问题:
      
      > **这种 search-space reduction 具体是怎么发生的?**
      
      mechanism 是:
      
      - policy network 为 tree search 提供 move prior,优先扩展更有希望的 moves;
      - value network 在 leaf position 直接估计 winning probability;
      - rollout 提供另一种快速评估;
      - MCTS 将这些信息 backup 到搜索树中,更新 action values 与 visit counts。
      
      因此:
      
      > **Cause:learned policy/value 缩小有效搜索空间。**  
      > **Mechanism:policy-guided selection + value evaluation + rollout + MCTS backup。**
      
      这很好地说明:
      
      > **Why asks what causes it; How asks through what mechanism the cause produces the effect.**
      
      ---
      
      ### 3.5 When → Boundary Conditions
      
      AlphaGo 原论文在这一维度上只给出**部分 evidence**:
      
      - full-sized Go;
      - 多个 computer Go programs;
      - 一位职业人类棋手;
      - 不同 search resources / system scales。
      
      但它没有系统画出完整 boundary map。
      
      沿 WIT 可以继续追问:
      
      - search budget 降到什么程度时优势消失?
      - policy network 足够弱时,MCTS 能否补偿?
      - value-network systematic bias 在什么条件下导致 search failure?
      - 对不同人类棋风和不同水平是否同样成立?
      - learning + search 是否能迁移到 Go 之外?
      
      这些是 **new questions**,不能伪装成原论文已经回答的结论。
      
      ---
      
      ### 3.6 To what extent → Magnitude
      
      论文提供了多层 magnitude:
      
      - 99.8% 对其他 Go programs;
      - 5–0 Fan Hui;
      - RL policy 对 SL policy >80%;
      - 无 search 的 RL policy 对 Pachi 85%;
      - value-network evaluation 接近强 rollout 的 accuracy,但单次计算量约低 15,000 倍。
      
      因此:
      
      > **Whether:有没有效果?**  
      > **Magnitude:效果到底有多大?**
      
      区分清楚。
      
      ---
      
      ## 4. Results:AlphaGo 对 WIT 的关键 falsification
      
      ### 4.1 Results subsection titles
      
      AlphaGo 的主要研究段落依次为:
      
      1. **Supervised learning of policy networks**
      2. **Reinforcement learning of policy networks**
      3. **Reinforcement learning of value networks**
      4. **Searching with policy and value networks**
      5. **Evaluating the playing strength of AlphaGo**
      
      这显然不是 finding-driven titles,而是:
      
      > **Component / Pipeline-driven**
      
      如果 WIT 规定:
      
      > “Results 必须按 scientific question / finding,而不能按 technique 组织。”
      
      那么 AlphaGo 就是一个明确 counterexample。
      
      ### 4.2 但是 small-essay test 成立
      
      虽然标题是 pipeline-driven,但连起来非常清楚:
      
      > **learn expert policy → improve policy by self-play → learn value → integrate policy/value with search → evaluate final system**
      
      所以真正的 invariant 不是标题形式,而是:
      
      > **Results should expose the logical progression of the scientific story.**
      
      因此 WIT 必须区分:
      
      - **Question / Finding-driven Results**
      - **Component / Pipeline-driven Results**
      
      二者都可以优秀。
      
      关键问题是:
      
      > **subsections 是 coherent progression,还是简单的 technique list?**
      
      AlphaGo 显然属于前者。
      
      ---
      
      ## 5. Results subsection:Fact → restrained 1-hop Opinion 成立吗?
      
      这一条经 AlphaGo 检验后反而更强。
      
      ### 5.1 例 1:RL policy
      
      Fact:
      
      > RL policy head-to-head 对 SL policy 的胜率超过 80%;无 search 时对 Pachi 胜率达到 85%。
      
      局部意义:
      
      > 优化最终 winning objective 能进一步提高 policy 的实际 playing strength。
      
      这是很近的 inference,没有跳得太远。
      
      ### 5.2 例 2:value network
      
      Fact:
      
      > 单次 value-network evaluation 接近使用强 RL policy 的 Monte Carlo rollouts 的 accuracy,但所需计算量约少 15,000 倍。
      
      局部意义:
      
      > learned value network 可以成为高效 position evaluator。
      
      仍然是局部能力解释。
      
      ### 5.3 例 3:mixed position evaluation
      
      Fact:
      
      > value-only 和 rollout-only 都能工作,但二者混合效果最好,对其他 variants 的胜率 ≥95%。
      
      作者随后给出局部解释:
      
      > 两种 position-evaluation mechanisms 是 complementary。
      
      这几乎就是 WIT 的标准例子:
      
      > **Fact → 1-hop Opinion**
      
      ### 5.4 WIT verdict
      
      **强支持。**
      
      但要加一句:
      
      > **Reasoning structure ≠ Surface prose structure.**
      
      AlphaGo 并没有机械地写:
      
      > “We next asked whether ...”
      
      例如作者直接说 “The second stage of the training pipeline ...”。
      
      所以 WIT 应要求:
      
      > **Motivation / Question → Test → Fact → 1-hop → Next Step 的 reasoning 可恢复**
      
      而不是要求 prose 必须模板化。
      
      ---
      
      ## 6. Discussion:AlphaGo 是否支持 WIT?
      
      AlphaGo 的 Discussion 很短,但非常适合检验 WIT。
      
      ### 6.1 Core Function 1:Integrated Interpretation
      
      第一部分不是重复 57%、85%、99.8%、5–0,而是把:
      
      - policy network;
      - value network;
      - supervised learning;
      - reinforcement learning;
      - tree search
      
      综合为一个整体 system。
      
      这符合:
      
      > **multiple local findings → integrated interpretation**
      
      或者用 WIT 的语言:
      
      > **multiple 1-hop Opinions → 2-hop Interpretation**
      
      ---
      
      ### 6.2 Core Function 2:Broader Meaning
      
      接下来作者把 AlphaGo 与 Deep Blue 比较。
      
      关键不是“谁赢得更多”,而是指出:
      
      > AlphaGo 搜索的 positions 少得多,却通过 policy network 更聪明地选择 positions,并通过 value network 更准确地评价它们。
      
      这已经从:
      
      > “各组件效果如何”
      
      上升到:
      
      > **为什么这种 computational strategy 有效。**
      
      可以压缩为:
      
      > **learn where to search and how to evaluate.**
      
      这是明显的 deeper interpretation。
      
      ---
      
      ### 6.3 Beyond This Study
      
      Discussion 最后把 Go 看作更一般的困难 AI decision/search problem,并指出 MCTS 过去已经扩散到 planning、scheduling、constraint satisfaction 等领域;AlphaGo 的 learning + search 组合可能对其他 seemingly intractable AI domains 有启发。
      
      这符合:
      
      > **this study → broader meaning / abstraction**
      
      ---
      
      ### 6.4 重要 counterexample:没有标准 Limitations / Future Studies
      
      AlphaGo Discussion 没有显式:
      
      > **New Questions → Limitations → Future Studies**
      
      但它仍然是一篇非常完整、有力的 Discussion。
      
      因此,它 falsify 了一个过强版本的 WIT:
      
      > “好 Discussion 必须具有固定六段式结构。”
      
      更合理的 WIT 是:
      
      > **Core Discussion:Integrated Interpretation → Broader Meaning**
      
      必要时再扩展:
      
      > **Optional:New Questions → Boundary / Limitations → Future Studies**
      
      ### 6.5 WIT verdict
      
      **强支持 core functions;反对 rigid surface template。**
      
      ---
      
      ## 7. Competing Hypotheses:AlphaGo 做得如何?
      
      从 WIT 角度,AlphaGo 实际上进行了不少 discriminating tests。
      
      ### 7.1 Hypothesis:只是 supervised imitation 就够了?
      
      Test:
      
      > RL policy vs SL policy。
      
      Result:
      
      > RL policy >80% 胜率。
      
      说明单纯 imitation 不是全部,direct optimization for winning 能继续提高能力。
      
      ### 7.2 Hypothesis:只要 policy,不需要 search?
      
      Test:
      
      > policy-only system 与加入 search 的 variants 比较。
      
      Result:
      
      > search 显著提升最终 system strength。
      
      ### 7.3 Hypothesis:value network 与 rollout 是替代关系?
      
      Test:
      
      > value-only、rollout-only、mixed evaluation。
      
      Result:
      
      > mixed best。
      
      因此结论从:
      
      > “哪一个更好?”
      
      变成:
      
      > **它们是 complementary mechanisms。**
      
      这是一种很漂亮的 competing-hypothesis refinement。
      
      ---
      
      ## 8. Falsification / Counterexample Check
      
      对 AlphaGo 的 central claim,可以构造潜在 falsifiers:
      
      - learned policy/value 并不能显著提高 playing strength;
      - search scaling 不能提高强度;
      - mixed evaluation 不优于单一 mechanism;
      - 对专业人类棋手失败;
      - system 只是在特定 computer opponents 上有效。
      
      论文中的实验没有发现这些核心 falsifiers。
      
      因此:
      
      > **central claim gains strong support, but is not universally proven.**
      
      仍然没有被系统检验的部分包括:
      
      - 不同类型 professional opponents;
      - 极低 compute regime;
      - Go 之外任务的 transferability。
      
      这些构成 boundary,而不是论文“错误”。
      
      ---
      
      ## 9. Claim–Evidence Mapping
      
      | Major claim | Main evidence | WIT judgment |
      |---|---|---|
      | Deep policy networks can predict strong moves | SL policy 57.0% expert-move accuracy;playing-strength analysis | 直接支持 |
      | RL improves actual playing strength | RL policy >80% vs SL;85% vs Pachi without search | 强支持 |
      | Value network can efficiently evaluate positions | accuracy vs rollouts;约 15,000× lower computation | 强支持 |
      | Policy/value + MCTS yields much stronger Go | component variants + tournament | 强支持 |
      | Mixed value + rollout evaluation is complementary | mixed variant ≥95% vs other variants | 很好的 Fact → 1-hop |
      | AlphaGo reaches professional level | 5–0 vs Fan Hui | 直接支持 |
      | The principle may matter beyond Go | conceptual analogy to other search/decision domains | 合理 broader implication,但不是实验验证 |
      
      最后一条特别重要:
      
      > **Broader implication ≠ experimentally demonstrated generalization.**
      
      AlphaGo 的措辞总体比较克制。
      
      ---
      
      ## 10. Reviewer Stress Test
      
      如果用 WIT 站在 reviewer 角度,最强的问题可能包括:
      
      (1)**Professional-level claim 的人类样本是否过少?**  
      只正式对一位 professional player 进行了五局比赛。
      
      (2)**强度提升中 compute scaling 与 algorithmic improvement 各占多少?**  
      论文有 single-machine / distributed 和 component comparisons,但完整因果分解仍可继续做。
      
      (3)**learning + search 的 general principle 是否真正能迁移到其他 domains?**  
      Discussion 提出希望,而没有把它当成已证明结果,因此不构成过度 claim。
      
      这些问题更多定义了后续 research space,而没有击穿论文 central claim。
      
      ---
      
      ## 11. WIT 对 AlphaGo 的最终评价
      
      ### 被 AlphaGo 强化的 WIT 原则
      
      > **Fact → restrained 1-hop Opinion**
      
      > **Multiple local findings → Integrated Interpretation → Broader Meaning**
      
      > **Finding / Component → Next logical step**
      
      > **Claim → Evidence**
      
      > **Competing explanations should be discriminated where possible.**
      
      ### 被 AlphaGo 修正的 WIT 原则
      
      过强版本:
      
      > **Results 必须 question/finding-driven。**
      
      修正为:
      
      > **Results 必须呈现 coherent logical progression;question-driven 与 pipeline-driven 都可以。**
      
      过强版本:
      
      > **Discussion 必须包含 New Questions、Limitations、Future Studies。**
      
      修正为:
      
      > **Discussion core = Integrated Interpretation → Broader Meaning;其他是 optional extensions。**
      
      ---
      
      ## 12. 这个案例对 WIT 最重要的贡献
      
      AlphaGo 是 WIT 的一个真正 external falsification test。
      
      它说明:
      
      > **WIT 最有价值的不是规定论文长什么样,而是识别论文必须完成哪些 reasoning functions。**
      
      因此可以得到两个 WIT 元原则:
      
      > **WIT specifies reasoning functions, not rigid prose forms.**
      
      以及:
      
      > **WIT is a question generator, not a checklist completer.**
      
      从这个意义上说,AlphaGo 不只是 WIT 的一个“应用示例”,也是帮助 WIT 自身变得更准确的一个 counterexample-driven refinement。
      
      ---
      
      ## Source
      
      Silver, D., Huang, A., Maddison, C. et al. *Mastering the game of Go with deep neural networks and tree search*. **Nature** 529, 484–489 (2016). DOI: 10.1038/nature16961.
      
    • Assess-WIT-using-AlphaGo.md 16.8 KB
      # WIT Case Study: Auditing WIT using the AlphaGo Paper
      
      **Paper:** Silver, D. et al. *Mastering the game of Go with deep neural networks and tree search*. Nature 529, 484–489 (2016).  
      **DOI:** 10.1038/nature16961  
      **Role in WIT:** An independent external validation case. Because the AlphaGo paper was published long before WIT was developed, it is particularly useful for falsifying and refining WIT rather than merely “demonstrating that WIT works.”
      
      ---
      
      ## 1. Why Choose AlphaGo?
      
      AlphaGo is an excellent stress test for WIT for three reasons:
      
      (1) It is a high-quality paper with an exceptionally strong scientific story.  
      (2) It is an AI method/system paper rather than a biological discovery paper.  
      (3) Its Results are clearly organized around components and pipeline stages, making it useful for testing whether WIT overstates the need for question-driven Results.
      
      Therefore, the goal of this case study is not to prove:
      
      > **AlphaGo conforms to WIT.**
      
      Instead, the question is:
      
      > **Which WIT reasoning principles survive the AlphaGo test, and which need to be revised?**
      
      This follows WIT's own principle:
      
      > **A successful falsification sharpens the claim.**
      
      ---
      
      ## 2. Introduction: Does the WIT Missing-Component Logic Hold?
      
      ### 2.1 Final Goal
      
      AlphaGo's final goal is clear:
      
      > **Achieve or exceed professional human performance in full-sized Go.**
      
      The difficulty comes from the enormous search space of Go. The paper notes that Go has an approximate branching factor of 250 and a typical game depth of 150, making exhaustive search infeasible.
      
      ### 2.2 Necessary Components
      
      The paper effectively decomposes the large-scale search problem into two core needs:
      
      - **Move selection / policy:** reduce search breadth;
      - **Position evaluation / value:** reduce search depth.
      
      In other words, solving Go is not simply about “searching more.” It requires:
      
      > **searching more selectively and evaluating positions more accurately.**
      
      ### 2.3 What Had Already Been Established?
      
      Previous work had already provided:
      
      - Monte Carlo tree search;
      - rollout;
      - shallow policy models;
      - handcrafted or relatively simple value functions.
      
      These components had pushed computer Go to strong amateur level, but not to professional level.
      
      ### 2.4 Missing Component
      
      What AlphaGo adds is not one isolated algorithm, but a set of previously missing learned capabilities:
      
      > **deep policy network + deep value network + learning from expert games and self-play + integration with MCTS**
      
      In WIT terms:
      
      > **Final Goal → Established Components → Missing Capability → This Study**
      
      The logic is very clear.
      
      ### 2.5 WIT Verdict
      
      **Strongly supported.**
      
      However, AlphaGo also suggests one refinement:
      
      > A “missing component” does not have to be a single module. It can also be a set of missing capabilities that must work together.
      
      ---
      
      ## 3. Opening AlphaGo Along the Six Scientific Dimensions
      
      WIT reframes problem opening from a list of interrogative words into six scientific dimensions:
      
      > **Existence → Determinants → Cause → Mechanism → Boundary Conditions → Magnitude**
      
      ### 3.1 Whether → Existence
      
      Question:
      
      > **Can deep networks + tree search actually reach professional Go level?**
      
      The paper provides direct evidence:
      
      - a 99.8% win rate against other Go programs;
      - a 5–0 match result against European Go champion Fan Hui.
      
      Thus, Existence is strongly established.
      
      ---
      
      ### 3.2 What → Determinants
      
      Question:
      
      > **What factors determine AlphaGo's playing strength?**
      
      The paper progressively analyzes:
      
      - supervised policy quality;
      - reinforcement learning;
      - value network;
      - rollout;
      - MCTS;
      - search computation / scaling.
      
      For example, the supervised policy network reaches 57.0% expert-move prediction accuracy; the RL policy defeats the SL policy in more than 80% of games; and even without search, the RL policy beats Pachi in 85% of games.
      
      So the paper does not merely ask whether AlphaGo is strong. It also asks:
      
      > **What determines how strong it becomes?**
      
      ---
      
      ### 3.3 Why → Cause
      
      Question:
      
      > **Why do traditional approaches struggle with Go, whereas AlphaGo breaks through?**
      
      The core causal explanation is:
      
      > **The raw Go search space is intractable; AlphaGo uses learned policy and value functions to reduce the effective breadth and depth of search.**
      
      Thus, Why is not simply:
      
      > “Because it uses deep learning.”
      
      A more informative causal statement is:
      
      > **Learning makes an otherwise intractable search problem tractable enough to search effectively.**
      
      ---
      
      ### 3.4 How → Mechanism
      
      Question:
      
      > **How does this reduction of the effective search space actually occur?**
      
      The mechanism is:
      
      - the policy network provides move priors to tree search and prioritizes promising moves;
      - the value network directly estimates winning probability at leaf positions;
      - rollout provides another fast evaluation signal;
      - MCTS backs these signals up through the search tree to update action values and visit counts.
      
      Thus:
      
      > **Cause:** learned policy/value functions reduce the effective search space.  
      > **Mechanism:** policy-guided selection + value evaluation + rollout + MCTS backup.
      
      This cleanly illustrates:
      
      > **Why asks what causes it; How asks through what mechanism the cause produces the effect.**
      
      ---
      
      ### 3.5 When → Boundary Conditions
      
      The original AlphaGo paper provides only **partial evidence** for this dimension:
      
      - full-sized Go;
      - multiple computer Go programs;
      - one professional human player;
      - different search resources / system scales.
      
      But it does not systematically map the complete boundary conditions.
      
      WIT naturally generates further questions:
      
      - At what search budget does the advantage disappear?
      - Can MCTS compensate when the policy network is weak?
      - Under what conditions does systematic bias in the value network cause search failure?
      - Does the advantage hold across human opponents with different styles and skill levels?
      - Can the learning + search principle transfer beyond Go?
      
      These are **new questions** and should not be presented as claims already answered by the original paper.
      
      ---
      
      ### 3.6 To What Extent → Magnitude
      
      The paper provides several quantitative scales:
      
      - 99.8% win rate against other Go programs;
      - 5–0 against Fan Hui;
      - RL policy vs. SL policy: >80%;
      - RL policy without search vs. Pachi: 85%;
      - value-network evaluation achieves accuracy comparable to strong rollouts while requiring roughly 15,000 times less computation per evaluation.
      
      Thus:
      
      > **Whether: Is there an effect?**  
      > **Magnitude: How large is the effect?**
      
      The distinction is clear.
      
      ---
      
      ## 4. Results: AlphaGo Provides a Key Falsification of WIT
      
      ### 4.1 Results Subsection Titles
      
      The main research sections of AlphaGo proceed as follows:
      
      1. **Supervised learning of policy networks**
      2. **Reinforcement learning of policy networks**
      3. **Reinforcement learning of value networks**
      4. **Searching with policy and value networks**
      5. **Evaluating the playing strength of AlphaGo**
      
      These are clearly not finding-driven titles. They are:
      
      > **Component / Pipeline-driven**
      
      If WIT stated:
      
      > “Results must be organized by scientific questions/findings rather than techniques,”
      
      then AlphaGo would be a direct counterexample.
      
      ### 4.2 But the Small-Essay Test Holds
      
      Although the titles are pipeline-driven, together they form a very clear progression:
      
      > **learn expert policy → improve policy by self-play → learn value → integrate policy/value with search → evaluate the final system**
      
      So the true invariant is not the title format. It is:
      
      > **Results should expose the logical progression of the scientific story.**
      
      WIT should therefore distinguish between:
      
      - **Question / Finding-driven Results**
      - **Component / Pipeline-driven Results**
      
      Both can be excellent.
      
      The key question is:
      
      > **Do the subsections form a coherent progression, or are they merely a list of techniques?**
      
      AlphaGo clearly belongs to the former.
      
      ---
      
      ## 5. Results Subsections: Does Fact → Restrained 1-hop Opinion Hold?
      
      This principle becomes even stronger after the AlphaGo test.
      
      ### 5.1 Example 1: RL Policy
      
      Fact:
      
      > The RL policy wins more than 80% of head-to-head games against the SL policy, and without search it beats Pachi in 85% of games.
      
      Local meaning:
      
      > Optimizing the final winning objective further improves actual playing strength beyond supervised imitation.
      
      This inference stays close to the data.
      
      ### 5.2 Example 2: Value Network
      
      Fact:
      
      > A single value-network evaluation achieves accuracy comparable to Monte Carlo rollouts using a strong RL policy, but with roughly 15,000 times less computation.
      
      Local meaning:
      
      > A learned value network can serve as an efficient position evaluator.
      
      Again, this is a local interpretation.
      
      ### 5.3 Example 3: Mixed Position Evaluation
      
      Fact:
      
      > Value-only and rollout-only evaluation both work, but combining them performs best, defeating the other variants in at least 95% of games.
      
      The authors then make a local interpretation:
      
      > The two position-evaluation mechanisms are complementary.
      
      This is almost a textbook WIT example:
      
      > **Fact → 1-hop Opinion**
      
      ### 5.4 WIT Verdict
      
      **Strongly supported.**
      
      But one important caveat should be added:
      
      > **Reasoning structure ≠ Surface prose structure.**
      
      AlphaGo does not mechanically write:
      
      > “We next asked whether ...”
      
      For example, it simply says that “the second stage of the training pipeline” aims to improve the policy network.
      
      Therefore WIT should require:
      
      > **The underlying Motivation / Question → Test → Fact → 1-hop → Next Step reasoning must be recoverable**
      
      rather than requiring formulaic prose.
      
      ---
      
      ## 6. Discussion: Does AlphaGo Support WIT?
      
      AlphaGo's Discussion is short, but it is an excellent test of WIT.
      
      ### 6.1 Core Function 1: Integrated Interpretation
      
      The first part does not repeat 57%, 85%, 99.8%, or 5–0.
      
      Instead, it integrates:
      
      - policy network;
      - value network;
      - supervised learning;
      - reinforcement learning;
      - tree search
      
      into one coherent system-level interpretation.
      
      This fits:
      
      > **multiple local findings → integrated interpretation**
      
      or in WIT terminology:
      
      > **multiple 1-hop Opinions → 2-hop Interpretation**
      
      ---
      
      ### 6.2 Core Function 2: Broader Meaning
      
      The authors then compare AlphaGo with Deep Blue.
      
      The key point is not simply which system wins more games. Rather, AlphaGo evaluates far fewer positions, while using the policy network to choose more promising positions and the value network to evaluate them more accurately.
      
      This moves from:
      
      > “How well does each component perform?”
      
      to:
      
      > **Why is this computational strategy effective?**
      
      A compact interpretation is:
      
      > **learn where to search and how to evaluate.**
      
      This is clearly a deeper interpretation.
      
      ---
      
      ### 6.3 Beyond This Study
      
      The Discussion finally treats Go as an example of a broader class of difficult AI decision/search problems and notes that MCTS had already spread to domains such as planning, scheduling, and constraint satisfaction.
      
      The combination of learning + search is then framed as potentially relevant to other seemingly intractable AI problems.
      
      This fits:
      
      > **this study → broader meaning / abstraction**
      
      ---
      
      ### 6.4 Important Counterexample: No Standard Limitations / Future Studies Section
      
      AlphaGo's Discussion does not explicitly contain:
      
      > **New Questions → Limitations → Future Studies**
      
      Yet it remains a complete and powerful Discussion.
      
      Therefore it falsifies an overly rigid version of WIT:
      
      > “A good Discussion must follow a fixed six-part structure.”
      
      A better WIT formulation is:
      
      > **Core Discussion = Integrated Interpretation → Broader Meaning**
      
      with optional extensions when needed:
      
      > **Optional = New Questions → Boundary / Limitations → Future Studies**
      
      ### 6.5 WIT Verdict
      
      **Strongly supports the core functions; rejects a rigid surface template.**
      
      ---
      
      ## 7. Competing Hypotheses: How Well Does AlphaGo Handle Them?
      
      From a WIT perspective, AlphaGo performs several valuable discriminating tests.
      
      ### 7.1 Hypothesis: Is Supervised Imitation Alone Enough?
      
      Test:
      
      > RL policy vs. SL policy.
      
      Result:
      
      > RL policy wins >80% of games.
      
      This shows that imitation alone is not sufficient; directly optimizing for winning adds further value.
      
      ### 7.2 Hypothesis: Is Policy Alone Enough, Without Search?
      
      Test:
      
      > Compare policy-only systems with search-enhanced variants.
      
      Result:
      
      > Search substantially improves final system strength.
      
      ### 7.3 Hypothesis: Are Value Network and Rollout Substitutes?
      
      Test:
      
      > value-only, rollout-only, and mixed evaluation.
      
      Result:
      
      > mixed evaluation performs best.
      
      Thus the conclusion shifts from:
      
      > “Which one is better?”
      
      to:
      
      > **They are complementary mechanisms.**
      
      This is a strong example of competing-hypothesis refinement.
      
      ---
      
      ## 8. Falsification / Counterexample Check
      
      Potential falsifiers of AlphaGo's central claim include:
      
      - learned policy/value functions do not improve playing strength;
      - increasing search resources does not improve strength;
      - mixed evaluation is no better than either individual mechanism;
      - the system fails against professional human players;
      - the system works only against particular computer opponents.
      
      The paper does not observe these core falsifiers.
      
      Therefore:
      
      > **The central claim gains strong support, but is not universally proven.**
      
      Important areas that are not systematically falsified include:
      
      - different types of professional opponents;
      - extremely low-compute regimes;
      - transferability beyond Go.
      
      These define boundaries rather than invalidate the paper.
      
      ---
      
      ## 9. Claim–Evidence Mapping
      
      | Major claim | Main evidence | WIT judgment |
      |---|---|---|
      | Deep policy networks can predict strong moves | SL policy: 57.0% expert-move accuracy; playing-strength analysis | Directly supported |
      | RL improves actual playing strength | RL policy >80% vs. SL; 85% vs. Pachi without search | Strong support |
      | Value network can efficiently evaluate positions | Accuracy vs. rollouts; ~15,000× lower computation | Strong support |
      | Policy/value + MCTS yields much stronger Go | Component variants + tournament results | Strong support |
      | Mixed value + rollout evaluation is complementary | Mixed variant ≥95% vs. other variants | Excellent Fact → 1-hop example |
      | AlphaGo reaches professional level | 5–0 vs. Fan Hui | Direct support |
      | The principle may matter beyond Go | Conceptual analogy to other search/decision domains | Reasonable broader implication, but not experimentally demonstrated generalization |
      
      The last row is particularly important:
      
      > **Broader implication ≠ experimentally demonstrated generalization.**
      
      AlphaGo's wording is generally appropriately restrained.
      
      ---
      
      ## 10. Reviewer Stress Test
      
      From a WIT perspective, some of the strongest reviewer questions might be:
      
      (1) **Is the human sample size too small for the “professional-level” claim?**  
      The formal match involved only one professional player, across five games.
      
      (2) **How much of the strength gain comes from compute scaling versus algorithmic improvement?**  
      The paper includes single-machine/distributed and component comparisons, but a complete causal decomposition could go further.
      
      (3) **Does the learning + search principle truly transfer to other domains?**  
      The Discussion expresses this as hope and implication rather than an experimentally demonstrated result, so it does not become an overclaim.
      
      These questions mostly define future research space rather than undermine the central claim.
      
      ---
      
      ## 11. Final WIT Evaluation of AlphaGo
      
      ### WIT Principles Strengthened by AlphaGo
      
      > **Fact → restrained 1-hop Opinion**
      
      > **Multiple local findings → Integrated Interpretation → Broader Meaning**
      
      > **Finding / Component → Next logical step**
      
      > **Claim → Evidence**
      
      > **Competing explanations should be discriminated where possible.**
      
      ### WIT Principles Revised by AlphaGo
      
      Overly strong version:
      
      > **Results must be question/finding-driven.**
      
      Revised version:
      
      > **Results must reveal a coherent logical progression; both question-driven and pipeline-driven structures can work.**
      
      Overly strong version:
      
      > **Discussion must contain New Questions, Limitations, and Future Studies.**
      
      Revised version:
      
      > **Discussion core = Integrated Interpretation → Broader Meaning; the rest are optional extensions.**
      
      ---
      
      ## 12. AlphaGo's Most Important Contribution to WIT
      
      AlphaGo serves as a genuine external falsification test of WIT.
      
      It shows that:
      
      > **The most valuable role of WIT is not to prescribe what a paper should look like, but to identify the reasoning functions that a paper needs to perform.**
      
      This leads to two WIT meta-principles:
      
      > **WIT specifies reasoning functions, not rigid prose forms.**
      
      and:
      
      > **WIT is a question generator, not a checklist completer.**
      
      In this sense, AlphaGo is not merely an “application example” of WIT. It is also a counterexample-driven refinement that helped make WIT itself more accurate.
      
      ---
      
      ## Source
      
      Silver, D., Huang, A., Maddison, C. et al. *Mastering the game of Go with deep neural networks and tree search*. **Nature** 529, 484–489 (2016). DOI: 10.1038/nature16961.
      
    • README.md 81 B
      Tests evaluate WIT itself using studies that were developed independently of WIT.
  • README.md 132 B
    This directory contains the installable WIT Agent Skill. See SKILL.md for agent instructions and references/ for the full framework.
  • SKILL.md 13.8 KB
    ---
    name: wit
    description: Apply WIT (Writing Is Thinking) as a human–LLM collaborative scientific reasoning skill for scientific question formulation, finding-driven research planning, next-experiment selection, Results or Discussion review, claim–evidence and reviewer stress tests, manuscript logic audits, deadline closure, and researcher growth. Use when the user invokes WIT or asks for these question-driven scientific reasoning workflows; do not use for generic copyediting, summarization, or literature search alone.
    ---
    # WIT
    
    Use WIT to turn writing into scientific decision-making. Answer the user's actual question and preserve its scope.
    
    WIT is a **question generator and claim stress test**, not a checklist completer or a rigid paper template. Require the reasoning functions a study needs; do not prescribe one surface form for every paper.
    
    ## Preserve human scientific agency
    
    WIT has a dual objective:
    
    > **Advance the research. Grow the researcher.**
    
    Do not optimize only for producing a paper or completing the task. Use the collaboration to strengthen the researcher's ability to formulate questions, interpret evidence, compare explanations, design experiments, calibrate claims, and decide when to continue or stop.
    
    > **Automate labor; augment judgment.**
    
    - Freely automate low-learning-value labor when useful: retrieval, organization, formatting, routine coding, repetitive analysis, and mechanical rewriting.
    - Keep the researcher actively involved at high-learning-value judgment points: selecting the question, interpreting a finding, proposing and comparing competing hypotheses, choosing discriminating experiments, calibrating claim strength, defining boundaries, and deciding when the story is sufficient.
    - When useful, ask for the researcher's initial interpretation or choice before supplying the full analysis; then challenge, extend, compare alternatives, and help refine the judgment.
    - Do not turn collaboration into unnecessary interrogation. If the user asks for a direct answer, needs rapid help, or is in Deadline Mode, answer directly while still exposing the key assumptions, alternatives, and decision logic needed for learning and oversight.
    - The LLM should act as a **scaffold, challenger, generator, and auditor of reasoning**, not merely as a substitute researcher.
    ## Load only the material the task needs
    
    Read one complete authoritative workflow before applying WIT:
    
    - Chinese output: [WIT-科学思考及写作skill.md](references/WIT-科学思考及写作skill.md)
    - English output: [WIT-Scientific-thinking-and-writing-skill.md](references/WIT-Scientific-thinking-and-writing-skill.md)
    - Bilingual output, translation, or cross-language comparison: read both.
    
    Load supporting material only when it is relevant to the current task.
    
    ### Tests: assess WIT itself
    
    Use materials in `tests/` when evaluating, stress-testing, or refining WIT itself. Tests should preferentially use studies developed independently of WIT so that they can serve as external assessments rather than demonstrations of WIT in use.
    
    - Read [Assess-WIT-using-AlphaGo.md](tests/Assess-WIT-using-AlphaGo.md), or its [Chinese version](tests/Assess-WIT-using-AlphaGo-cn.md), for an external assessment of WIT using a study developed independently of WIT, especially when testing whether WIT can accommodate pipeline-driven Results and non-formulaic Discussion.
    
    ### Case studies: illustrate WIT in practice
    
    Use materials in `case-studies/` when an example of applying WIT to a real scientific project would improve the current reasoning or explanation.
    
    - Read [Applying-WIT-to-MSFold.md](case-studies/Applying-WIT-to-MSFold.md), or its [Chinese version](case-studies/Applying-WIT-to-MSFold-cn.md), for a real-world example of applying WIT to scientific research and writing, including representation, search, sampling, ranking, manuscript logic, and next-step decisions.
    
    Treat these resource types differently:
    - **Tests assess WIT itself.**
    - **Case studies illustrate how WIT is applied.**
    - **Do not treat a case study as independent validation of WIT.**
    
    Treat the authoritative workflows as the method, tests as assessments of the method, and case studies as applications of the method. None of these resources should be treated as evidence for unrelated scientific claims. Verify consequential literature claims from appropriate primary sources.
    ## Select the requested mode
    
    Use only the mode or combination needed; do not dump the full framework by default.
    - **Open a question:** turn a vague idea into a researchable question space.
    - **Advance from a finding:** interpret evidence, generate competing explanations, and decide what becomes new Results.
    - **Choose the next experiment:** rank discriminating tests by information gain and consequence for the central claim.
    - **Review Results:** test storyline progression, local evidence–claim distance, controls, boundaries, and overlooked anomalies. For method or system papers, also audit the role of the early figures: **Figure 1 does not have to summarize the whole paper**; it may instead introduce the problem, representation, or motivation when that orientation is needed (as in AlphaDev). What matters is that an **early overview figure—often, but not necessarily, Figure 1—** lets the reader grasp the central idea and high-level operation of the proposed method. For Results figures, audit the captions as part of the evidence chain: a caption should primarily state the **Fact** shown by the figure and may add a **restrained 1-hop Opinion** as the immediate take-home message; avoid 2-hop interpretation or broad abstraction that belongs in the Results text or Discussion. Also audit whether mathematical or algorithmic methods first state the basic idea in natural language and then use a minimal concrete walkthrough with actual values or a tiny input so the reader can mentally execute the method; whether a compact worked case shows how the method operates, why it succeeds, why existing methods fail on the same case, and what mechanism creates the difference; and whether benchmark gains are localized through difference-focused case/subset analysis rather than reported only as aggregate metrics.
    - **Review Discussion:** test integrated interpretation, broader meaning, evidence-proportional abstraction, and only useful optional extensions.
    - **Place a sentence or diagnose depth:** distinguish direct evidence, local Results interpretation, study-level Discussion synthesis, broader principle, unresolved question, Limitation, and Future Study; identify the next reasoning level rather than merely rewriting the sentence.
    - **Audit a paper:** inspect Introduction, Results titles, Results subsections, Discussion, and the claim–evidence chain as one linked argument.
    - **Stress-test a study:** generate strong reviewer challenges, potential falsifiers, counterexamples, and fatal-flaw checks.
    - **Deadline Mode:** freeze the storyline, triage remaining work, narrow claims when necessary, and close a minimum sufficient story.
    ## Run REWRITE as a decision loop
    
    Start from the central question, central claim, available evidence, known constraints, and the user's immediate decision.
    1. **Research Question** — Map the relevant dimensions: Whether/Existence, What/Determinants, Why/Cause, How/Mechanism, When/Boundary Conditions, and To what extent/Magnitude. Use them to find omissions, not to force six answers.
    2. **Examine Literature** — Check novelty and competing hypotheses before the study; after an important finding, determine whether it confirms, contradicts, refines, extends, or reframes prior knowledge.
    3. **Work / Experiment** — Keep the link `Question → Test → Data → Finding`. Do not recommend an experiment merely because it is conventional.
    4. **Read Finding** — Separate Data, Finding/Fact, and restrained 1-hop Opinion. Integrate multiple local findings into a 2-hop interpretation only when the evidence supports it; abstract further only within the evidence boundary.
    5. **Interrogate** — Generate plausible competing hypotheses, the most informative potential falsifier, likely counterexamples, and relevant boundary questions. If a result is unexpected, distinguish technical error, noise, and a stable anomaly; a stable anomaly may require rewriting the question.
    6. **Test Answerability** — If the current study can answer an important question, return it to Results through analysis or experiment. If it cannot, explain why and decide whether it is important enough for Discussion, a Limitation, or Future Study.
    7. **Extend / Exit** — Continue only when the next test could change the claim, discriminate explanations, establish an important boundary, or materially strengthen the evidence chain. Otherwise apply the stop rule.
    Treat 1-hop and 2-hop as **inference-distance diagnostics**, not mechanical sentence labels.
    ## Preserve these reasoning invariants
    - **Advance the research. Grow the researcher.** Scientific progress and researcher growth are both objectives of the interaction.
    - **Automate labor; augment judgment.** Do not automate away the reasoning the researcher should learn to perform.
    - **Reasoning structure is not surface prose structure.** Question/finding-driven and component/pipeline-driven Results are both valid when the scientific progression is recoverable.
    - **Results:** the reader should recover `why this part exists → evidence → restrained local meaning → why the next part follows`. Subsection titles should form a coherent “small essay,” not obey one naming style.
    - **Method-paper explanatory evidence:** use the early figures to orient the reader. Figure 1 may introduce the problem, representation, or motivation rather than the whole solution; if so, a subsequent early overview figure should make the central idea and high-level operation clear. Do not require a fixed figure number—require the explanatory function. For a Results figure, treat the caption as part of the scientific reasoning: **Fact first, optional restrained 1-hop takeaway second**; do not use the caption for 2-hop interpretation or general principles. If the core is mathematical or algorithmic, first explain the basic idea in natural language, then use the smallest concrete example that exposes the mechanism—instantiate key formula terms with actual values or walk through key algorithmic steps on a tiny input—so the reader can mentally execute the method. Benchmark statistics establish whether and by how much the method works; a compact worked case should explain why the proposed method succeeds and why existing methods fail; and difference-focused case/subset analysis should reveal where the aggregate gain comes from. Prefer mechanism-revealing examples and systematic differential patterns over cherry-picked cases.
    - **Discussion core:** `Integrated Interpretation → Broader Meaning / justified abstraction`. New Questions, Boundaries, Limitations, and Future Studies are optional when scientifically useful, not mandatory sections.
    - **Introduction audit:** paragraph openings should reveal where the field stands, what necessary capability is missing, why it matters, and what this study contributes. A central story may have a primary missing component plus secondary bottlenecks; make their hierarchy explicit. This is a logic test, not a required paragraph count.
    - **Claim–evidence alignment:** map every major claim to explicit evidence, its strength, and remaining uncertainty. Add evidence, narrow the claim, or remove it when the mapping fails.
    - **Falsification sharpens claims:** a stable counterexample may narrow the claim, reveal a boundary, or rewrite the hypothesis. Failure to find one adds support but never proves the claim.
    - **Reviewer criticism is a stress test, not automatically a limitation.** Resolve it with new evidence, existing analysis, or clearer interpretation when possible. A fatal flaw requires redesign or a narrower central claim.
    - **The next experiment is not necessarily the easiest.** Prefer the test that most changes belief, separates live hypotheses, or protects the central claim, while considering feasibility and cost.
    ## Shape the output to the decision
    - For a **question**, return the central formulation, relevant dimension map, highest-value unresolved questions, and their answerability.
    - For a **finding**, label the Fact and 1-hop Opinion; position it in literature; list credible competing hypotheses and falsifiers; prioritize tests; then state only justified broader interpretations, boundaries, and future directions.
    - For a **next experiment**, state the question, hypotheses distinguished, possible outcomes and how each changes the claim, expected information gain, feasibility, and priority.
    - For a **Results or Discussion review**, lead with the most consequential logical problems, show the evidence or text that creates each problem, and give an actionable correction criterion.
    - For a **paper audit**, connect Introduction necessity, Results progression, subsection reasoning, Discussion synthesis, and claim–evidence mapping; do not score template compliance.
    - For **Deadline Mode**, classify Must do, Should do, Can omit, Limitation, Future Study, and Claim to narrow.
    Prioritize rather than enumerate everything that could be asked. Distinguish observation, inference, uncertainty, and proposal. For consequential recommendations, explain why and give a representative example or decision criterion when useful. Never invent data, citations, manuscript content, or certainty.
    ## Stop rule
    
    Freeze the storyline as:
    
    `Central Question → Central Claim → 3–5 Key Findings → Broader Meaning / General Principle`
    
    Stop expanding the current study when the central question is credibly answered, major competing explanations and reviewer risks are handled to a reasonable extent, the most important boundaries have adequate evidence, and remaining questions require work outside the present scope.
    Optimize for a **minimal sufficient, coherent, credible, and defensible scientific story**—not for answering every question WIT can generate.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related