The Skill Is Real, the Ritual Is Not Yet Measured

Curated skills measurably improve coding agents where they carry procedure the model lacks; the process disciplines the frameworks sell remain unmeasured — and the nearest tests lean the other way.

Outline
THE BRIEF
THE QUESTION

Skills frameworks like superpowers load process disciplines into a coding agent's harness. Do they improve what an agent can do, or teach it to perform the shape of a process — and does that depend on model capability?

WHY IT MATTERS

The skills mechanism is a vendor-supported feature of the major coding-agent harnesses, the frameworks built on it are spreading fast, and teams are deciding now whether to standardize on one.

THE ANSWER

Skill loading as a mechanism is real: curated skills lift mean pass rates from 33.9% to 50.5%, largest where the skill carries procedure the model lacks [9]. The process disciplines themselves have no such record: no published same-model ablation shows a statistically supported task-outcome improvement from a superpowers-class framework, on any tier. The closest measurements lean the other way: superpowers indistinguishable from baseline on process quality [13], one discipline delivered as pure instruction making a small model worse [14]. The capability gradient is not one gradient but an inverted-U, which the report walks; how process-discipline layers behave across it remains, precisely, unmeasured.

What a skills framework actually installs

A skill is prose, not code. In the canonical implementation, each skill's name and a short description sit permanently in the agent's system prompt; the full instruction body loads only when the model itself judges that the description matches the task at hand [8]. Everything downstream of that judgment is persuasion. The superpowers framework — at v6.3.0 as of August 12, 2026, shipping both of the skills at issue here, brainstorming and systematic-debugging [1, 2] — enforces its methodology with capitalized imperatives, not gates: the entry skill instructs that "if you think there is even a 1% chance a skill might apply to what you are doing, you absolutely must invoke the skill" [3]. The brainstorming skill classifies every request into three paths and holds implementation behind an explicit human-approval gate [4]. The systematic-debugging skill opens with an iron law — "NO FIXES WITHOUT ROOT CAUSE INVESTIGATION FIRST" — and rebuts anticipated rationalizations in a red-flags table [5]. No mechanism compels compliance with any of it.

Two properties of the artifact matter for everything that follows. The first is its own theory of value: the framework's circulated self-description is that it "doesn't give Claude new capabilities. It gives Claude discipline" [7] — a claim about behavior shaping, not capability creation, and one made without measurement. The systematic-debugging skill goes further and asserts, also without measurement, that the discipline "is FASTER than guess-and-check thrashing" [5]. Second, the framework's own quality bar tests process, not outcomes: its author describes "basic end to end tests that make sure that the agent runs the full brainstorming - planning - implementing flow and verifies skill usage," explicitly "not a proper evals suite" [6]. The artifact checks whether the ritual ran, not whether the ritual helped.

The loading mechanism is also, by the author's own account, sensitive to model version: beginning with Opus 4.5, Claude became likelier to guess a skill's content from its description — "claiming it was going to use a skill and then....winging it without actually reading the skill" — forcing a rewrite of every description into pure when-to-use triggers [6]. A framework whose delivery mechanism must be re-tuned per model generation is not a model-independent investment; this is first-party, testimonial-class evidence, but it comes from the party with the least incentive to volunteer it.

Skill loading in general: measured, real, and two-shaped

Whatever is true of process rituals, skill loading as a mechanism now has a substantial controlled record, and it is consistent across independent measurements.

TABLE 1 The 2026 controlled record on skill loading — each study's design, headline result (deltas in percentage points, pp), and main caveat.
Study (2026)DesignHeadline resultMain caveat
SkillsBench [9]87 tasks, 18 model–harness configs, paired no-skill vs curated-skill, deterministic verifiersMean pass rate 33.9% → 50.5% (+16.6 pp; range +4.1 to +25.7)Peer-review status unconfirmed; curated skills, idealized loading
Tessl skill-evaluation framework [10]~500 real skills, ~1,000 generated tasks, 19 configs, LLM-judgedOverall-score gains on every model; instruction-following moves far more than goal completionAuthors affiliated with the registry the skills came from; tasks generated from skill content — the design "inherently favors the with-skill setup"
Agentic skills in the wild [11]Progressive realism on SkillsBench tasks, three model–harness stacksGains degrade monotonically toward baseline as loading becomes realisticSame task suite as SkillsBench
Harness-Bench [12]106 tasks × 6 harnesses × 8 model backends, 5,194 trajectories23.8-point aggregate gap between best and worst harness under identical models and tasksMeasures whole configurations; explicitly not a skill-layer ablation

Two findings inside this record do most of the analytical work.

Gains are capability-shaped where procedural knowledge is missing

SkillsBench's gains are largest precisely where the model could not have known the procedure: a prefix-cache-replay task jumps from a 1.9% to a 94.4% pass rate with the skill, and the domain gradient runs from Natural Science (+28.8 pp) down to Software Engineering (+11.6 pp) and Mathematics/OR (+9.7 pp) — the domains where pretraining and tooling coverage are already strong benefit least [9]. These are deterministic pass/fail verifiers, so this is capability in the strict sense: tasks that failed now pass. The mechanism has sharp edges, though. Thirteen of 87 tasks get worse with skills; one to three focused skills gain +18.0 to +19.0 pp while four or more gain only +10.1 pp; compact and standard-length skills gain +19.0 and +21.5 pp, detailed documentation +14.5 pp, and comprehensive prose documentation +0.7 pp — essentially nothing [9]. Skills written by the agent itself land below the no-skill baseline (−8.1 to −11.5 pp across three configurations) [9]. The value is in curated, compact, verifier-relevant procedure — not in volume of instruction, which is what the most ceremony-heavy process skills are.

Gains are compliance-shaped where the model can already do the task

The vendor-affiliated Tessl study is the only measurement that decomposes the gain, and its decomposition gives the ritual-versus-substance question a measured vocabulary. On its ~1,000 tasks, goal completion is near-saturated with or without skills for almost every strong model — Opus 4.8 moves from 93.3 to 97.5 — while instruction following moves enormously, 59.8 to 88.0 for the same model [10]. On work the model can already do, a skill mostly changes how it does it: conventions honored, opinionated workflows followed. That is not worthless — the study argues, plausibly, that ignoring encoded instructions degrades quality, safety, and maintainability even when the task nominally succeeds — but it is compliance, and it should not be booked as capability. Read jointly, the two studies say: skills deliver capability where they carry absent procedure, and deliver conformance where they encode preference. Process-discipline skills are, by construction, almost entirely the second kind.

The process disciplines themselves: a record close to empty

For the frameworks the question actually names, the efficacy record is close to empty, and what exists points the wrong way for the strong claim.

The superpowers repository ships no benchmark and no quantitative efficacy claim; its efficiency numbers for the 6.0 release ("roughly twice as fast and while spending almost 50% fewer tokens") come from the project's own eval harness with the authors' own caveat that "these numbers won't hold on every harness and for every workload" [2]. The surrounding enthusiasm — hundreds of thousands of GitHub stars, endorsements, experience essays — is testimonial throughout, and the author's README concedes there is no usage telemetry: "we have no idea how many of you are using Superpowers" [1]. Anthropic's own engineering material on the skills mechanism offers authoring and testing guidance but no controlled measurement in its published discussion of developing and evaluating skills [8]. None of this is evidence of absence of effect; it is absence of evidence, from every party with an interest in the effect existing.

The three nearest controlled datapoints, in ascending order of weight:

1. A registry-run ablation of the brainstorming skill: no effect. The Tessl registry's evaluation of superpowers' flagship skill reports 1.00x agent success versus baseline — on a single scenario, far too small to carry a verdict in either direction, but the only ablation-shaped number that exists for the focal framework's focal skill [15].

2. A same-model harness comparison: superpowers ≈ baseline. RigorBench, a 2026 benchmark of engineering process discipline, ran superpowers, another skills collection, the authors' own harness, and a bare ReAct control on the same pinned model (Gemini 3.5 Flash, medium thinking, 100 tasks, ~410 executions) [13]. Superpowers scored 0.41 on process quality against the ReAct baseline's 0.40 — the paper's own reading is that "tool-enforced frameworks like Agent-Skills (0.39) and Superpowers (0.41) perform similarly to the baseline" — with a modest outcome-score edge, 0.70 versus 0.64 [13]. The result carries heavy caveats: the winning harness belongs to the paper's authors, the paper characterizes superpowers loosely, and the model is a single pinned mid-tier configuration. But as the only same-model comparison in print, it shows the flagship process framework failing to move the very construct — process discipline — it exists to install, while the authors' planning-enforced harness (0.53 process, 0.83 outcome) did move both [13]. Enforcement in the control loop, not prose in the context window, is what separated the winner from the baseline.

3. The only direct test of a discipline-as-prompt: negative. TDAD (2026, single author, SWE-bench Verified, small local models) is the one controlled measurement of a process discipline delivered the way skills deliver it — as loaded instruction. TDD prompting alone increased test-level regressions from 6.08% to 9.94%, worse than no discipline at all (Qwen3-Coder 30B, n=100). The same discipline delivered as context — a dependency-graph test map telling the agent which tests its change might break — cut regressions to 1.82%, and deployed as an agent skill improved resolution from 24% to 32% (Qwen3.5-35B-A3B, n=25) [14]. The author's summary is the sharpest sentence in this literature: agents "do not need to be told how to do TDD; they need to be told which tests to check," because "surfacing contextual information outperforms prescribing procedural workflows" [14]. One small-N, small-model study cannot settle the question — but its direction is exactly the ritual hypothesis, measured.

The finding, stated at full strength and no further: as of August 2026, the cited literature contains no published same-model ablation showing a statistically supported task-outcome improvement from a superpowers-class process-discipline framework, on any model tier — RigorBench's outcome edge (0.70 versus 0.64, overlapping intervals, no significance test reported for the pair) included [13]. The closest measurements show approximately no effect on the process-discipline construct itself on a pinned mid-tier model, no effect at N=1 for the flagship skill, and active harm when one such discipline was tested as pure instruction on a small model. A load-bearing efficacy claim with no measured support is a measured absence, and this is the central one in the field.

Which way does the capability gradient run?

The question's second half — do weak models gain what strong models don't need, or the reverse — presumes one gradient. The measured record contains at least two, running in opposite directions, plus a frontier-era inversion, and the direction depends on what the scaffold does.

The 2022–24 ancestry: procedure-execution had a floor, in that era

The prompting-technique lineage is the measured ancestry of the skills intervention class, and its era must stay attached to its findings. Chain-of-thought's founding measurement (Wei et al., 2022; PaLM, GPT-3, and LaMDA model series) found the technique "does not positively impact performance until used with a model of sufficient scale" and "actually hurts performance for most models smaller than 10B parameters" — of that generation [16]. Self-Refine (2023; GPT-3.5, ChatGPT, GPT-4) found gains ordered by model strength — "GPT-4 + Self-Refine performs better than GPT-3.5 + Self-Refine" across all tasks — and its sharpest datum is a weak model failing to execute the ritual at all: Vicuna-13B, a 2023-era open 13B model, "was not able to consistently generate the feedback in the required format" — and even when handed oracle feedback, it failed to execute the refinement step [18]. Huang et al. (2023; GPT-3.5, GPT-4) found intrinsic self-correction degrading reasoning: models "struggle to self-correct their responses without external feedback, and at times, their performance even degrades" [19]. These are findings about that era's models, and they do not license present-tense predictions in either direction: no published study runs the prompting-era scaffolds and skill-file loading on the same model grid, so nothing permits using the 2022–24 curve to forecast skill-era behavior. The current stratum must be read on its own measurements — and it now has them.

The reasoning-model inversion: imposed scaffolds began to harm the strong

By late 2024 the gradient inverted at the top. The Medprompt-to-o1 study (November 2024, o1-preview) found the best 2023-era prompt scaffold obsolete against a reasoning-native model — o1-preview without prompting techniques largely outperformed scaffolded GPT-4 — and found that "few-shot prompting hinders o1's performance, suggesting that in-context learning may no longer be an effective steering approach for reasoning-native models" [20]. Mind Your Step (2024) measured the harm directly on adverse task classes: on tasks where deliberation hurts humans, o1-preview — whose inference-time reasoning is built in and cannot be removed — lost up to 36.3 percentage points of absolute accuracy against zero-shot GPT-4o, a cross-model measure of reasoning harm rather than of an imposed scaffold [21]. Meanwhile the largest audit of the technique (110 papers, 1,218 comparisons, plus fresh runs on 20 datasets and 14 models) found chain-of-thought's benefit concentrated on symbolic, math, and logic tasks (average +14.2, +12.3, +6.9 points) with little elsewhere — the interaction lives in the task type more than the model tier [17]. Scaffold value, in this lineage, migrated into the model.

The 2026 stratum: an inverted-U with a floor — and a sign flip under deployment realism

The current era's harness measurements tell a defect-compensation story at the top. A 2026 grid holding one harness-evolution recipe fixed across eight languages and three rollout models found that "gains compensate recoverable execution defects" and vanish where discipline is already present — "GPT-5-mini commits few such defects at all," so its cells show nulls [22]. A Nature Machine Intelligence study across 260 configurations found "single-agent baseline performance emerges as the most robust predictor of whether coordination improves or decreases performance": weaker models benefit from coordination scaffolds and capable models outgrow them [23]. For scaffolds that compensate defects, the ceiling edge is measured: strong models take little from what they already do.

The skill-stratified data itself, though, is not monotone in either direction. Across SkillsBench's 18 configurations, the largest lifts go to mid-tier stacks — GLM 5.1 (+25.7 pp), Gemini 3.1 Pro (+24.8 pp), DeepSeek V4 Pro (+23.2 pp) — while the flattest gains sit at both ends: Gemini 3.1 Flash Lite manages +4.1 pp, and Gemini 3.5 Flash, starting from a strong 41.1% no-skill baseline, gains only +7.1 pp; the paper's own reading is that "high base capability does not imply high Skill leverage" [9]. The Tessl study finds "the relative impact of skills is larger for smaller models than for bigger ones" within a family — Haiku and Sonnet versus Opus — but with a hard floor beneath it: its weakest models (the Nemotron family) "barely benefit at all," and Kimi K2.6 gains only 7.1 points because it "does not utilize the skill's content properly" [10]. The two results compose into an inverted-U: within capable families, smaller variants gain relatively more; below an instruction-following floor, gains collapse regardless of parameter count. At the frontier, the current version of the benchmark states the ordering as a finding in its own right: "The strongest absolute systems are not always the largest beneficiaries." The highest with-skills pass rates belong to frontier stacks — GPT-5.5 reaches 67.3% — while Claude Opus 4.8, starting from the strongest Claude no-skill baseline (45.7%), gains only +8.4 pp under the same harness where GLM 5.1 gains +25.7 [31]. That ordering is version-bound and recent: the February 2026 release (v1, 84 tasks, 7 configurations) had reported the largest within-model lift at the frontier — Claude Code with Opus 4.5, at +23.3 pp [33] — and the current v4 (87 tasks, 18 configurations) reverses it [31]. Frontier configurations still gain, most in double digits, but which tier gains most has proven to be a property of the measured grid, not of the mechanism.

The widely circulated "equalizer effect" has to be read against that table, because it comes from the same one. In the current version, MiniMax M2.7 with curated skills passes 34.9% of tasks — above stronger no-skill baselines under the same harness: GLM 5.1 at 32.7% and MiniMax M3 at 29.7% [31]. (The illustration itself has churned across the paper's revisions — the February 2026 v1 made the same point with Claude Haiku 4.5 with skills beating Opus 4.5 without — while the shape of the claim held.) That is a real and decision-relevant result — but it is a claim about the cost-performance frontier: a cheaper model plus curated skills can outscore stronger models without them. It is not a claim that small models gain most: the same table gives its largest lifts to mid-tier stacks, and the paper's own finding warns that the strongest systems are not always the largest beneficiaries [31]. Circulated as a claim that skills close the capability gap, the equalizer promises more than its table delivers — the with-skills frontier still sits far above every cheaper stack.

A factorial measurement sharpens the ceiling edge into statistics. ClawsBench (44 productivity tasks, 6 models, 4 harnesses) varies skills and a meta-prompt independently: for the weakest model tested, the two scaffolds are approximately additive; for more capable models, "either scaffold alone lifts TSR from near-zero to ∼55–60%" (TSR: task success rate) and adding the second yields little — strong negative interactions (p ≤ .003) [32]. Two scope caveats keep this in its lane: the paper calls its near-zero baselines "an information floor, not a capability floor," and its skills are designed to inject API knowledge [32]. Read as a scope statement, this measures knowledge-scaffold saturation, not process-discipline enforcement. Read correctly, it gives "strong models don't need it" its only statistically tested form — as scaffold redundancy saturating, not scaffolding being useless.

Deployment conditions then break the picture asymmetrically. When skills must be retrieved from a large pool rather than hand-delivered, the sign flips against the weak: Kimi K2.5 (19.8% versus 21.8% baseline) and Qwen3.5-397B-A17B (19.7% versus 20.5% baseline) fall below their no-skill baselines on irrelevant retrieved skills, while Claude Opus 4.6 degrades from 55.4% to 38.4% but keeps a +3.0-point margin over its baseline — the study's reading is that "stronger models can better ignore irrelevant skills, while weaker models are more likely to be hurt by low-quality retrieved skills" [11]. This is the current stratum's sharpest datapoint against the thesis that skills are a compatibility shim for weak models: under deployment-shaped conditions, the models the thesis says should benefit are the ones harmed, and only the frontier model retains any gain. The measured 2025–26 shape, in one sentence: an inverted-U with an instruction-following floor, largest relative gains to smaller-but-capable variants, frontier gains that persist under curated loading but no longer lead — and a realism penalty that lands almost entirely on the weak.

What does not exist, anywhere on that curve, is a capability-stratified measurement of a process-discipline framework itself: the benchmark built for process discipline pins a single model by design to isolate the harness effect [13], and the one prompt-delivered discipline test ran on a single small model [14] — so the model axis for superpowers-class frameworks is not merely unmeasured; the field's own instruments removed it.

Deployment: where measured gains leak away

Every number in the paired-ablation record assumes the right skill reaches the context window. The measured degradation from that assumption is steep. On SkillsBench tasks with Claude Opus 4.6, force-loading curated skills yields 55.4%; letting the agent choose from the same skills drops it to 51.2%; adding distractor skills, 43.5%; requiring retrieval from a 34k-skill pool, 40.1%; removing curated skills from the pool entirely, 38.4% — three points above never having skills at all [11]. Only 49% of trajectories load all curated skills even when directly provided [11]. Engineering can claw much of this back — retrieval plus query-specific skill refinement lifted the same model from 57.7% to 65.5% on Terminal-Bench 2.0 [11] — but the headline is that idealized deltas do not survive realistic loading, and a growing library is a liability: expanding skill collections measurably degrades performance through skill shadowing [30], consistent with SkillsBench's four-or-more-skills falloff [9].

The failure record around the flagship framework is congruent, and mostly primary-sourced from its own trackers. The dominant reported failure is not hollow performance but silent non-invocation: the mandatory brainstorming skill un-invokable due to a configuration flag [26], skills never firing when the harness's Plan Mode takes precedence [25], marketplace skills silently failing to load on one platform while working on another [27]. Nothing errors when a skill fails to trigger — which also poisons the testimonial record, since a user praising the framework cannot know whether it was active. When skills do load, the mandate is documented as non-binding: the framework's own issue tracker records the agent writing implementation first and tests after despite an explicit CLAUDE.md TDD requirement, with documentation-only enforcement marked "already tried, agent still skips TDD" [24], and the harness vendor's tracker shows a checklist marked CRITICAL skipped wholesale [28]. The framework's own release notes rewrote its review process to be "cheaper, stricter, and harder to game" and rescaled the brainstorming ceremony so "small tasks skip the two-document ritual" [2] — first-party concessions that the prior mandatory process was being gamed and that uniform ceremony was overweight. Practitioner counter-testimony inside the enthusiasm venues ("I found superpowers a huge token guzzler. And more generic a skill is the worse it seemed to perform" [29]) conflicts with the mechanism's token-light design intent [8]; neither side of that conflict has been measured on superpowers itself.

What this means in practice

The evidence supports conditional adoption, not a verdict on the ecosystem.

  • Adopt curated, compact, domain-procedural skills where a verifiable procedure gap exists — brittle formats, niche pipelines, tooling the model predictably mishandles. This is the measured sweet spot: double-digit gains under paired evaluation, largest off the beaten path of pretraining [9]. Keep skills few and short; the measured penalty for volume and prose is severe [9].
  • Do not book capability gains from process-discipline frameworks, on any model tier. Their efficacy claims are currently unmeasured; the nearest measurements show nothing or harm [13, 14, 15]. Legitimate reasons to adopt one anyway exist — standardized specs, review gates, and audit trails have organizational value independent of pass rates — but that is a workflow-tooling decision and should be justified as one.
  • If you deploy skills at all, instrument invocation. The dominant failure mode is silent [25, 26, 27], loading is unreliable even when skills are provided [11], and without telemetry your team's testimonials measure belief, not behavior.
  • On frontier models, prefer context and enforcement over prescribed procedure. The 2024 reasoning-model record shows imposed reasoning scaffolds harming strong models (o1-preview era) [20, 21]; 2026 harness gains vanish where the model already has the discipline [22]; and the one benchmark where a process layer beat baseline did it by enforcing phases in the control loop, not by persuading in the prompt [13]. TDAD's direction — surface which tests to check rather than instruct how to do TDD — is the transferable design lesson [14].
  • On small and mid-tier models, buy the cost frontier, not a capability myth — and pilot before fleet rollout. Current-era paired data supports curated-skill gains for smaller-but-capable models, including a cheaper model with skills outscoring stronger models without (34.9% against 32.7% and 29.7% — a cost-frontier result, not a gap closure) [31]. But gains collapse below an instruction-following floor [10], and in retrieval-shaped deployment the weakest stacks fall below their no-skill baselines [11].
  • What would change this analysis: an independent, paired, same-model ablation of a named process-discipline framework on a frontier model at N≥100 with invocation telemetry; and a capability-stratified rerun varying model strength under a fixed skill layer. Either result, in either direction, would outweigh most of the testimony written to date.

Conclusion

Do skills frameworks fundamentally improve what an agent can do, or teach it to perform the shape of a process? For domain-procedural knowledge, the improvement is real, measured twice independently, and bounded: it appears where the skill carries procedure the model lacks, shrinks where the model already knows the domain, decays sharply under realistic loading, and inverts when libraries bloat. For the process disciplines that superpowers-class frameworks actually sell, the honest answer is that efficacy has not been measured. The perimeter of nearby evidence — a 1.00x single-scenario ablation, a same-model comparison indistinguishable from baseline on process quality, a discipline-as-prompt test that made things worse, and a framework whose own tests check ritual execution rather than outcomes — leans toward the performance-of-process reading for outcomes, while leaving the door open for enforcement-based designs that move the loop rather than the prompt. On the capability gradient, the era-disciplined record refuses a single direction: 2022–24 procedures gave that generation's weak models nothing or made them worse — the sharpest case could not run them at all — and 2024's strongest reasoning models were harmed by imposed ones. The 2026 skill-era measurements trace an inverted-U: collapse below an instruction-following floor, largest relative gains to smaller-but-capable variants, frontier gains that persist under curated loading but no longer lead, and a sign flip against the weakest models under deployment realism. The question that decides the investment thesis — how process-skill layers specifically behave across that curve — remains, precisely and consequentially, unmeasured.

Appendix: Limitations

Most of the controlled record here consists of 2026 arXiv preprints whose peer-review status is unconfirmed; the sole journal-attested result is the Nature Machine Intelligence collaboration study. Two load-bearing measurements carry conflicts of interest in opposite directions: the Tessl study's authors operate the registry its skills came from and note themselves that the task-generation design favors the with-skill condition, so its deltas are best read as upper bounds; RigorBench's authors benchmark their own harness against superpowers, and its characterization of that framework is loose enough that its near-null result should be weighed as one datapoint, not a verdict. The TDAD result — the only direct test of a process discipline as loaded instruction — is a single-author study on small local models at N=100 and N=25, without confidence intervals, and its direction needs independent replication before the specific figures are leaned on.

The evidence is concentrated in coding and terminal-agentic tasks; whether any of it generalizes to non-software agent work is not established by the cited studies. Temporal instability is structural: the loading mechanism is documented by its own author as behaving differently across model versions, several benchmarks pin models that will be obsolete within quarters, and the 2022–24 findings invoked as ancestry are explicitly era-bound. Finally, the strongest positive association between process discipline and outcomes in this record is correlational at the benchmark level; no study cited here demonstrates a causal path from loaded process instruction to improved task outcomes on a frontier model, and that remains the single measurement the field most needs.

How we verified

3% of the key figures in this report sit in sentences without a verified per-claim binding. We show this openly — trust comes from disclosing the checks that did not pass, too.

Per-claim audit · support verdicts

81 of 81 marker instances bound & audited: 32 stated · 49 grounded · 3 verified, shown via source excerpt

32 stated49 grounded
Figures traced to source · per-claim audit
97% 74 of 76 figures
Automated checks
Citation markers reconciled against the reference list Every cited URL verified against the evidence store Fact-to-citation attribution overlap checked Every named system grounded in a retrieved source Section citations confined to their pre-bound evidence set 2 figure(s) in clauses without a verified per-claim binding
retrieved 392
passed relevance screening 198
in the writer's working set 168
cited 33

Evidence reflects sources as of publication (2026-08-20); citations last re-verified 2026-08-25.

These checks establish citation traceability and internal consistency. They do not independently reproduce the underlying experiments, guarantee that third-party figures are correct, or ensure that volatile values — prices, model versions, benchmark results — have not changed since retrieval.

References
  1. obra/superpowers: An agentic skills framework (GitHub repository) github.com · captured 2026-08-19
  2. Superpowers RELEASE-NOTES.md (v6.0.0–v6.3.0) github.com · captured 2026-08-19
  3. superpowers/skills/using-superpowers/SKILL.md github.com · captured 2026-08-19
  4. superpowers/skills/brainstorming/SKILL.md github.com · captured 2026-08-19
  5. superpowers/skills/systematic-debugging/SKILL.md github.com · captured 2026-08-19
  6. Jesse Vincent, "Superpowers 4" blog.fsck.com · captured 2026-08-19
  7. Massively Parallel Procrastination, Mentions: 2026 blog.fsck.com · captured 2026-08-19
  8. Anthropic, "Equipping agents for the real world with Agent Skills" www.anthropic.com · captured 2026-08-19
  9. Li et al., "SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks" (arXiv 2602.12670v4, 2026) arxiv.org · captured 2026-08-19
  10. Shaposhnikov et al., "A Framework for Evaluating Agentic Skills at Scale" (arXiv 2606.17819v1, 2026) arxiv.org · captured 2026-08-19
  11. Liu et al., "How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings" (arXiv 2604.04323v1, 2026) arxiv.org · captured 2026-08-19
  12. "Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows" (arXiv 2605.27922v1, 2026) arxiv.org · captured 2026-08-19
  13. "RigorBench: Benchmarking Engineering Process Discipline in Autonomous AI Coding Agents" (arXiv 2606.22678v2, 2026) arxiv.org · captured 2026-08-19
  14. Alonso, "TDAD: Test-Driven Agentic Development arxiv.org · captured 2026-08-19
  15. Tessl skill registry, obra/superpowers evaluation page tessl.io · captured 2026-08-19
  16. Wei et al., "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models" (arXiv 2201.11903, 2022) arxiv.org · captured 2026-08-19
  17. Sprague et al., "To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning" (arXiv 2409.12183, 2024) arxiv.org · captured 2026-08-19
  18. Madaan et al., "Self-Refine: Iterative Refinement with Self-Feedback" (arXiv 2303.17651, 2023) arxiv.org · captured 2026-08-19
  19. Huang et al., "Large Language Models Cannot Self-Correct Reasoning Yet" (arXiv 2310.01798, 2023) arxiv.org · captured 2026-08-19
  20. "From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond" (arXiv 2411.03590, 2024) arxiv.org · captured 2026-08-19
  21. "Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse" (arXiv 2410.21333, 2024) arxiv.org · captured 2026-08-19
  22. "One Recipe, Many Harnesses: What Self-Evolution Encodes Across Languages and Models" (arXiv 2608.10178v1, 2026) arxiv.org · captured 2026-08-19
  23. Kim et al., "Capable language models can outgrow the benefits of collaboration" (Nature Machine Intelligence 8, 2026) www.nature.com · captured 2026-08-19
  24. obra/superpowers issue #384, "Automatic TDD Skill Enforcement Before Implementation" github.com · captured 2026-08-19
  25. obra/superpowers issue #1667, "Skills not triggered when Claude Code Plan Mode is active" github.com · captured 2026-08-19
  26. obra/superpowers issue #345, "Skill superpowers:brainstorm cannot be used with Skill tool due to disable-model-invocation" github.com · captured 2026-08-19
  27. anthropics/claude-code issue #64763, "Desktop (Windows): marketplace plugin skills never appear" github.com · captured 2026-08-19
  28. anthropics/claude-code issue #32198, "Claude Code skips mandatory rules in CLAUDE.md (Definition of Done)" github.com · captured 2026-08-19
  29. Hacker News, "Superpowers 6" discussion thread news.ycombinator.com · captured 2026-08-19
  30. "More Skills, Worse Agents? Skill Shadowing Degrades Performance When Expanding Skill Libraries" (arXiv 2605.24050v2, 2026) arxiv.org · captured 2026-08-19
  31. Li et al., "SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks" (v4, 87 tasks, 18 configurations; arXiv 2602.12670v4, 2026) arxiv.org · captured 2026-08-19
  32. "ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces" (arXiv 2604.05172v2, 2026) arxiv.org · captured 2026-08-19
  33. Li et al., "SkillsBench" v1 release (February 2026; 84 tasks, 7 configurations; superseded by v4 [31]) arxiv.org · captured 2026-08-19