Files
dotfiles/claude/skills/tooling-scan/SKILL.md
T

6.1 KiB


name: tooling-scan description: Survey the dev-tooling landscape for tools that could help US build CapableMind better — agent runtimes, memory substrates, inference infra, skill/diagnostic tooling. The build-lens sibling of /landscape-scan (which scans competitors). Generative but anti-shiny: a tool earns adoption only if it addresses a named bottleneck at acceptable switching cost with governance-grain fit. Surfaces candidates and bounded trials; never adopts unilaterally. Records to the tooling register.

Tooling Scan

Survey dev tooling that could help us build CapableMind — and keep the tooling register current. Sibling of /landscape-scan (which scans comparable/competing projects). The lenses are opposite: landscape asks "is someone building what we build?"; tooling asks "what helps us build it?"

Principle

Research is important, but it must push us to make something better — not defeat us. (steward)

Generative, but anti-shiny. The honest prior: our bottlenecks are usually not tool-shaped — they're hard engineering bugs, collaborator bandwidth, and the deliberately governed pace the work demands. A shinier tool is often a displacement activity — the question we reach for because the real work is slow. So the bar is high: a tool earns a look only if it maps to a named bottleneck, and earns adoption only if the gain clears the switching cost and the tool runs with our governance grain, not against it.

§0 — Resolve the workstream lens

This skill is workstream-parametrized — one method, many lenses ("simple exterior, complex underneath"). Default workstream: capablemind.

If invoked as /tooling-scan <workstream> (e.g. studium-engine), FIRST read ~/_Dev/CapableMind-AI/docs/thinking/David/research/lens-<workstream>.md and use its values — the bet, the named bottlenecks (the step-2 gate), the governance-grain test, the register path, and workstream-specific sources — in place of the CapableMind specifics named below (the capablemind lens, also in lens-capablemind.md). If no card exists for the named workstream, say so and stop — never scan against a guessed lens.

The rubric (per candidate)

Assess in order — stop early if it fails an upstream test:

  1. What it is + maturity. Verify from source/docs, not marketing. Mark [SPECULATIVE].
  2. Which named bottleneck does it address? Map to our actual list — e.g.: hard-bug diagnosis · MemPalace reliability · inference reliability/cost (the teacher/slot tier) · procedural-knowledge capture · parallelism/throughput · spec↔code workflow. If it maps to none → PASS. "More capable" is not a bottleneck.
  3. Switching/integration cost vs marginal gain over the current stack (Claude Code + MemPalace + our skills + the governance loop). Be concrete about what we'd give up.
  4. Governance-grain fit. Does it run with our model or against it? Test against: the loop is load-bearing (does it push autonomy past propose?), data sovereignty (does our content leave the machine?), contamination (does it pressure toward pleasing outputs?). Hermes-as-runtime ran against (autonomous self-write); OpenRouter was configurable-compatible (default-no-log + ZDR + BYOK).
  5. Verdict: ADOPT-candidate · HARVEST (borrow the idea into our existing tools, don't switch — as we did turning Hermes's self-improvement into the governed skill-harvest step) · TRY (bounded experiment) · PASS · WATCH.
  6. Incumbent-bias flag. For runtime/substrate-class candidates (things that would replace Claude Code or MemPalace), the assessing agent is the incumbent — say so, discount the "don't switch" reflex, and prefer a bounded empirical trial by the steward over the agent's assertion.

Procedure

  1. Open the tooling register — ~/_Dev/CapableMind-AI/docs/thinking/David/research/tooling-register.md (create on first run; seed with the stack in play). Read current entries + last-scanned. Append in place.
  2. Re-check the stack in play for movement (incumbents + already-surfaced tools: Claude Code, MemPalace, OpenRouter, Hermes-as-runtime). Has anything changed that shifts a prior verdict? (Movement matters as much as novelty — cf. Hermes's self-improvement loop appearing between April and May.)
  3. Scan for new candidates mapped to a named bottleneck. Sources: GitHub (trending dev-tools, MCP servers, agent harnesses), HN, the landscape register's watch list, release notes of tools we already use.
  4. Run each through the rubric. Be ruthless at step 2 (named bottleneck) — most finds PASS there, and that's the point, not a failure of the scan.
  5. Record + recommend. Append verdicts to the register; spin a dated deep-dive only for an ADOPT-candidate/TRY. Surface a digest: the few that matter, each with bottleneck-addressed + verdict, and any bounded trial worth the steward running.

Output

A tight digest:

  • Movement in incumbents/known tools (or "none").
  • New candidates that cleared step 2 (named bottleneck), each one line: tool — bottleneck — verdict.
  • Recommended bounded trials (for runtime/substrate-class, steward-run, with the incumbent-bias caveat).
  • Register updated: what changed.

Constraints

  • Anti-shiny. If a scan only produces "X is more capable," it found nothing — re-aim at the named-bottleneck test. A clean "nothing worth adopting" is the common, valid result.
  • Verify from source. No claims from marketing or memory; read docs/code, cite, mark speculation.
  • Surface, don't adopt. Adoption is the steward's call — and for anything touching the product (inference infra, substrate) or the runtime, it's steward + Seb + jurist. The scan informs; it never switches tools.
  • Name the incumbent bias. The assessing agent benefits from the status quo; say so on runtime/substrate calls and lean on bounded trials over assertion.
  • Proportion. A couple-of-days pulse. Don't let tool-shopping become the work instead of the building.
  • Pool with Seb. Shared register; he scans too (the OpenRouter and Hermes finds were his).