Compare commits

..

2 Commits

Author SHA1 Message Date
Jesse Vincent c7b456dd0e tests: update SDD assertions and prompts to current skill behavior
The SDD skill tests still asserted on pre-rename skill text, so correct
model answers failed the suite:

- "Read at beginning" asserted "Step 1|beginning|start|Load Plan";
  the current skill has no numbered steps or Load Plan phase (setup
  covers it: "Read the plan once" during Setup). A live run today
  failed this assertion when the model correctly answered "during
  setup". Pattern now accepts setup/before-dispatch paraphrases while
  still requiring an at-the-start answer.
- "Provides text directly" asserted the removed provide-full-task-text
  behavior; SDD now routes task requirements through brief files
  (scripts/task-brief). The test now asks brief-file-vs-whole-plan and
  asserts the brief-based flow.
- The integration test's prompt and summary told the agent to
  "provide full task text to subagents (don't make them read files)",
  contradicting the skill it verifies; reworded to the task-brief flow.
  Also gave its direct `timeout 1800 claude -p` the same </dev/null
  stdin guard as run_claude.

Live-LLM tests; verified with bash -n on every touched file.

Part of #2130; defects documented in PR #2071 by @ericyen97903-lab.
2026-08-13 00:32:49 +00:00
Jesse Vincent 72ee5bbc5e tests: redirect stdin from /dev/null when spawning claude CLI
run_claude ran `timeout "$timeout" "${cmd[@]}"` with the suite's stdin
inherited by the spawned CLI. When the suite is run from a terminal (or
any open stdin), claude -p can block reading stdin and each test stalls
for its full timeout instead of completing.

Verified deterministically with a stub `claude` that reads stdin (cat):
with an open stdin pipe the old helper blocks until timeout kills it
(exit 124); with the redirect it exits immediately with output.

Part of #2130; defect documented in PR #2071 by @ericyen97903-lab.
2026-08-13 00:32:39 +00:00
5 changed files with 24 additions and 24 deletions
+1 -1
View File
@@ -101,7 +101,7 @@ Skills are not prose — they are code that shapes agent behavior. If you modify
## Eval harness ## Eval harness
Skill-behavior evals live in [superpowers-evals](https://github.com/prime-radiant-inc/superpowers-evals/), cloned into `evals/` — see `evals/README.md` for setup. Quorum (the harness CLI, one part of that eval lab) drives real coding-agent CLIs — Claude Code, Codex, Gemini, and others — through a Gauntlet QA agent and grades them against scenario acceptance criteria plus deterministic post-checks. Plugin-infrastructure tests still live at `tests/`. Skill-behavior evals live in [superpowers-evals](https://github.com/prime-radiant-inc/superpowers-evals/), cloned into `evals/` — see `evals/README.md` for setup. Drill (the harness) drives real tmux sessions of Claude Code / Codex / Gemini CLI and judges skill compliance with an LLM verifier. Plugin-infrastructure tests still live at `tests/`.
## Understand the Project Before Contributing ## Understand the Project Before Contributing
+9 -10
View File
@@ -14,23 +14,22 @@ Live in `tests/`. Currently:
- `tests/codex-plugin-sync/` — bash sync verification. - `tests/codex-plugin-sync/` — bash sync verification.
- `tests/kimi/` — bash/Python checks for Kimi plugin manifest wiring. - `tests/kimi/` — bash/Python checks for Kimi plugin manifest wiring.
- `tests/claude-code/test-helpers.sh`, `analyze-token-usage.py` — utilities used by remaining bash tests. - `tests/claude-code/test-helpers.sh`, `analyze-token-usage.py` — utilities used by remaining bash tests.
- `tests/claude-code/test-subagent-driven-development.sh` — agent-can-describe-SDD test (no quorum counterpart; tests description-recall, not behavior). - `tests/claude-code/test-subagent-driven-development.sh` — agent-can-describe-SDD test (no drill counterpart; tests description-recall, not behavior).
- `tests/claude-code/test-subagent-driven-development-integration.sh` — extended SDD integration with token analysis (quorum covers the YAGNI subset; bash adds commit-count, Claude Code task-tracking, and token telemetry assertions). - `tests/claude-code/test-subagent-driven-development-integration.sh` — extended SDD integration with token analysis (drill covers the YAGNI subset; bash adds commit-count, Claude Code task-tracking, and token telemetry assertions).
- `tests/claude-code/test-worktree-native-preference.sh` — RED-GREEN-REFACTOR validation for worktree skill (quorum covers the PRESSURE phase; bash also covers RED/GREEN baselines). - `tests/claude-code/test-worktree-native-preference.sh` — RED-GREEN-REFACTOR validation for worktree skill (drill covers the PRESSURE phase; bash also covers RED/GREEN baselines).
- `tests/explicit-skill-requests/` — Haiku-specific, multi-turn, and skill-name-prompted tests not covered by quorum. - `tests/explicit-skill-requests/` — Haiku-specific, multi-turn, and skill-name-prompted tests not covered by drill.
Run plugin tests via the relevant directory's `run-*.sh` or `npm test`. Run plugin tests via the relevant directory's `run-*.sh` or `npm test`.
## Skill behavior evals ## Skill behavior evals
Live in `evals/` (the [superpowers-evals](https://github.com/prime-radiant-inc/superpowers-evals/) eval lab, since renamed from Drill). Quorum is the harness CLI — one part of the system: it drives real coding-agent CLIs through a Gauntlet QA agent and grades them against each scenario's acceptance criteria plus deterministic post-checks. Scenarios live at `evals/scenarios/<name>/`. See `evals/README.md` for setup, the container runtime, and the safety model. Quick start (local break-glass run): Live in `evals/`. Drill is the harness; scenarios live at `evals/scenarios/*.yaml`. See `evals/README.md` for setup. Quick start:
```bash ```bash
cd evals cd evals
bun install uv sync --extra dev
export SUPERPOWERS_ROOT=/path/to/superpowers export ANTHROPIC_API_KEY=sk-...
bun run quorum run scenarios/triggering-test-driven-development --coding-agent claude uv run drill run triggering-test-driven-development -b claude
bun run quorum show <run-dir>
``` ```
Quorum scenarios are slow (3-30+ minutes each) and run real LLM sessions in permissive modes — read `evals/README.md`'s Live Eval Risk section first. Only the static gates (`bun run check`, `bun run quorum check`) are safe for public CI; the natural follow-up remains a tiered model (static gates on PR, live sweep nightly + on-demand). Drill scenarios are slow (3-30+ minutes each) and run real LLM sessions. They are not part of CI today; the natural follow-up is a tiered model (fast subset on PR, full sweep nightly + on-demand).
+3 -2
View File
@@ -15,8 +15,9 @@ run_claude() {
cmd+=(--allowed-tools="$allowed_tools") cmd+=(--allowed-tools="$allowed_tools")
fi fi
# Run Claude in headless mode with timeout # Run Claude in headless mode with timeout. Redirect stdin from
if timeout "$timeout" "${cmd[@]}" > "$output_file" 2>&1; then # /dev/null so the CLI can't block waiting for input and hang the suite.
if timeout "$timeout" "${cmd[@]}" > "$output_file" 2>&1 < /dev/null; then
cat "$output_file" cat "$output_file"
rm -f "$output_file" rm -f "$output_file"
return 0 return 0
@@ -23,7 +23,7 @@ echo "========================================"
echo "" echo ""
echo "This test executes a real plan using the skill and verifies:" echo "This test executes a real plan using the skill and verifies:"
echo " 1. Plan is read once (not per task)" echo " 1. Plan is read once (not per task)"
echo " 2. Full task text provided to subagents" echo " 2. Task requirements routed to subagents via brief files"
echo " 3. Subagents perform self-review" echo " 3. Subagents perform self-review"
echo " 4. Spec compliance review before code quality" echo " 4. Spec compliance review before code quality"
echo " 5. Review loops when issues found" echo " 5. Review loops when issues found"
@@ -136,7 +136,7 @@ I want you to execute the implementation plan at docs/superpowers/plans/implemen
IMPORTANT: Follow the skill exactly. I will be verifying that you: IMPORTANT: Follow the skill exactly. I will be verifying that you:
1. Read the plan once at the beginning 1. Read the plan once at the beginning
2. Provide full task text to subagents (don't make them read files) 2. Route each task's requirements to subagents via a task brief file (don't make them read the whole plan)
3. Ensure subagents do self-review before reporting 3. Ensure subagents do self-review before reporting
4. Run spec compliance review before code quality review 4. Run spec compliance review before code quality review
5. Use review loops when issues are found 5. Use review loops when issues are found
@@ -150,7 +150,7 @@ PROMPT="Execute the implementation plan at docs/superpowers/plans/implementation
IMPORTANT: Follow the skill exactly. I will be verifying that you: IMPORTANT: Follow the skill exactly. I will be verifying that you:
1. Read the plan once at the beginning 1. Read the plan once at the beginning
2. Provide full task text to subagents (don't make them read files) 2. Route each task's requirements to subagents via a task brief file (don't make them read the whole plan)
3. Ensure subagents do self-review before reporting 3. Ensure subagents do self-review before reporting
4. Run spec compliance review before code quality review 4. Run spec compliance review before code quality review
5. Use review loops when issues are found 5. Use review loops when issues are found
@@ -164,7 +164,7 @@ PLUGIN_DIR=$(cd "$SCRIPT_DIR/../.." && pwd)
# other concurrent claude sessions. # other concurrent claude sessions.
echo "Running Claude (plugin-dir: $PLUGIN_DIR, cwd: $TEST_PROJECT)..." echo "Running Claude (plugin-dir: $PLUGIN_DIR, cwd: $TEST_PROJECT)..."
echo "================================================================================" echo "================================================================================"
cd "$TEST_PROJECT" && timeout 1800 claude -p "$PROMPT" --plugin-dir "$PLUGIN_DIR" --allowed-tools=all --permission-mode bypassPermissions 2>&1 | tee "$OUTPUT_FILE" || { cd "$TEST_PROJECT" && timeout 1800 claude -p "$PROMPT" --plugin-dir "$PLUGIN_DIR" --allowed-tools=all --permission-mode bypassPermissions < /dev/null 2>&1 | tee "$OUTPUT_FILE" || {
echo "" echo ""
echo "================================================================================" echo "================================================================================"
echo "EXECUTION FAILED (exit code: $?)" echo "EXECUTION FAILED (exit code: $?)"
@@ -316,7 +316,7 @@ if [ $FAILED -eq 0 ]; then
echo "" echo ""
echo "The subagent-driven-development skill correctly:" echo "The subagent-driven-development skill correctly:"
echo " ✓ Reads plan once at start" echo " ✓ Reads plan once at start"
echo " ✓ Provides full task text to subagents" echo " ✓ Routes task requirements via brief files"
echo " ✓ Enforces self-review" echo " ✓ Enforces self-review"
echo " ✓ Runs spec compliance before code quality" echo " ✓ Runs spec compliance before code quality"
echo " ✓ Spec reviewer verifies independently" echo " ✓ Spec reviewer verifies independently"
@@ -4,7 +4,7 @@
# #
# No drill coverage: this test asks the agent to *describe* SDD (string- # No drill coverage: this test asks the agent to *describe* SDD (string-
# matches its verbal explanation against expected keywords like # matches its verbal explanation against expected keywords like
# "self-review", "skeptical", "worktree", "Step 1", "loop"). Drill scenarios # "self-review", "skeptical", "worktree", "setup", "loop"). Drill scenarios
# test behavior (real subagent dispatch, plan-following, review loops), # test behavior (real subagent dispatch, plan-following, review loops),
# not description-recall. Kept by design. # not description-recall. Kept by design.
set -euo pipefail set -euo pipefail
@@ -83,7 +83,7 @@ else
exit 1 exit 1
fi fi
if assert_contains "$output" "Step 1\|beginning\|start\|Load Plan" "Read at beginning"; then if assert_contains "$output" "beginning\|start\|setup\|before.*dispatch\|before.*task" "Read at beginning"; then
: # pass : # pass
else else
exit 1 exit 1
@@ -133,16 +133,16 @@ echo ""
echo "Test 7: Task context provision..." echo "Test 7: Task context provision..."
output=$(run_claude "In subagent-driven-development, how does the controller provide task information to the implementer subagent? Answer using exactly this structure: output=$(run_claude "In subagent-driven-development, how does the controller provide task information to the implementer subagent? Answer using exactly this structure:
Controller provides: <directly or by file> Controller provides: <brief file or whole plan file>
Implementer must read plan file: <yes or no>" "$CLAUDE_PROMPT_TIMEOUT") Implementer must read whole plan file: <yes or no>" "$CLAUDE_PROMPT_TIMEOUT")
if assert_contains "$output" "provide.*directly\|full.*text\|paste\|include.*prompt" "Provides text directly"; then if assert_contains "$output" "task-brief\|brief file\|brief.*path\|Controller provides:.*brief" "Provides task brief file"; then
: # pass : # pass
else else
exit 1 exit 1
fi fi
if assert_contains "$output" "Implementer must read plan file:.*no" "Doesn't make subagent read file"; then if assert_contains "$output" "Implementer must read whole plan file:.*no" "Doesn't make subagent read whole plan"; then
: # pass : # pass
else else
exit 1 exit 1