tests: update SDD assertions and prompts to current skill behavior

The SDD skill tests still asserted on pre-rename skill text, so correct
model answers failed the suite:

- "Read at beginning" asserted "Step 1|beginning|start|Load Plan";
  the current skill has no numbered steps or Load Plan phase (setup
  covers it: "Read the plan once" during Setup). A live run today
  failed this assertion when the model correctly answered "during
  setup". Pattern now accepts setup/before-dispatch paraphrases while
  still requiring an at-the-start answer.
- "Provides text directly" asserted the removed provide-full-task-text
  behavior; SDD now routes task requirements through brief files
  (scripts/task-brief). The test now asks brief-file-vs-whole-plan and
  asserts the brief-based flow.
- The integration test's prompt and summary told the agent to
  "provide full task text to subagents (don't make them read files)",
  contradicting the skill it verifies; reworded to the task-brief flow.
  Also gave its direct `timeout 1800 claude -p` the same </dev/null
  stdin guard as run_claude.

Live-LLM tests; verified with bash -n on every touched file.

Part of #2130; defects documented in PR #2071 by @ericyen97903-lab.
This commit is contained in:
Jesse Vincent
2026-08-13 00:32:49 +00:00
parent 72ee5bbc5e
commit c7b456dd0e
2 changed files with 11 additions and 11 deletions
@@ -23,7 +23,7 @@ echo "========================================"
echo ""
echo "This test executes a real plan using the skill and verifies:"
echo " 1. Plan is read once (not per task)"
echo " 2. Full task text provided to subagents"
echo " 2. Task requirements routed to subagents via brief files"
echo " 3. Subagents perform self-review"
echo " 4. Spec compliance review before code quality"
echo " 5. Review loops when issues found"
@@ -136,7 +136,7 @@ I want you to execute the implementation plan at docs/superpowers/plans/implemen
IMPORTANT: Follow the skill exactly. I will be verifying that you:
1. Read the plan once at the beginning
2. Provide full task text to subagents (don't make them read files)
2. Route each task's requirements to subagents via a task brief file (don't make them read the whole plan)
3. Ensure subagents do self-review before reporting
4. Run spec compliance review before code quality review
5. Use review loops when issues are found
@@ -150,7 +150,7 @@ PROMPT="Execute the implementation plan at docs/superpowers/plans/implementation
IMPORTANT: Follow the skill exactly. I will be verifying that you:
1. Read the plan once at the beginning
2. Provide full task text to subagents (don't make them read files)
2. Route each task's requirements to subagents via a task brief file (don't make them read the whole plan)
3. Ensure subagents do self-review before reporting
4. Run spec compliance review before code quality review
5. Use review loops when issues are found
@@ -164,7 +164,7 @@ PLUGIN_DIR=$(cd "$SCRIPT_DIR/../.." && pwd)
# other concurrent claude sessions.
echo "Running Claude (plugin-dir: $PLUGIN_DIR, cwd: $TEST_PROJECT)..."
echo "================================================================================"
cd "$TEST_PROJECT" && timeout 1800 claude -p "$PROMPT" --plugin-dir "$PLUGIN_DIR" --allowed-tools=all --permission-mode bypassPermissions 2>&1 | tee "$OUTPUT_FILE" || {
cd "$TEST_PROJECT" && timeout 1800 claude -p "$PROMPT" --plugin-dir "$PLUGIN_DIR" --allowed-tools=all --permission-mode bypassPermissions < /dev/null 2>&1 | tee "$OUTPUT_FILE" || {
echo ""
echo "================================================================================"
echo "EXECUTION FAILED (exit code: $?)"
@@ -316,7 +316,7 @@ if [ $FAILED -eq 0 ]; then
echo ""
echo "The subagent-driven-development skill correctly:"
echo " ✓ Reads plan once at start"
echo " ✓ Provides full task text to subagents"
echo " ✓ Routes task requirements via brief files"
echo " ✓ Enforces self-review"
echo " ✓ Runs spec compliance before code quality"
echo " ✓ Spec reviewer verifies independently"