diff --git a/agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/CODE_REVIEW-cloud-G06.md b/agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/CODE_REVIEW-cloud-G06.md new file mode 100644 index 00000000..a87d8105 --- /dev/null +++ b/agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/CODE_REVIEW-cloud-G06.md @@ -0,0 +1,149 @@ + + +# Code Review Reference - TEST + +> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.** +> The task is NOT complete until every implementation-owned section below is filled in. +> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving. +> Fill implementation-owned sections, then stop with active files in place and report ready for review. +> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt. +> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields. +> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state. +> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume. +> Follow the ownership table at the bottom of this file for which sections you own. + +## Overview + +date=2026-08-12 +task=m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest, plan=2, tag=TEST + +## Archive Evidence Snapshot + +- Prior plan/review before Epic self-review: `plan_local_G07_0.log`, `code_review_cloud_G07_0.log`. +- Refinement source pair: `plan_cloud_G09_1.log`, `code_review_cloud_G09_1.log`. +- Neither prior pair contains implementation evidence or an official verdict. The refinement preserves the original manifest, skill-contract, and regression scope while moving only the external readiness preflight into child `05+04_readiness_preflight`. +- Baseline fixture checksum/shape and the 178-test manifest/attempt/connectivity suite were green at `b197e5db70637f87017a024a847e3e53fdc72e8b` in the source plan evidence. + +## For the Review Agent + +> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section. + +Compare implementation of each item against source files. Run the applicable verification commands directly and record fresh output in `Verification Results`; implementation-owned output is handoff evidence, not a substitute for reviewer verification. If implementation is present, repair missing or stale verification output instead of failing solely for insufficient recorded evidence. When verification exposes a defect, collect the necessary data, determine the exact root cause, and select one concrete fix before generating the follow-up plan; never delegate investigation or remedy selection to the worker. +Review completion means the following steps are finished: + +1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals. +2. Archive `CODE_REVIEW-cloud-G06.md` → `code_review_cloud_G06_2.log` and `PLAN-local-G06.md` → `plan_local_G06_2.log`. +3. If PASS, write `complete.log` and move active task directory to `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/`. If WARN/FAIL, fully write the next filesystem state required by the code-review skill. +4. If PASS, preserve the first-line `milestone-task` metadata in `complete.log` and report it for runtime aggregation. Roadmap state evaluation belongs to `sync-milestone-workstate`. +5. Check applicable `Review-Only Checklist` items at the final `.log` location before reporting. + +--- + +## Implementation Item Completion + +| Item | Status | +|------|---------| +| TEST-1 Create one exact immutable readiness manifest | [ ] | +| TEST-3 Preserve the full benchmark baseline | [ ] | + +## Implementation Checklist + +- [ ] [TEST-1] Add the exact tracked bench-02 manifest and a static regression that locks fixture checksum/version, viewports, rubric, policies, C01-C09 bindings, seed, and seeded order; atomically update project-skill scoring language and its contract test to the manifest-selected rubric catalog. +- [ ] [TEST-3] Run final schema, focused regression, full benchmark suite, and whitespace verification while leaving live preflight and scored execution untouched. +- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output. + +## Review-Only Checklist + +> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent. +> Implementing agents must not modify or check this section. + +- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`. +- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match. +- [ ] Run applicable required verification and record fresh command/output; repair reviewer-reconstructable evidence gaps instead of forwarding them to another plan. +- [ ] For every Required/Suggested finding, record reviewer-collected `Evidence`, exact `Root Cause`, and one `Selected Fix` with affected files/symbols/tests and acceptance commands before creating a follow-up plan. +- [ ] Archive active `CODE_REVIEW-cloud-G06.md` to `code_review_cloud_G06_2.log`. +- [ ] Archive active `PLAN-local-G06.md` to `plan_local_G06_2.log`. +- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`. +- [ ] If PASS, write `complete.log` based on `agent-ops/skills/common/code-review/templates/complete-log-template.md` and leave no active `.md` files. +- [ ] If PASS, move active task directory `agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/` to `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/` and update this checklist at the final archive path. +- [ ] If PASS, preserve and report `milestone-task` metadata for runtime aggregation, without modifying roadmap or directly calling `update-roadmap`. +- [ ] If PASS for split work, keep parent `agent-task/m-iop-one-shot-agent-model-comparison/` because child `05+04_readiness_preflight` remains active until its own review completes. +- [ ] If WARN/FAIL, write the next filesystem state matching code-review verdict and do not write `complete.log`. + +## Deviations from Plan + +_Record any deviations from the plan and the rationale here._ + +## Key Design Decisions + +_Record key design decisions here._ + +## Reviewer Checkpoints + +- Confirm both predecessor completion files exist before reviewing implementation. +- Confirm the complete C01-C09 map, seed-derived order, fixture checksum, two images, policies, and new rubric version are exact. +- Confirm hybrid routes use route kind `execution_preset`, include `repair`, bind every cloud stage to its exact model with high effort, and bind work to `ornith-fast` with no invented effort. +- Confirm project-skill scoring uses the exact manifest-selected rubric, names the closed two-version set, and preserves predecessor 02's all-cell preflight contract. +- Confirm no public live preflight or scored run/score/report command was executed in this child. + +## Verification Results + +Record actual stdout/stderr under each command. If a command changes, document the replacement and reason in `Deviations from Plan`. + +### Dependency Verification + +```bash +python3 - <<'PY' +from pathlib import Path +root = Path("agent-task") +group = "m-iop-one-shot-agent-model-comparison" +for subtask in ("02_route_preflight_contract", "03+01_rubric_version_contract"): + active = root / group / subtask / "complete.log" + archived = sorted((root / "archive").glob(f"*/*/{group}/{subtask}/complete.log")) + candidates = [path for path in (active, *archived) if path.is_file()] + assert len(candidates) == 1, (subtask, candidates) + print(f"ok: predecessor complete {candidates[0]}") +PY +``` + +_Actual output:_ + +### TEST-1 Verification + +```bash +python3 scripts/agent_comparison_benchmark.py validate --manifest scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json +python3 -m unittest scripts.agent_benchmark.manifest_test.ManifestValidationTest.test_iop_one_shot_manifest_locks_benchmark_readiness +python3 -m unittest scripts.agent_benchmark.skill_contract_test +``` + +_Actual output:_ + +### TEST-3 and Final Verification + +```bash +for manifest in scripts/fixtures/agent-comparison-benchmark-manifest.example.json scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json; do python3 scripts/agent_comparison_benchmark.py validate --manifest "$manifest"; done +python3 -m unittest discover -s scripts/agent_benchmark -p '*_test.py' +git diff --check +``` + +_Actual output:_ + +--- + +> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?** +> If anything is blank, go back and fill it in before saving this file. +> Leave review-agent-only sections unchanged. + +## Section Ownership + +| Section | Owner | Note | +|---------|-------|------| +| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these (archive, complete.log, and task-directory archive move are review-agent only) | +| Archive Evidence Snapshot | Fixed at stub creation from plan when present | Implementing agent uses it as default prior-loop context; read only the specific archive files cited there when more detail is required | +| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only | +| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only | +| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section | +| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content | +| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan | +| Verification Results (section headings + commands) | Implementing agent, then review agent | Implementing agent records initial output; review agent reruns applicable commands and may fill, replace, or append fresh verified output before verdict. Implementing-agent command changes require a `Deviations from Plan` entry | +| Code Review Result | Review agent appends | Not included in stub | diff --git a/agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/PLAN-local-G06.md b/agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/PLAN-local-G06.md new file mode 100644 index 00000000..f2495de9 --- /dev/null +++ b/agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/PLAN-local-G06.md @@ -0,0 +1,244 @@ + + +# Plan - Locked C01-C09 Benchmark Manifest + +## For the Implementing Agent + +Do not start until the dependency command resolves exactly one active or archived `complete.log` for both predecessor indices 02 and 03. Filling the implementation-owned sections in `CODE_REVIEW-cloud-G06.md` is the mandatory last implementation step. Run every verification command, record actual notes and stdout/stderr, keep both active files in place, and report ready for review; only the code-review skill may finalize or archive this task. If blocked, record only the exact blocker, attempted commands/output, and resume condition in the implementation-owned evidence fields. Do not ask the user, call user-input tools, create control-plane stop files, classify the next state, archive logs, or write `complete.log`. + +## Background + +The repository already contains the reusable two-image vanilla landing-page fixture, but no immutable manifest combines it with the approved C01-C09 matrix, explicit seed, new rubric version, and exact policies. This packet creates and statically validates that tracked input before the dependent readiness-preflight packet contacts any external caller or runtime. + +## Archive Evidence Snapshot + +- Prior plan/review before Epic self-review: `plan_local_G07_0.log`, `code_review_cloud_G07_0.log`. +- Refinement source pair: `plan_cloud_G09_1.log`, `code_review_cloud_G09_1.log`. +- Neither prior pair contains implementation evidence or an official verdict. The refinement preserves the original manifest, skill-contract, and regression scope while moving only the external readiness preflight into child `05+04_readiness_preflight`. +- Baseline fixture checksum/shape and the 178-test manifest/attempt/connectivity suite were green at `b197e5db70637f87017a024a847e3e53fdc72e8b` in the source plan evidence. + +## Analysis + +### Files Read + +- `scripts/agent_benchmark/manifest.py` +- `scripts/agent_benchmark/manifest_test.py` +- `scripts/agent_benchmark/rubric.py` +- `scripts/agent_benchmark/rubric_test.py` +- `scripts/agent_benchmark/scoring.py` +- `scripts/agent_benchmark/scoring_test.py` +- `scripts/agent_benchmark/skill_contract_test.py` +- `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json` +- `scripts/fixtures/agent-comparison-benchmark-manifest.example.json` +- `scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json` +- `scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json` +- `scripts/fixtures/agent-comparison-benchmark/prompt.md` +- `scripts/fixtures/agent-comparison-benchmark/reference.txt` +- `scripts/fixtures/agent-comparison-benchmark/aurora-grid.svg` +- `scripts/fixtures/agent-comparison-benchmark/orbit-rings.svg` +- `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md` +- `agent-spec/testing/agent-comparison-benchmark.md` +- `agent-contract/outer/anthropic-compatible-api.md` +- `agent-contract/outer/openai-compatible-api.md` +- `agent-contract/inner/edge-config-runtime-refresh.md` +- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/iop-one-shot-agent-model-comparison.md` +- `agent-roadmap/sdd/knowledge-tool-optimization-extension/iop-one-shot-agent-model-comparison/SDD.md` + +### SDD Criteria + +- SDD status is `[승인됨]`, implementation lock is released, and no unresolved user review exists. +- First-line scope is narrowed to `milestone-task=fixture-lock,matrix-lock`. +- S01 requires identical prompt/assets/workspace/viewports/rubric checksum/version. S03 requires the immutable nine-cell manifest, repetitions 1, explicit seed, fresh/isolated policy, timeout, and bindings. +- The static manifest assertions and complete benchmark regression provide this child’s S01/S03 evidence. Live S02 evidence belongs to child 05. + +### Verification Context + +- The source plan fixed fixture checksum `sha256:7dc1be6ed4a9f2f873016b708b99d827b0249c74f2ac287e1fcf8deade8dcd98`, viewports 1920x1080 and 375x812, and seeded order C02, C05, C03, C06, C08, C09, C01, C07, C04. +- The approved hybrid route ids are `gemini-hybrid` and `gpt-hybrid`; stage order is selector, plan, work, review, repair. Cloud stages use `high`, while `ornith-fast` work omits effort. +- The manifest/schema route-kind contract accepts only `direct` and `execution_preset`; the locked hybrid cells must use `execution_preset` exactly. +- Fresh static validation and regression output is required; cached output is not accepted. No external runner is required for this child. + +### Test Coverage Gaps + +- No shipped manifest contains exactly C01-C09 with one explicit seed and the approved new rubric version. +- No static regression asserts fixture checksum, exact two-image paths, policies, full cell map, seeded order, and hybrid binding table together. +- The project benchmark skill still hardcodes the legacy worksheet, and its contract test does not require manifest-selected scoring with the closed rubric catalog. + +### Symbol References + +- No symbol is renamed or removed. +- The new fixture path is consumed by `load_manifest`, public CLI commands, `RunStore`, and the manifest regression test; no new API is introduced. + +### Split Judgment + +- This child owns the complete production/documentation change and every repository regression required for its PASS. +- Child `05+04_readiness_preflight` depends on this reviewed manifest and owns only the additional external closure verification. It does not modify the manifest or duplicate the static regression. +- Existing indices 01-04 remain fixed because their task directories contain logs; the new closure child uses the next free index 05. + +### Scope Rationale + +- Do not modify fixture prompt/reference/image bytes, benchmark runtime code, scoring/reporting logic, caller adapters, `../iop-s2`, credentials, or external registrations. +- Update only the project skill's scoring wording and matching contract assertions; preserve predecessor 02's all-cell preflight wording. +- Do not run public live preflight or any scored `run`, `resume`, `score`, or `report` command in this child. + +### Final Routing + +- `evaluation_mode=isolated-reassessment`; `finalizer=finalize-task-policy.sh`, `finalizer_mode=pair`. +- Build closures are all closed. Scores `2+0+2+1+1=G06`; base/final route `local-fit`, lane `local`, catalog `worker/local/G06`, filename `PLAN-local-G06.md`. +- Review closures are all closed. Scores `2+0+2+1+1=G06`; route `official-review`, lane `cloud`, catalog `review/cloud/G06`, filename `CODE_REVIEW-cloud-G06.md`. +- `large_indivisible_context=false`; positive loop risks: `boundary_contract`, `structured_interpretation`, `variant_product` (`count=3`); `review_rework_count=0`; `evidence_integrity_failure=false`; no capability gap. + +## Dependencies and Execution Order + +1. Resolve exactly one active-sibling or matching archived `complete.log` for `02_route_preflight_contract`. +2. Resolve exactly one active-sibling or matching archived `complete.log` for `03+01_rubric_version_contract`. +3. Create and statically validate the locked manifest and project-skill scoring contract. +4. On PASS, child `05+04_readiness_preflight` may run the external readiness gate. + +## Implementation Checklist + +- [ ] [TEST-1] Add the exact tracked bench-02 manifest and a static regression that locks fixture checksum/version, viewports, rubric, policies, C01-C09 bindings, seed, and seeded order; atomically update project-skill scoring language and its contract test to the manifest-selected rubric catalog. +- [ ] [TEST-3] Run final schema, focused regression, full benchmark suite, and whitespace verification while leaving live preflight and scored execution untouched. +- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output. + +### [TEST-1] Create one exact immutable readiness manifest + +**Problem** + +The generic fixture manifest has only three homogeneous presets (`scripts/fixtures/agent-comparison-benchmark-manifest.example.json:44`), while the direct preflight example has only five direct cells (`scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json:44`). Neither can satisfy S01 and S03, and the current route-kind contract accepts `direct` or `execution_preset` only (`scripts/agent_benchmark/manifest.py:33`). + +**Solution** + +Create `scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json` with these immutable common fields: + +```json +{ + "pipeline_version": "2", + "environment": "dev", + "testbed": "../iop-s2", + "execution_order_seed": "bench-02-c01-c09-v1", + "repetitions": 1, + "session_policy": "fresh", + "setup_cache_policy": "isolated", + "timeout": {"run_seconds": 300, "idle_seconds": 30, "quiet_seconds": 10, "cleanup_grace_seconds": 5}, + "rubric_version": "one-shot-agent-comparison-v1", + "output_root": "agent-test/runs/bench-02" +} +``` + +Reuse the existing evaluator, fixture paths/checksum, and viewports byte-for-byte. Use the exact matrix below; stage order is selector, plan, work, review, repair. Cloud selector/plan/review/repair effort is `high`; `ornith-fast` work omits effort. + +| ID | Cell id | Caller | Route kind/id | Request model/effort | Expected bindings | +|----|---------|--------|---------------|----------------------|-------------------| +| C01 | `c01-claude-sonnet-direct` | claude | direct / `claude-sonnet-5` | `claude-sonnet-5` / max | request `claude-sonnet-5`/max | +| C02 | `c02-claude-gemini-direct` | claude | direct / `gemini-3.6-flash` | `gemini-3.6-flash` / high | request `gemini-3.6-flash`/high | +| C03 | `c03-agy-gemini-direct` | agy | direct / `gemini-3.6-flash` | `gemini-3.6-flash` / high | request `gemini-3.6-flash`/high | +| C04 | `c04-claude-gpt-direct` | claude | direct / `gpt-5.6-luna` | `gpt-5.6-luna` / xhigh | request `gpt-5.6-luna`/xhigh | +| C05 | `c05-codex-gpt-direct` | codex | direct / `gpt-5.6-luna` | `gpt-5.6-luna` / xhigh | request `gpt-5.6-luna`/xhigh | +| C06 | `c06-claude-gemini-hybrid` | claude | execution_preset / `gemini-hybrid` | `gemini-hybrid` / high | selector=`gemini-3.6-flash`/high; plan=`gemini-3.6-flash`/high; work=`ornith-fast`/omitted; review=`gemini-3.6-flash`/high; repair=`gemini-3.6-flash`/high | +| C07 | `c07-agy-gemini-hybrid` | agy | execution_preset / `gemini-hybrid` | `gemini-hybrid` / high | same as C06 | +| C08 | `c08-claude-gpt-hybrid` | claude | execution_preset / `gpt-hybrid` | `gpt-hybrid` / xhigh | selector=`gpt-5.6-terra`/high; plan=`gpt-5.6-terra`/high; work=`ornith-fast`/omitted; review=`gpt-5.6-terra`/high; repair=`gpt-5.6-terra`/high | +| C09 | `c09-codex-gpt-hybrid` | codex | execution_preset / `gpt-hybrid` | `gpt-hybrid` / xhigh | same as C08 | + +Add `test_iop_one_shot_manifest_locks_benchmark_readiness` to `manifest_test.py`. Assert exact fixture version/checksum/assets, exactly two image workspace paths, viewports, rubric, timeout/session/cache/repetition/output root, seed, complete cell payloads, digest shape, and loaded seeded order: + +```text +c02-claude-gemini-direct +c05-codex-gpt-direct +c03-agy-gemini-direct +c06-claude-gemini-hybrid +c08-claude-gpt-hybrid +c09-codex-gpt-hybrid +c01-claude-sonnet-direct +c07-agy-gemini-hybrid +c04-claude-gpt-direct +``` + +In the same change, replace the project skill's fixed `landing-quality-v1` scoring sentence with a requirement to use the exact immutable manifest-selected rubric. Name the closed supported set (`landing-quality-v1`, `one-shot-agent-comparison-v1`) and forbid fallback or reinterpretation. Extend `skill_contract_test.py` so its base contract requires that wording and both version literals, while a mutation back to fixed-legacy wording fails. Preserve predecessor 02's all-cell preflight assertions unchanged. + +**Modified Files and Checklist** + +- [ ] `scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json`: add the exact locked readiness manifest. +- [ ] `scripts/agent_benchmark/manifest_test.py`: add the exact static regression and recompute fixture checksum from resolved inputs. +- [ ] `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md`: document manifest-selected exact scoring and the closed two-version rubric set. +- [ ] `scripts/agent_benchmark/skill_contract_test.py`: require that scoring contract without weakening the all-cell preflight contract. + +**Test Strategy** + +Write the named manifest test. Compare complete nested cell payloads, not just counts. Recalculate the fixture checksum with production helpers and assert the manifest digest is `sha256:` plus 64 lowercase hex characters. In the skill contract test, assert both exact supported rubric versions and manifest selection, then mutation-test the legacy-only regression. + +**Verification** + +```bash +python3 scripts/agent_comparison_benchmark.py validate --manifest scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json +python3 -m unittest scripts.agent_benchmark.manifest_test.ManifestValidationTest.test_iop_one_shot_manifest_locks_benchmark_readiness +python3 -m unittest scripts.agent_benchmark.skill_contract_test +``` + +Expected: the manifest validates, the exact lock test passes, and the skill contract suite accepts only manifest-selected rubric wording. + +### [TEST-3] Preserve the full benchmark baseline + +**Problem** + +The final file consumes the new order, rubric, and all-cell contracts through manifest loading (`scripts/agent_benchmark/manifest.py:651`), run storage/execution (`scripts/agent_benchmark/attempts.py:426`), scoring (`scripts/agent_benchmark/scoring.py:2006`), and the public CLI (`scripts/agent_comparison_benchmark.py:273`). It must not regress legacy fixtures or any other benchmark behavior. + +**Solution** + +Validate all four manifests, run the focused static test and the full benchmark discovery suite, and check the diff. Do not run live preflight in this child. + +**Modified Files and Checklist** + +- [ ] `CODE_REVIEW-cloud-G06.md`: record final static command output. + +**Test Strategy** + +No additional files beyond TEST-1. Existing suites are the regression oracle. + +**Verification** + +```bash +for manifest in scripts/fixtures/agent-comparison-benchmark-manifest.example.json scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json; do python3 scripts/agent_comparison_benchmark.py validate --manifest "$manifest"; done +python3 -m unittest discover -s scripts/agent_benchmark -p '*_test.py' +git diff --check +``` + +Expected: four manifests validate, the complete discovered benchmark suite reports `OK`, and whitespace check is silent. + +## Modified Files Summary + +| File | Item | +|------|------| +| `scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json` | TEST-1 | +| `scripts/agent_benchmark/manifest_test.py` | TEST-1 | +| `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md` | TEST-1 | +| `scripts/agent_benchmark/skill_contract_test.py` | TEST-1 | +| `agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/CODE_REVIEW-cloud-G06.md` | TEST-3 | + +## Final Verification + +Run from `/config/workspace/iop-s0` after both predecessors complete; fresh output is required and cached output is not acceptable. + +```bash +python3 - <<'PY' +from pathlib import Path +root = Path("agent-task") +group = "m-iop-one-shot-agent-model-comparison" +for subtask in ("02_route_preflight_contract", "03+01_rubric_version_contract"): + active = root / group / subtask / "complete.log" + archived = sorted((root / "archive").glob(f"*/*/{group}/{subtask}/complete.log")) + candidates = [path for path in (active, *archived) if path.is_file()] + assert len(candidates) == 1, (subtask, candidates) + print(f"ok: predecessor complete {candidates[0]}") +PY +python3 scripts/agent_comparison_benchmark.py validate --manifest scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json +python3 -m unittest scripts.agent_benchmark.manifest_test.ManifestValidationTest.test_iop_one_shot_manifest_locks_benchmark_readiness +python3 -m unittest scripts.agent_benchmark.skill_contract_test +for manifest in scripts/fixtures/agent-comparison-benchmark-manifest.example.json scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json; do python3 scripts/agent_comparison_benchmark.py validate --manifest "$manifest"; done +python3 -m unittest discover -s scripts/agent_benchmark -p '*_test.py' +git diff --check +``` + +Expected: dependency check and every command exit 0, four manifests validate, all suites report `OK`, and `git diff --check` is silent. No external preflight or scored benchmark command is run. + +After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`. diff --git a/agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/CODE_REVIEW-cloud-G09.md b/agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/code_review_cloud_G09_1.log similarity index 100% rename from agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/CODE_REVIEW-cloud-G09.md rename to agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/code_review_cloud_G09_1.log diff --git a/agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/PLAN-cloud-G09.md b/agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/plan_cloud_G09_1.log similarity index 100% rename from agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/PLAN-cloud-G09.md rename to agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/plan_cloud_G09_1.log diff --git a/agent-task/m-iop-one-shot-agent-model-comparison/05+04_readiness_preflight/CODE_REVIEW-cloud-G09.md b/agent-task/m-iop-one-shot-agent-model-comparison/05+04_readiness_preflight/CODE_REVIEW-cloud-G09.md new file mode 100644 index 00000000..bed31878 --- /dev/null +++ b/agent-task/m-iop-one-shot-agent-model-comparison/05+04_readiness_preflight/CODE_REVIEW-cloud-G09.md @@ -0,0 +1,220 @@ + + +# Code Review Reference - TEST + +> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.** +> The task is NOT complete until every implementation-owned section below is filled in. +> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving. +> Fill implementation-owned sections, then stop with active files in place and report ready for review. +> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt. +> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields. +> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state. +> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume. +> Follow the ownership table at the bottom of this file for which sections you own. + +## Overview + +date=2026-08-12 +task=m-iop-one-shot-agent-model-comparison/05+04_readiness_preflight, plan=0, tag=TEST + +## Archive Evidence Snapshot + +- Refinement source pair: `agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/plan_cloud_G09_1.log` and matching review log. +- The source pair contains no implementation evidence or official verdict. This child preserves only its TEST-2 external readiness scope and depends on the reviewed static manifest child. +- Source-plan read-only evidence found the current testbed blocked by host-incompatible/stale artifacts, missing benchmark environment references, and an incompatible agy version/help surface. Those facts are planning evidence, not a substitute for execution-time checks. + +## For the Review Agent + +> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section. + +Compare the execution evidence against the fixed gate. Rerun only when the plan's setup prerequisites are present and a fresh readiness observation is required; do not run scored benchmark commands. If the gate is blocked, verify that the exact safe blocker and resume condition are recorded without leaking runtime values. +Review completion means the following steps are finished: + +1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals. +2. Archive `CODE_REVIEW-cloud-G09.md` → `code_review_cloud_G09_0.log` and `PLAN-cloud-G09.md` → `plan_cloud_G09_0.log`. +3. If PASS, write `complete.log` and move active task directory to `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/05+04_readiness_preflight/`. If WARN/FAIL, fully write the next filesystem state required by the code-review skill. +4. If PASS, preserve the first-line `milestone-task` metadata in `complete.log` and report it for runtime aggregation. Roadmap state evaluation belongs to `sync-milestone-workstate`. +5. Check applicable `Review-Only Checklist` items at the final `.log` location before reporting. + +--- + +## Implementation Item Completion + +| Item | Status | +|------|---------| +| TEST-2 Produce one redacted nine-cell readiness record | [ ] | + +## Implementation Checklist + +- [ ] [TEST-2] Run the secret-safe caller/testbed/environment gate and public all-cell preflight; require nine ready results or record the exact blocker and resume condition without substitution. +- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output. + +## Review-Only Checklist + +> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent. +> Implementing agents must not modify or check this section. + +- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`. +- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match. +- [ ] Run applicable required verification and record fresh command/output; repair reviewer-reconstructable evidence gaps instead of forwarding them to another plan. +- [ ] For every Required/Suggested finding, record reviewer-collected `Evidence`, exact `Root Cause`, and one `Selected Fix` with affected files/symbols/tests and acceptance commands before creating a follow-up plan. +- [ ] Archive active `CODE_REVIEW-cloud-G09.md` to `code_review_cloud_G09_0.log`. +- [ ] Archive active `PLAN-cloud-G09.md` to `plan_cloud_G09_0.log`. +- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`. +- [ ] If PASS, write `complete.log` based on `agent-ops/skills/common/code-review/templates/complete-log-template.md` and leave no active `.md` files. +- [ ] If PASS, move active task directory `agent-task/m-iop-one-shot-agent-model-comparison/05+04_readiness_preflight/` to `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/05+04_readiness_preflight/` and update this checklist at the final archive path. +- [ ] If PASS, preserve and report `milestone-task` metadata for runtime aggregation, without modifying roadmap or directly calling `update-roadmap`. +- [ ] If PASS for split work, remove empty active parent `agent-task/m-iop-one-shot-agent-model-comparison/` or verify it was kept due to remaining siblings/files. +- [ ] If WARN/FAIL, write the next filesystem state matching code-review verdict and do not write `complete.log`. + +## Deviations from Plan + +_Record any deviations from the plan and the rationale here._ + +## Key Design Decisions + +_Record key design decisions here._ + +## Reviewer Checkpoints + +- Confirm predecessor 04 has exactly one active or archived `complete.log` before accepting any readiness evidence. +- Confirm the external gate captures the clean independent `../iop-s2` dev HEAD and accepts only the production adapter's exact known agy version/help contract. +- Confirm both iop-s2 artifacts are current-HEAD AArch64 ELF executables before accepting endpoint readiness; `test -x` alone is insufficient. +- Confirm live readiness came only from public preflight and durable evidence contains no endpoint, secret, or raw config value. +- Confirm every locked manifest cell is represented exactly once and all nine are ready. +- Confirm no tracked production/config/test file and no scored run/score/report surface was modified or executed in this child. + +## Verification Results + +Record actual safe stdout/stderr below. If a command changes, document the replacement and reason in `Deviations from Plan`. The full external block is fixed and must be recorded with actual safe output or the exact blocker and resume condition. + +### TEST-2 External Preflight + +```bash +set -euo pipefail +blocked() { + printf 'blocked: %s\n' "$1" >&2 + exit 69 +} +python3 - <<'PY' || blocked "predecessor 04 must have exactly one active or archived complete.log" +from pathlib import Path +root = Path("agent-task") +group = "m-iop-one-shot-agent-model-comparison" +subtask = "04+02,03_locked_benchmark_manifest" +active = root / group / subtask / "complete.log" +archived = sorted((root / "archive").glob(f"*/*/{group}/{subtask}/complete.log")) +candidates = [path for path in (active, *archived) if path.is_file()] +assert len(candidates) == 1, (subtask, candidates) +print(f"ok: predecessor complete {candidates[0]}") +PY +test "$(git -C ../iop-s2 branch --show-current)" = "dev" || blocked "../iop-s2 must be on branch dev" +test -z "$(git -C ../iop-s2 status --short)" || blocked "../iop-s2 must be clean" +testbed_head="$(git -C ../iop-s2 rev-parse HEAD)" || blocked "../iop-s2 HEAD must resolve" +test -n "$testbed_head" || blocked "../iop-s2 HEAD must be non-empty" +printf 'ok: testbed branch=dev head=%s clean=true\n' "$testbed_head" +command -v readelf >/dev/null || blocked "readelf must be installed" +test -x ../iop-s2/build/bin/iop-edge || blocked "current iop-edge artifact must be executable" +test -x ../iop-s2/build/dev/iop-node || blocked "current iop-node artifact must be executable" +test -f ../iop-s2/configs/edge.yaml || blocked "../iop-s2/configs/edge.yaml must exist" +commit_epoch="$(git -C ../iop-s2 show -s --format=%ct HEAD)" || blocked "testbed commit time must resolve" +for binary in ../iop-s2/build/bin/iop-edge ../iop-s2/build/dev/iop-node; do + test "$(stat -c %Y "$binary")" -ge "$commit_epoch" || blocked "$binary must be built from the current testbed HEAD" + readelf -h "$binary" | rg 'Machine:\s+AArch64' >/dev/null || blocked "$binary must be a Linux AArch64 ELF artifact" +done +../iop-s2/build/bin/iop-edge --help >/dev/null || blocked "iop-edge help must execute on this host" +test -n "$(../iop-s2/build/dev/iop-node version)" || blocked "iop-node version must execute on this host" +printf 'ok: current-HEAD Linux AArch64 Edge/Node artifacts are executable\n' +claude --version || blocked "claude caller must be installed" +agy --version || blocked "agy caller must be installed" +codex --version || blocked "codex caller must be installed" +python3 - <<'PY' || blocked "agy must match the adapter-known version and transport help contract" +import subprocess +from scripts.agent_benchmark.agy_iop import AGY_KNOWN_VERSION, inspect_agy_iop_capability +version_run = subprocess.run(["agy", "--version"], check=True, capture_output=True, text=True) +help_run = subprocess.run(["agy", "--help"], check=True, capture_output=True, text=True) +version = (version_run.stdout + version_run.stderr).strip() +help_text = help_run.stdout + help_run.stderr +capability = inspect_agy_iop_capability(version, help_text) +assert capability.version == AGY_KNOWN_VERSION, (capability.version, AGY_KNOWN_VERSION) +assert capability.iop_transport_supported, capability +print(f"ok: adapter-known agy {AGY_KNOWN_VERSION} transport is documented") +PY +python3 - <<'PY' || blocked "benchmark endpoint, secret, and config-observation references must be present" +import json, os, re +name = re.compile(r"^[A-Za-z_][A-Za-z0-9_]*$") +for caller in ("CLAUDE", "AGY", "CODEX"): + assert os.environ.get(f"IOP_BENCH_{caller}_BASE_URL") + ref = os.environ.get(f"IOP_BENCH_{caller}_SECRET_ENV", "") + assert name.fullmatch(ref) and os.environ.get(ref) +config_ref = os.environ.get("IOP_BENCH_CONFIG_OBSERVATION_ENV", "") +assert name.fullmatch(config_ref) and os.environ.get(config_ref) +value = json.loads(os.environ[config_ref]) +assert value.get("schema_version") == "1" and isinstance(value.get("routes"), list) +print("ok: benchmark environment references present") +PY +set +e +preflight_output="$(python3 scripts/agent_comparison_benchmark.py preflight --manifest scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json 2>&1)" +preflight_status=$? +set -e +printf '%s\n' "$preflight_output" +case "$preflight_status" in + 0) case "$preflight_output" in *"status=ready ready=9 registration_required=0 implementation_gap=0"*) ;; *) exit 1 ;; esac ;; + 69) case "$preflight_output" in *"error: preflight blocked "*) ;; *) exit 1 ;; esac ;; + *) exit "$preflight_status" ;; +esac +run_id="${preflight_output#*run_id=}"; run_id="${run_id%% *}" +test -n "$run_id" +python3 - "$run_id" "$preflight_status" <<'PY' +import json, os, sys +from pathlib import Path +from scripts.agent_benchmark.manifest import load_manifest +manifest_path = Path("scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json") +manifest = load_manifest(manifest_path, repo_root=Path.cwd()) +run_root = Path(manifest.output_root) / sys.argv[1] +record = json.loads((run_root / "preflight/preflight-000001.json").read_text(encoding="ascii")) +assert [item["cell"]["id"] for item in record["results"]] == [cell.id for cell in manifest.matrix] +assert len(record["results"]) == 9 +if sys.argv[2] == "0": + assert record["status"] == "ready" and all(item["status"] == "ready" for item in record["results"]) +else: + assert record["status"] in {"registration_required", "implementation_gap"} + assert any(item["status"] != "ready" for item in record["results"]) +sensitive = [] +for caller in ("CLAUDE", "AGY", "CODEX"): + sensitive.append(os.environ[f"IOP_BENCH_{caller}_BASE_URL"].encode()) + sensitive.append(os.environ[os.environ[f"IOP_BENCH_{caller}_SECRET_ENV"]].encode()) +config_ref = os.environ["IOP_BENCH_CONFIG_OBSERVATION_ENV"] +sensitive.append(os.environ[config_ref].encode()) +durable = b"".join(path.read_bytes() for path in run_root.rglob("*") if path.is_file()) +assert all(value and value not in durable for value in sensitive) +print("ok: nine preflight results are manifest-bound and runtime values are absent") +for item in record["results"]: + issues = ",".join(f'{issue["code"]}:{issue["resume_code"]}' for issue in item["issues"]) or "none" + print(f'cell={item["cell"]["id"]} status={item["status"]} issues={issues}') +PY +if [ "$preflight_status" -eq 69 ]; then + exit 69 +fi +``` + +_Actual safe output or exact blocker and resume condition:_ + +--- + +> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?** +> If anything is blank, go back and fill it in before saving this file. +> Leave review-agent-only sections unchanged. + +## Section Ownership + +| Section | Owner | Note | +|---------|-------|------| +| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these (archive, complete.log, and task-directory archive move are review-agent only) | +| Archive Evidence Snapshot | Fixed at stub creation from plan when present | Implementing agent uses it as default prior-loop context; read only the specific archive files cited there when more detail is required | +| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only | +| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only | +| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section | +| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content | +| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan | +| Verification Results (section headings + commands) | Implementing agent, then review agent | Implementing agent records initial output; review agent reruns applicable commands and may fill, replace, or append fresh verified output before verdict. Implementing-agent command changes require a `Deviations from Plan` entry | +| Code Review Result | Review agent appends | Not included in stub | diff --git a/agent-task/m-iop-one-shot-agent-model-comparison/05+04_readiness_preflight/PLAN-cloud-G09.md b/agent-task/m-iop-one-shot-agent-model-comparison/05+04_readiness_preflight/PLAN-cloud-G09.md new file mode 100644 index 00000000..d975d4f4 --- /dev/null +++ b/agent-task/m-iop-one-shot-agent-model-comparison/05+04_readiness_preflight/PLAN-cloud-G09.md @@ -0,0 +1,239 @@ + + +# Plan - Redacted Nine-Cell Readiness Preflight + +## For the Implementing Agent + +Do not start until the dependency command resolves exactly one active or archived `04+02,03_locked_benchmark_manifest/complete.log`. Filling the implementation-owned sections in `CODE_REVIEW-cloud-G09.md` is the mandatory last implementation step. Run the fixed external gate exactly once after its setup checks pass, record safe output or the exact blocker and resume condition, keep both active files in place, and report ready for review; only the code-review skill may finalize or archive this task. Do not ask the user, call user-input tools, create control-plane stop files, classify the next state, archive logs, or write `complete.log`. + +## Background + +The reviewed static child creates the immutable C01-C09 input. This packet performs the remaining execution-day readiness closure through the public benchmark preflight without substituting missing callers, credentials, models, presets, runtime artifacts, or effort settings. + +## Archive Evidence Snapshot + +- Refinement source pair: `agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/plan_cloud_G09_1.log` and matching review log. +- The source pair contains no implementation evidence or official verdict. This child preserves only its TEST-2 external readiness scope and depends on the reviewed static manifest child. +- Source-plan read-only evidence found the current testbed blocked by host-incompatible/stale artifacts, missing benchmark environment references, and an incompatible agy version/help surface. Those facts are planning evidence, not a substitute for execution-time checks. + +## Analysis + +### Files Read + +- `scripts/agent_comparison_benchmark.py` +- `scripts/agent_benchmark/manifest.py` +- `scripts/agent_benchmark/attempts.py` +- `scripts/agent_benchmark/connectivity.py` +- `scripts/agent_benchmark/live_iop.py` +- `scripts/agent_benchmark/claude_iop.py` +- `scripts/agent_benchmark/agy_iop.py` +- `scripts/agent_benchmark/codex_iop.py` +- `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json` +- `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md` +- `agent-spec/testing/agent-comparison-benchmark.md` +- `agent-contract/outer/anthropic-compatible-api.md` +- `agent-contract/outer/openai-compatible-api.md` +- `agent-contract/inner/edge-config-runtime-refresh.md` +- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/iop-one-shot-agent-model-comparison.md` +- `agent-roadmap/sdd/knowledge-tool-optimization-extension/iop-one-shot-agent-model-comparison/SDD.md` + +### SDD Criteria + +- SDD status is `[승인됨]`, implementation lock is released, and no unresolved user review exists. +- First-line scope is `milestone-task=route-readiness`. +- S02 requires execution-day auth, model/preset, effort, stream/finish/idle checks for all C01-C09 or an exact blocker. +- The required evidence is one redacted C01-C09 preflight matrix with auth/route/effort/terminal status, or the exact safe blocker and resume condition. + +### Verification Context + +- Runner/workdir: local Linux `aarch64`, `/config/workspace/iop-s0`. +- Testbed: `/config/workspace/iop-s2`, branch `dev`; execution must capture its exact current clean HEAD as independent read-only provenance. +- Source-plan evidence found `../iop-s2/build/bin/iop-edge` to be Mach-O and both runtime artifacts older than the observed testbed HEAD. It also found agy `1.1.12` while the reviewed adapter accepted `1.1.11`, and all benchmark environment references were missing. +- Exact resume condition remains: rebuild both iop-s2 artifacts from its exact clean `dev` HEAD for Linux AArch64, provide all benchmark env references and their non-empty referenced secret/config values, register all direct/hybrid routes with exact bindings, run the rebuilt endpoints, and install the adapter-known agy version whose help satisfies every documented transport token. +- No raw endpoint, secret, or config-observation value may be written to the repository or review file. + +### Test Coverage Gaps + +- Static fixtures cannot prove execution-day runtime identity, artifact compatibility, credentials, route registration, caller transport, or finish/idle readiness. +- No durable nine-cell all-ready record exists for the locked manifest. + +### Symbol References + +- No production symbol or tracked config is changed. +- The public `preflight` command and durable record are the only acceptance boundary; never invoke a caller/provider directly. + +### Split Judgment + +- This is the allowed closure-verification child from the source pair. It has no production write set and consumes the exact reviewed manifest from child 04. +- Static manifest construction and all repository regressions remain in predecessor 04; this child does not duplicate them. +- No child created in this refinement pass is split again. + +### Scope Rationale + +- Do not edit `../iop-s2`, install tools, change credentials/registrations, or substitute models, presets, callers, or effort. +- Do not modify the locked manifest, fixture bytes, benchmark runtime, adapters, project skill, tests, scoring, or reporting. +- Do not run scored `run`, `resume`, `score`, or `report`; this child ends at readiness preflight evidence. + +### Final Routing + +- `evaluation_mode=isolated-reassessment`; `finalizer=finalize-task-policy.sh`, `finalizer_mode=pair`. +- Build closures are all closed. Scores `2+1+2+2+2=G09`; base/final route `grade-boundary`, lane `cloud`, catalog `worker/cloud/G09`, filename `PLAN-cloud-G09.md`. +- Review closures are all closed. Scores `2+1+2+2+2=G09`; route `official-review`, lane `cloud`, catalog `review/cloud/G09`, filename `CODE_REVIEW-cloud-G09.md`. +- `large_indivisible_context=false`; positive loop risks: `temporal_state`, `boundary_contract`, `structured_interpretation`, `variant_product` (`count=4`; grade remains the route basis); `review_rework_count=0`; `evidence_integrity_failure=false`; no capability gap. + +## Dependencies and Execution Order + +1. Resolve exactly one active-sibling or matching archived `complete.log` for `04+02,03_locked_benchmark_manifest`. +2. Run dependency, clean-testbed, artifact, caller-version/help, and secret-reference checks in the fixed order. +3. Only after every setup check passes, invoke the public benchmark `preflight` once and inspect its durable nine-result record. +4. Stop on the first exact blocker or after recording the all-ready redacted evidence. Do not proceed to scored execution. + +## Implementation Checklist + +- [ ] [TEST-2] Run the secret-safe caller/testbed/environment gate and public all-cell preflight; require nine ready results or record the exact blocker and resume condition without substitution. +- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output. + +### [TEST-2] Produce one redacted nine-cell readiness record + +**Problem** + +Static validation cannot establish live readiness. The public CLI returns closed preflight blockers with exit 69 (`scripts/agent_comparison_benchmark.py:208`), while canonical per-cell issue/resume evidence and preflight collection are owned by `scripts/agent_benchmark/connectivity.py:406` and `scripts/agent_benchmark/attempts.py:386`. Planning-time external preflight was blocked by stale/host-incompatible artifacts, missing environment references, and an incompatible agy capability surface. + +**Solution** + +First run dependency, clean-testbed, artifact, caller-version/help, and secret-reference checks. The environment check dereferences names without printing names or values. Then invoke only the public benchmark `preflight` command once. Parse its run id, require `status=ready ready=9 registration_required=0 implementation_gap=0`, and inspect the durable result list against the loaded manifest. Search durable bytes for exact runtime base URLs, dereferenced secrets, and raw config observation JSON without printing them. + +If any setup check fails or preflight returns 69, record the exact safe output and resume condition in `CODE_REVIEW-cloud-G09.md` and stop. The shell must capture preflight stderr and status without allowing `set -e` to exit before the blocker can be recorded. Do not mutate external state. + +**Modified Files and Checklist** + +- [ ] `CODE_REVIEW-cloud-G09.md`: record safe preflight output, run id, nine-result inspection, or exact blocker/resume evidence. + +**Test Strategy** + +No unit test is added. The public CLI and durable record are the acceptance oracle. Never call a caller/provider directly. + +**Verification** + +```bash +set -euo pipefail +blocked() { + printf 'blocked: %s\n' "$1" >&2 + exit 69 +} +python3 - <<'PY' || blocked "predecessor 04 must have exactly one active or archived complete.log" +from pathlib import Path +root = Path("agent-task") +group = "m-iop-one-shot-agent-model-comparison" +subtask = "04+02,03_locked_benchmark_manifest" +active = root / group / subtask / "complete.log" +archived = sorted((root / "archive").glob(f"*/*/{group}/{subtask}/complete.log")) +candidates = [path for path in (active, *archived) if path.is_file()] +assert len(candidates) == 1, (subtask, candidates) +print(f"ok: predecessor complete {candidates[0]}") +PY +test "$(git -C ../iop-s2 branch --show-current)" = "dev" || blocked "../iop-s2 must be on branch dev" +test -z "$(git -C ../iop-s2 status --short)" || blocked "../iop-s2 must be clean" +testbed_head="$(git -C ../iop-s2 rev-parse HEAD)" || blocked "../iop-s2 HEAD must resolve" +test -n "$testbed_head" || blocked "../iop-s2 HEAD must be non-empty" +printf 'ok: testbed branch=dev head=%s clean=true\n' "$testbed_head" +command -v readelf >/dev/null || blocked "readelf must be installed" +test -x ../iop-s2/build/bin/iop-edge || blocked "current iop-edge artifact must be executable" +test -x ../iop-s2/build/dev/iop-node || blocked "current iop-node artifact must be executable" +test -f ../iop-s2/configs/edge.yaml || blocked "../iop-s2/configs/edge.yaml must exist" +commit_epoch="$(git -C ../iop-s2 show -s --format=%ct HEAD)" || blocked "testbed commit time must resolve" +for binary in ../iop-s2/build/bin/iop-edge ../iop-s2/build/dev/iop-node; do + test "$(stat -c %Y "$binary")" -ge "$commit_epoch" || blocked "$binary must be built from the current testbed HEAD" + readelf -h "$binary" | rg 'Machine:\s+AArch64' >/dev/null || blocked "$binary must be a Linux AArch64 ELF artifact" +done +../iop-s2/build/bin/iop-edge --help >/dev/null || blocked "iop-edge help must execute on this host" +test -n "$(../iop-s2/build/dev/iop-node version)" || blocked "iop-node version must execute on this host" +printf 'ok: current-HEAD Linux AArch64 Edge/Node artifacts are executable\n' +claude --version || blocked "claude caller must be installed" +agy --version || blocked "agy caller must be installed" +codex --version || blocked "codex caller must be installed" +python3 - <<'PY' || blocked "agy must match the adapter-known version and transport help contract" +import subprocess +from scripts.agent_benchmark.agy_iop import AGY_KNOWN_VERSION, inspect_agy_iop_capability +version_run = subprocess.run(["agy", "--version"], check=True, capture_output=True, text=True) +help_run = subprocess.run(["agy", "--help"], check=True, capture_output=True, text=True) +version = (version_run.stdout + version_run.stderr).strip() +help_text = help_run.stdout + help_run.stderr +capability = inspect_agy_iop_capability(version, help_text) +assert capability.version == AGY_KNOWN_VERSION, (capability.version, AGY_KNOWN_VERSION) +assert capability.iop_transport_supported, capability +print(f"ok: adapter-known agy {AGY_KNOWN_VERSION} transport is documented") +PY +python3 - <<'PY' || blocked "benchmark endpoint, secret, and config-observation references must be present" +import json, os, re +name = re.compile(r"^[A-Za-z_][A-Za-z0-9_]*$") +for caller in ("CLAUDE", "AGY", "CODEX"): + assert os.environ.get(f"IOP_BENCH_{caller}_BASE_URL") + ref = os.environ.get(f"IOP_BENCH_{caller}_SECRET_ENV", "") + assert name.fullmatch(ref) and os.environ.get(ref) +config_ref = os.environ.get("IOP_BENCH_CONFIG_OBSERVATION_ENV", "") +assert name.fullmatch(config_ref) and os.environ.get(config_ref) +value = json.loads(os.environ[config_ref]) +assert value.get("schema_version") == "1" and isinstance(value.get("routes"), list) +print("ok: benchmark environment references present") +PY +set +e +preflight_output="$(python3 scripts/agent_comparison_benchmark.py preflight --manifest scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json 2>&1)" +preflight_status=$? +set -e +printf '%s\n' "$preflight_output" +case "$preflight_status" in + 0) case "$preflight_output" in *"status=ready ready=9 registration_required=0 implementation_gap=0"*) ;; *) exit 1 ;; esac ;; + 69) case "$preflight_output" in *"error: preflight blocked "*) ;; *) exit 1 ;; esac ;; + *) exit "$preflight_status" ;; +esac +run_id="${preflight_output#*run_id=}"; run_id="${run_id%% *}" +test -n "$run_id" +python3 - "$run_id" "$preflight_status" <<'PY' +import json, os, sys +from pathlib import Path +from scripts.agent_benchmark.manifest import load_manifest +manifest_path = Path("scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json") +manifest = load_manifest(manifest_path, repo_root=Path.cwd()) +run_root = Path(manifest.output_root) / sys.argv[1] +record = json.loads((run_root / "preflight/preflight-000001.json").read_text(encoding="ascii")) +assert [item["cell"]["id"] for item in record["results"]] == [cell.id for cell in manifest.matrix] +assert len(record["results"]) == 9 +if sys.argv[2] == "0": + assert record["status"] == "ready" and all(item["status"] == "ready" for item in record["results"]) +else: + assert record["status"] in {"registration_required", "implementation_gap"} + assert any(item["status"] != "ready" for item in record["results"]) +sensitive = [] +for caller in ("CLAUDE", "AGY", "CODEX"): + sensitive.append(os.environ[f"IOP_BENCH_{caller}_BASE_URL"].encode()) + sensitive.append(os.environ[os.environ[f"IOP_BENCH_{caller}_SECRET_ENV"]].encode()) +config_ref = os.environ["IOP_BENCH_CONFIG_OBSERVATION_ENV"] +sensitive.append(os.environ[config_ref].encode()) +durable = b"".join(path.read_bytes() for path in run_root.rglob("*") if path.is_file()) +assert all(value and value not in durable for value in sensitive) +print("ok: nine preflight results are manifest-bound and runtime values are absent") +for item in record["results"]: + issues = ",".join(f'{issue["code"]}:{issue["resume_code"]}' for issue in item["issues"]) or "none" + print(f'cell={item["cell"]["id"]} status={item["status"]} issues={issues}') +PY +if [ "$preflight_status" -eq 69 ]; then + exit 69 +fi +``` + +Expected: every setup command exits 0. A ready preflight prints one ready summary, verifies nine ready results, and exits 0. A closed preflight blocker is captured despite `set -e`, verifies and prints only per-cell status plus canonical issue/resume codes from the durable record, and exits 69 so implementation records that exact safe blocker and the resume condition before stopping. + +## Modified Files Summary + +| File | Item | +|------|------| +| `agent-task/m-iop-one-shot-agent-model-comparison/05+04_readiness_preflight/CODE_REVIEW-cloud-G09.md` | TEST-2 | + +## Final Verification + +Run from `/config/workspace/iop-s0` after predecessor 04 completes. Run the TEST-2 verification block exactly once after its setup checks pass. No other validation or scored benchmark command belongs to this child. + +Expected: dependency and setup checks pass, one public preflight reports nine ready results, and durable evidence contains none of the exact runtime values. Otherwise the review stub records the exact safe blocker and resume condition. + +After completing the readiness check, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.