feat(epic): benchmark-readiness 작업을 준비한다
This commit is contained in:
parent
b197e5db70
commit
4c16f95c6d
16 changed files with 3650 additions and 0 deletions
|
|
@ -0,0 +1,138 @@
|
|||
<!-- task=m-iop-one-shot-agent-model-comparison/01_execution_order_contract plan=1 tag=API milestone-task=matrix-lock -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-12
|
||||
task=m-iop-one-shot-agent-model-comparison/01_execution_order_contract, plan=1, tag=API
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior plan: `agent-task/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/plan_local_G08_0.log`.
|
||||
- Prior review: `agent-task/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/code_review_cloud_G08_0.log`.
|
||||
- The prior pair contains no implementation evidence or official verdict. Epic self-review found one stale dependency statement: it named a nonexistent integration child and incorrectly implied that child depended directly on both indices 01 and 02.
|
||||
- Replan keeps the implementation contract unchanged and corrects the DAG to `01 -> 03+01`, `02` independent, and `(02,03) -> 04+02,03`. Baseline manifest validation and the 178-test manifest/attempt/connectivity suite were green at `b197e5db70637f87017a024a847e3e53fdc72e8b`.
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files. Run the applicable verification commands directly and record fresh output in `Verification Results`; implementation-owned output is handoff evidence, not a substitute for reviewer verification. If implementation is present, repair missing or stale verification output instead of failing solely for insufficient recorded evidence. When verification exposes a defect, collect the necessary data, determine the exact root cause, and select one concrete fix before generating the follow-up plan; never delegate investigation or remedy selection to the worker.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals.
|
||||
2. Archive `CODE_REVIEW-cloud-G08.md` → `code_review_cloud_G08_1.log` and `PLAN-local-G08.md` → `plan_local_G08_1.log`.
|
||||
3. If PASS, write `complete.log` and move active task directory to `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/`. If WARN/FAIL, fully write the next filesystem state required by the code-review skill.
|
||||
4. If PASS, preserve the first-line `milestone-task` metadata in `complete.log` and report it for runtime aggregation. Roadmap state evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check applicable `Review-Only Checklist` items at the final `.log` location before reporting.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|---------|
|
||||
| API-1 Add one canonical seeded-order contract | [ ] |
|
||||
| API-2 Bind slot allocation to canonical matrix order | [ ] |
|
||||
| API-3 Preserve the complete benchmark baseline | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] [API-1] Add the optional, backward-compatible execution-order seed to loader normalization, schema validation, and manifest digest behavior; verify normal, boundary, omission, and permutation cases.
|
||||
- [ ] [API-2] Prove `RunStore.slots` consumes the seeded canonical matrix order with repetitions nested per cell and does not add a second ordering source.
|
||||
- [ ] [API-3] Run the focused and full benchmark regression commands and confirm all shipped example manifests still validate.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Run applicable required verification and record fresh command/output; repair reviewer-reconstructable evidence gaps instead of forwarding them to another plan.
|
||||
- [ ] For every Required/Suggested finding, record reviewer-collected `Evidence`, exact `Root Cause`, and one `Selected Fix` with affected files/symbols/tests and acceptance commands before creating a follow-up plan.
|
||||
- [ ] Archive active `CODE_REVIEW-cloud-G08.md` to `code_review_cloud_G08_1.log`.
|
||||
- [ ] Archive active `PLAN-local-G08.md` to `plan_local_G08_1.log`.
|
||||
- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write `complete.log` based on `agent-ops/skills/common/code-review/templates/complete-log-template.md` and leave no active `.md` files.
|
||||
- [ ] If PASS, move active task directory `agent-task/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/` to `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/` and update this checklist at the final archive path.
|
||||
- [ ] If PASS, preserve and report `milestone-task` metadata for runtime aggregation, without modifying roadmap or directly calling `update-roadmap`.
|
||||
- [ ] If PASS for split work, remove empty active parent `agent-task/m-iop-one-shot-agent-model-comparison/` or verify it was kept due to remaining siblings/files.
|
||||
- [ ] If WARN/FAIL, write the next filesystem state matching code-review verdict and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record any deviations from the plan and the rationale here._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record key design decisions here._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Confirm omission preserves legacy id ordering and does not inject a default into canonical bytes.
|
||||
- Confirm explicit seeds use the documented domain-separated rank and stable cell-id tie-breaker.
|
||||
- Confirm the JSON schema and Python loader reject the same invalid seed shapes.
|
||||
- Confirm the slot allocator has no independent sorting rule.
|
||||
|
||||
## Verification Results
|
||||
|
||||
Record actual stdout/stderr under each command. If a command changes, document the replacement and reason in `Deviations from Plan`.
|
||||
|
||||
### API-1 Verification
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
### API-2 Verification
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.attempts_test.AttemptStoreTest.test_slots_follow_seeded_manifest_order_before_repetitions
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
### API-3 and Final Verification
|
||||
|
||||
```bash
|
||||
for manifest in scripts/fixtures/agent-comparison-benchmark-manifest.example.json scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json; do python3 scripts/agent_comparison_benchmark.py validate --manifest "$manifest"; done
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these (archive, complete.log, and task-directory archive move are review-agent only) |
|
||||
| Archive Evidence Snapshot | Fixed at stub creation from plan when present | Implementing agent uses it as default prior-loop context; read only the specific archive files cited there when more detail is required |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results (section headings + commands) | Implementing agent, then review agent | Implementing agent records initial output; review agent reruns applicable commands and may fill, replace, or append fresh verified output before verdict. Implementing-agent command changes require a `Deviations from Plan` entry |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,237 @@
|
|||
<!-- task=m-iop-one-shot-agent-model-comparison/01_execution_order_contract plan=1 tag=API milestone-task=matrix-lock -->
|
||||
|
||||
# Plan - Seeded Benchmark Execution Order Contract
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections in `CODE_REVIEW-cloud-G08.md` is the mandatory last implementation step. Run every verification command, record actual notes and stdout/stderr, keep both active files in place, and report ready for review; only the code-review skill may finalize or archive this task. If blocked, record only the exact blocker, attempted commands/output, and resume condition in the implementation-owned evidence fields. Do not ask the user, call user-input tools, create control-plane stop files, classify the next state, archive logs, or write `complete.log`.
|
||||
|
||||
## Background
|
||||
|
||||
The approved benchmark SDD requires an immutable execution-order seed for C01-C09, but the loader currently canonicalizes every matrix only by cell id. This packet adds a backward-compatible optional seed whose deterministic order is part of the normalized manifest and is consumed by the existing slot allocator.
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior plan: `agent-task/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/plan_local_G08_0.log`.
|
||||
- Prior review: `agent-task/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/code_review_cloud_G08_0.log`.
|
||||
- The prior pair contains no implementation evidence or official verdict. Epic self-review found one stale dependency statement: it named a nonexistent integration child and incorrectly implied that child depended directly on both indices 01 and 02.
|
||||
- Replan keeps the implementation contract unchanged and corrects the DAG to `01 -> 03+01`, `02` independent, and `(02,03) -> 04+02,03`. Baseline manifest validation and the 178-test manifest/attempt/connectivity suite were green at `b197e5db70637f87017a024a847e3e53fdc72e8b`.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `scripts/agent_benchmark/manifest.py`
|
||||
- `scripts/agent_benchmark/manifest_test.py`
|
||||
- `scripts/agent_benchmark/attempts.py`
|
||||
- `scripts/agent_benchmark/attempts_test.py`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-manifest.example.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/iop-one-shot-agent-model-comparison.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/iop-one-shot-agent-model-comparison/SDD.md`
|
||||
- `agent-spec/testing/agent-comparison-benchmark.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- SDD: `agent-roadmap/sdd/knowledge-tool-optimization-extension/iop-one-shot-agent-model-comparison/SDD.md`, status `[승인됨]`, implementation lock released, no unresolved user review.
|
||||
- First-line scope: `milestone-task=matrix-lock`.
|
||||
- Target scenario: S03 requires `repetitions=1`, an execution-order seed, fresh session/setup-cache policy, timeout, and expected bindings in an immutable scored manifest.
|
||||
- Evidence Map row: S03 requires an immutable scored manifest and order seed, linked to `matrix-lock` evidence.
|
||||
- This packet implements only the reusable seed/order contract. The dependent locked-manifest packet supplies the concrete nine-cell S03 evidence.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- No separate handoff was supplied. Repository-native evidence came from the loader/schema/tests above and from fresh commands run at starting HEAD `b197e5db70637f87017a024a847e3e53fdc72e8b`.
|
||||
- Current baseline passed all three example `validate` commands and `python3 -m unittest scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test` (`Ran 178 tests`, `OK`).
|
||||
- `Manifest.matrix` is sorted by id at `manifest.py:625-627`; `RunStore.slots` iterates that tuple at `attempts.py:679-681`. Therefore the loader can own the seeded order without a second scheduler or changes to run-state code.
|
||||
- The manifest schema is a public tracked config contract, so normal, omission-compatibility, permutation, invalid-token, digest, and slot-order tests are required. Fresh test output is required; cached output is not accepted.
|
||||
- No external verification is required for this packet.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
- Existing tests cover cell-id canonicalization and repetitions, but no field represents an execution-order seed.
|
||||
- No test proves that two differently ordered JSON inputs with the same explicit seed produce the same matrix order and digest.
|
||||
- No test proves that omitted seed behavior and existing example digests remain backward compatible.
|
||||
- No test binds seeded loader order to the slot sequence.
|
||||
|
||||
### Symbol References
|
||||
|
||||
- No symbol is renamed or removed.
|
||||
- `Manifest(...)` constructor call sites in `manifest.py` and tests must be updated only as required by the new defaulted field; keep the field optional to avoid unrelated call-site churn.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
- This is split child `01_execution_order_contract`. Its stable contract is: one optional seed deterministically canonicalizes `Manifest.matrix`, enters the digest only when explicitly present, and is consumed unchanged by `RunStore.slots`.
|
||||
- PASS evidence is loader/schema regression coverage plus the slot-order test.
|
||||
- It is independent of `02_route_preflight_contract`; neither packet claims the other's files. Only `03+01_rubric_version_contract` depends directly on this packet. The final `04+02,03_locked_benchmark_manifest` packet depends on indices 02 and 03.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
- Do not create the C01-C09 manifest here; the dependent integration packet owns it.
|
||||
- Do not change preflight observation coverage, caller adapters, run-record schemas, pipeline version, fixture bytes, contracts, or IOP configuration.
|
||||
- Do not reorder manifests that omit the field and do not rewrite the existing example manifests solely to exercise an optional field.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- `evaluation_mode=isolated-reassessment`; `finalizer=finalize-task-policy.sh`, `finalizer_mode=pair`.
|
||||
- Build closures: scope/context/verification/evidence/ownership/decision all closed. Scores `2+1+2+1+2=G08`; base/final route `local-fit`, lane `local`, catalog `worker/local/G08`, filename `PLAN-local-G08.md`.
|
||||
- Review closures: all closed. Scores `2+1+2+1+2=G08`; route `official-review`, lane `cloud`, catalog `review/cloud/G08`, filename `CODE_REVIEW-cloud-G08.md`.
|
||||
- `large_indivisible_context=false`; positive loop risks: `boundary_contract`, `structured_interpretation`, `variant_product` (`count=3`); `review_rework_count=0`; `evidence_integrity_failure=false`; no capability gap.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] [API-1] Add the optional, backward-compatible execution-order seed to loader normalization, schema validation, and manifest digest behavior; verify normal, boundary, omission, and permutation cases.
|
||||
- [ ] [API-2] Prove `RunStore.slots` consumes the seeded canonical matrix order with repetitions nested per cell and does not add a second ordering source.
|
||||
- [ ] [API-3] Run the focused and full benchmark regression commands and confirm all shipped example manifests still validate.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Add one canonical seeded-order contract
|
||||
|
||||
**Problem**
|
||||
|
||||
`Manifest` has no seed field (`manifest.py:112-126`), the top-level allowlist accepts only `repetitions` as optional (`manifest.py:681-694`), and `_validate_matrix` always sorts by id (`manifest.py:611-627`). Consequently S03 cannot persist or reproduce a deliberately seeded C01-C09 order.
|
||||
|
||||
**Solution**
|
||||
|
||||
Keep pipeline version 2 backward compatible. Append `execution_order_seed: str | None = None` after the existing required `digest` field in `Manifest` so dataclass default ordering stays valid, accept it as an optional top-level field using the existing cell-id token constraints, and add the same optional constraint to the JSON schema. Pass an explicit seed into both loader-created `Manifest` values before digesting. When absent, preserve exact cell-id sorting and omit the field from canonical serialization so prior normalized digests remain unchanged. When present, order cells by a fixed domain-separated SHA-256 rank with cell id as the collision tie-breaker, include the seed in canonical serialization, and therefore bind it into `digest_manifest_and_resolved_inputs`.
|
||||
|
||||
Before (`scripts/agent_benchmark/manifest.py:625-627`):
|
||||
|
||||
```python
|
||||
# Sort cells by id for canonical ordering
|
||||
cells.sort(key=lambda c: c.id)
|
||||
return tuple(cells)
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
def _execution_order_key(seed: str, cell: MatrixCell) -> tuple[bytes, str]:
|
||||
rank = hashlib.sha256(
|
||||
b"iop-benchmark-order-v1\0" + seed.encode("ascii") + b"\0" + cell.id.encode("ascii")
|
||||
).digest()
|
||||
return rank, cell.id
|
||||
|
||||
# No seed preserves the pipeline-v2 cell-id order. An explicit seed is stable
|
||||
# across JSON input permutations and platforms.
|
||||
cells.sort(
|
||||
key=(lambda cell: cell.id)
|
||||
if execution_order_seed is None
|
||||
else (lambda cell: _execution_order_key(execution_order_seed, cell))
|
||||
)
|
||||
```
|
||||
|
||||
Construct the canonical dict first, then conditionally add `execution_order_seed`; never serialize a synthetic default. Add comments documenting the domain string and compatibility behavior.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `scripts/agent_benchmark/manifest.py`: model, validate, normalize, order, and digest the explicit seed.
|
||||
- [ ] `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json`: accept only the optional bounded seed token while retaining `additionalProperties: false`.
|
||||
- [ ] `scripts/agent_benchmark/manifest_test.py`: add exact regression tests named `test_execution_order_seed_is_canonical_and_permutation_stable`, `test_omitted_execution_order_seed_preserves_legacy_order_and_digest_contract`, and schema/loader invalid-seed parity coverage.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
Write tests. Use at least three cells whose SHA-256 seeded order differs from lexical id order. Assert input permutation independence, same-seed repeatability, different-seed digest/order differentiation, invalid empty/oversize/non-token rejection, and omission compatibility. Existing example validation must remain unchanged.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test
|
||||
```
|
||||
|
||||
Expected: all manifest loader/schema/digest tests pass freshly.
|
||||
|
||||
### [API-2] Bind slot allocation to the canonical matrix order
|
||||
|
||||
**Problem**
|
||||
|
||||
`RunStore.slots` already iterates `manifest.matrix` (`attempts.py:679-681`), but the current test only checks repetitions for a one-cell manifest (`attempts_test.py:457`). There is no regression proof that a seeded multi-cell order reaches attempt allocation without being re-sorted.
|
||||
|
||||
**Solution**
|
||||
|
||||
Do not modify `RunStore.slots`. Extend the attempts-test manifest helper to accept an explicit seed and multiple cells, then assert the exact sequence is the loader's seeded `Manifest.matrix` order with repetitions `1..N` adjacent for each cell.
|
||||
|
||||
Before (`scripts/agent_benchmark/attempts.py:679-681`):
|
||||
|
||||
```python
|
||||
@staticmethod
|
||||
def slots(manifest: Manifest) -> tuple[Slot, ...]:
|
||||
return tuple(Slot(cell.id, repetition) for cell in manifest.matrix for repetition in range(1, manifest.repetitions + 1))
|
||||
```
|
||||
|
||||
After: keep this production code unchanged; lock its behavior with `test_slots_follow_seeded_manifest_order_before_repetitions`.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `scripts/agent_benchmark/attempts_test.py`: add the seeded multi-cell slot-order regression without changing production attempts code.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
Write the named test. Assert both `(cell_id, repetition)` tuples and that the first occurrence order equals `manifest.matrix`. This prevents a future alphabetical sort in the store from silently defeating the seed.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.attempts_test.AttemptStoreTest.test_slots_follow_seeded_manifest_order_before_repetitions
|
||||
```
|
||||
|
||||
Expected: the exact seeded slot sequence passes.
|
||||
|
||||
### [API-3] Preserve the complete benchmark baseline
|
||||
|
||||
**Problem**
|
||||
|
||||
The field touches canonical JSON and a public schema; focused tests alone could miss example or integration drift.
|
||||
|
||||
**Solution**
|
||||
|
||||
Run fresh benchmark suites and all shipped validation fixtures after the code and tests are complete. Do not accept cached output.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `CODE_REVIEW-cloud-G08.md`: record actual implementation decisions and full command output.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
No additional test file beyond API-1/API-2. The full existing suite is the regression oracle.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
for manifest in scripts/fixtures/agent-comparison-benchmark-manifest.example.json scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json; do python3 scripts/agent_comparison_benchmark.py validate --manifest "$manifest"; done
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Expected: three `ok: manifest is valid` lines, all tests `OK`, and no whitespace errors.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Item |
|
||||
|------|------|
|
||||
| `scripts/agent_benchmark/manifest.py` | API-1 |
|
||||
| `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json` | API-1 |
|
||||
| `scripts/agent_benchmark/manifest_test.py` | API-1 |
|
||||
| `scripts/agent_benchmark/attempts_test.py` | API-2 |
|
||||
| `agent-task/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/CODE_REVIEW-cloud-G08.md` | API-3 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
Run from `/config/workspace/iop-s0`; fresh output is required and cached test output is not acceptable.
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test
|
||||
python3 -m unittest scripts.agent_benchmark.attempts_test.AttemptStoreTest.test_slots_follow_seeded_manifest_order_before_repetitions
|
||||
for manifest in scripts/fixtures/agent-comparison-benchmark-manifest.example.json scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json; do python3 scripts/agent_comparison_benchmark.py validate --manifest "$manifest"; done
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Expected: every command exits 0, the three manifests validate, all benchmark tests report `OK`, and `git diff --check` is silent.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,131 @@
|
|||
<!-- task=m-iop-one-shot-agent-model-comparison/01_execution_order_contract plan=0 tag=API milestone-task=matrix-lock -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-12
|
||||
task=m-iop-one-shot-agent-model-comparison/01_execution_order_contract, plan=0, tag=API
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files. Run the applicable verification commands directly and record fresh output in `Verification Results`; implementation-owned output is handoff evidence, not a substitute for reviewer verification. If implementation is present, repair missing or stale verification output instead of failing solely for insufficient recorded evidence. When verification exposes a defect, collect the necessary data, determine the exact root cause, and select one concrete fix before generating the follow-up plan; never delegate investigation or remedy selection to the worker.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals.
|
||||
2. Archive `CODE_REVIEW-cloud-G08.md` → `code_review_cloud_G08_0.log` and `PLAN-local-G08.md` → `plan_local_G08_0.log`.
|
||||
3. If PASS, write `complete.log` and move active task directory to `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/`. If WARN/FAIL, fully write the next filesystem state required by the code-review skill.
|
||||
4. If PASS, preserve the first-line `milestone-task` metadata in `complete.log` and report it for runtime aggregation. Roadmap state evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check applicable `Review-Only Checklist` items at the final `.log` location before reporting.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|---------|
|
||||
| API-1 Add one canonical seeded-order contract | [ ] |
|
||||
| API-2 Bind slot allocation to canonical matrix order | [ ] |
|
||||
| API-3 Preserve the complete benchmark baseline | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] [API-1] Add the optional, backward-compatible execution-order seed to loader normalization, schema validation, and manifest digest behavior; verify normal, boundary, omission, and permutation cases.
|
||||
- [ ] [API-2] Prove `RunStore.slots` consumes the seeded canonical matrix order with repetitions nested per cell and does not add a second ordering source.
|
||||
- [ ] [API-3] Run the focused and full benchmark regression commands and confirm all shipped example manifests still validate.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Run applicable required verification and record fresh command/output; repair reviewer-reconstructable evidence gaps instead of forwarding them to another plan.
|
||||
- [ ] For every Required/Suggested finding, record reviewer-collected `Evidence`, exact `Root Cause`, and one `Selected Fix` with affected files/symbols/tests and acceptance commands before creating a follow-up plan.
|
||||
- [ ] Archive active `CODE_REVIEW-cloud-G08.md` to `code_review_cloud_G08_0.log`.
|
||||
- [ ] Archive active `PLAN-local-G08.md` to `plan_local_G08_0.log`.
|
||||
- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write `complete.log` based on `agent-ops/skills/common/code-review/templates/complete-log-template.md` and leave no active `.md` files.
|
||||
- [ ] If PASS, move active task directory `agent-task/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/` to `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/` and update this checklist at the final archive path.
|
||||
- [ ] If PASS, preserve and report `milestone-task` metadata for runtime aggregation, without modifying roadmap or directly calling `update-roadmap`.
|
||||
- [ ] If PASS for split work, remove empty active parent `agent-task/m-iop-one-shot-agent-model-comparison/` or verify it was kept due to remaining siblings/files.
|
||||
- [ ] If WARN/FAIL, write the next filesystem state matching code-review verdict and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record any deviations from the plan and the rationale here._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record key design decisions here._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Confirm omission preserves legacy id ordering and does not inject a default into canonical bytes.
|
||||
- Confirm explicit seeds use the documented domain-separated rank and stable cell-id tie-breaker.
|
||||
- Confirm the JSON schema and Python loader reject the same invalid seed shapes.
|
||||
- Confirm the slot allocator has no independent sorting rule.
|
||||
|
||||
## Verification Results
|
||||
|
||||
Record actual stdout/stderr under each command. If a command changes, document the replacement and reason in `Deviations from Plan`.
|
||||
|
||||
### API-1 Verification
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
### API-2 Verification
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.attempts_test.AttemptStoreTest.test_slots_follow_seeded_manifest_order_before_repetitions
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
### API-3 and Final Verification
|
||||
|
||||
```bash
|
||||
for manifest in scripts/fixtures/agent-comparison-benchmark-manifest.example.json scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json; do python3 scripts/agent_comparison_benchmark.py validate --manifest "$manifest"; done
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these (archive, complete.log, and task-directory archive move are review-agent only) |
|
||||
| Archive Evidence Snapshot | Fixed at stub creation from plan when present | Implementing agent uses it as default prior-loop context; read only the specific archive files cited there when more detail is required |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results (section headings + commands) | Implementing agent, then review agent | Implementing agent records initial output; review agent reruns applicable commands and may fill, replace, or append fresh verified output before verdict. Implementing-agent command changes require a `Deviations from Plan` entry |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,230 @@
|
|||
<!-- task=m-iop-one-shot-agent-model-comparison/01_execution_order_contract plan=0 tag=API milestone-task=matrix-lock -->
|
||||
|
||||
# Plan - Seeded Benchmark Execution Order Contract
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections in `CODE_REVIEW-cloud-G08.md` is the mandatory last implementation step. Run every verification command, record actual notes and stdout/stderr, keep both active files in place, and report ready for review; only the code-review skill may finalize or archive this task. If blocked, record only the exact blocker, attempted commands/output, and resume condition in the implementation-owned evidence fields. Do not ask the user, call user-input tools, create control-plane stop files, classify the next state, archive logs, or write `complete.log`.
|
||||
|
||||
## Background
|
||||
|
||||
The approved benchmark SDD requires an immutable execution-order seed for C01-C09, but the loader currently canonicalizes every matrix only by cell id. This packet adds a backward-compatible optional seed whose deterministic order is part of the normalized manifest and is consumed by the existing slot allocator.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `scripts/agent_benchmark/manifest.py`
|
||||
- `scripts/agent_benchmark/manifest_test.py`
|
||||
- `scripts/agent_benchmark/attempts.py`
|
||||
- `scripts/agent_benchmark/attempts_test.py`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-manifest.example.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/iop-one-shot-agent-model-comparison.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/iop-one-shot-agent-model-comparison/SDD.md`
|
||||
- `agent-spec/testing/agent-comparison-benchmark.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- SDD: `agent-roadmap/sdd/knowledge-tool-optimization-extension/iop-one-shot-agent-model-comparison/SDD.md`, status `[승인됨]`, implementation lock released, no unresolved user review.
|
||||
- First-line scope: `milestone-task=matrix-lock`.
|
||||
- Target scenario: S03 requires `repetitions=1`, an execution-order seed, fresh session/setup-cache policy, timeout, and expected bindings in an immutable scored manifest.
|
||||
- Evidence Map row: S03 requires an immutable scored manifest and order seed, linked to `matrix-lock` evidence.
|
||||
- This packet implements only the reusable seed/order contract. The dependent locked-manifest packet supplies the concrete nine-cell S03 evidence.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- No separate handoff was supplied. Repository-native evidence came from the loader/schema/tests above and from fresh commands run at starting HEAD `b197e5db70637f87017a024a847e3e53fdc72e8b`.
|
||||
- Current baseline passed all three example `validate` commands and `python3 -m unittest scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test` (`Ran 178 tests`, `OK`).
|
||||
- `Manifest.matrix` is sorted by id at `manifest.py:625-627`; `RunStore.slots` iterates that tuple at `attempts.py:679-681`. Therefore the loader can own the seeded order without a second scheduler or changes to run-state code.
|
||||
- The manifest schema is a public tracked config contract, so normal, omission-compatibility, permutation, invalid-token, digest, and slot-order tests are required. Fresh test output is required; cached output is not accepted.
|
||||
- No external verification is required for this packet.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
- Existing tests cover cell-id canonicalization and repetitions, but no field represents an execution-order seed.
|
||||
- No test proves that two differently ordered JSON inputs with the same explicit seed produce the same matrix order and digest.
|
||||
- No test proves that omitted seed behavior and existing example digests remain backward compatible.
|
||||
- No test binds seeded loader order to the slot sequence.
|
||||
|
||||
### Symbol References
|
||||
|
||||
- No symbol is renamed or removed.
|
||||
- `Manifest(...)` constructor call sites in `manifest.py` and tests must be updated only as required by the new defaulted field; keep the field optional to avoid unrelated call-site churn.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
- This is split child `01_execution_order_contract`. Its stable contract is: one optional seed deterministically canonicalizes `Manifest.matrix`, enters the digest only when explicitly present, and is consumed unchanged by `RunStore.slots`.
|
||||
- PASS evidence is loader/schema regression coverage plus the slot-order test.
|
||||
- It is independent of `02_route_preflight_contract`; neither packet claims the other's files. `03+01,02_locked_benchmark_manifest` depends on both.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
- Do not create the C01-C09 manifest here; the dependent integration packet owns it.
|
||||
- Do not change preflight observation coverage, caller adapters, run-record schemas, pipeline version, fixture bytes, contracts, or IOP configuration.
|
||||
- Do not reorder manifests that omit the field and do not rewrite the existing example manifests solely to exercise an optional field.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- `evaluation_mode=first-pass`; `finalizer=finalize-task-policy.sh`, `finalizer_mode=pair`.
|
||||
- Build closures: scope/context/verification/evidence/ownership/decision all closed. Scores `2+1+2+1+2=G08`; base/final route `local-fit`, lane `local`, catalog `worker/local/G08`, filename `PLAN-local-G08.md`.
|
||||
- Review closures: all closed. Scores `2+1+2+1+2=G08`; route `official-review`, lane `cloud`, catalog `review/cloud/G08`, filename `CODE_REVIEW-cloud-G08.md`.
|
||||
- `large_indivisible_context=false`; positive loop risks: `boundary_contract`, `structured_interpretation`, `variant_product` (`count=3`); `review_rework_count=0`; `evidence_integrity_failure=false`; no capability gap.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] [API-1] Add the optional, backward-compatible execution-order seed to loader normalization, schema validation, and manifest digest behavior; verify normal, boundary, omission, and permutation cases.
|
||||
- [ ] [API-2] Prove `RunStore.slots` consumes the seeded canonical matrix order with repetitions nested per cell and does not add a second ordering source.
|
||||
- [ ] [API-3] Run the focused and full benchmark regression commands and confirm all shipped example manifests still validate.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Add one canonical seeded-order contract
|
||||
|
||||
**Problem**
|
||||
|
||||
`Manifest` has no seed field (`manifest.py:112-126`), the top-level allowlist accepts only `repetitions` as optional (`manifest.py:681-694`), and `_validate_matrix` always sorts by id (`manifest.py:611-627`). Consequently S03 cannot persist or reproduce a deliberately seeded C01-C09 order.
|
||||
|
||||
**Solution**
|
||||
|
||||
Keep pipeline version 2 backward compatible. Add `execution_order_seed: str | None = None` to `Manifest`, accept it as an optional top-level field using the existing cell-id token constraints, and add the same optional constraint to the JSON schema. When absent, preserve exact cell-id sorting and omit the field from canonical serialization so prior normalized digests remain unchanged. When present, order cells by a fixed domain-separated SHA-256 rank with cell id as the collision tie-breaker, include the seed in canonical serialization, and therefore bind it into `digest_manifest_and_resolved_inputs`.
|
||||
|
||||
Before (`scripts/agent_benchmark/manifest.py:625-627`):
|
||||
|
||||
```python
|
||||
# Sort cells by id for canonical ordering
|
||||
cells.sort(key=lambda c: c.id)
|
||||
return tuple(cells)
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
def _execution_order_key(seed: str, cell: MatrixCell) -> tuple[bytes, str]:
|
||||
rank = hashlib.sha256(
|
||||
b"iop-benchmark-order-v1\0" + seed.encode("ascii") + b"\0" + cell.id.encode("ascii")
|
||||
).digest()
|
||||
return rank, cell.id
|
||||
|
||||
# No seed preserves the pipeline-v2 cell-id order. An explicit seed is stable
|
||||
# across JSON input permutations and platforms.
|
||||
cells.sort(
|
||||
key=(lambda cell: cell.id)
|
||||
if execution_order_seed is None
|
||||
else (lambda cell: _execution_order_key(execution_order_seed, cell))
|
||||
)
|
||||
```
|
||||
|
||||
Construct the canonical dict first, then conditionally add `execution_order_seed`; never serialize a synthetic default. Add comments documenting the domain string and compatibility behavior.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `scripts/agent_benchmark/manifest.py`: model, validate, normalize, order, and digest the explicit seed.
|
||||
- [ ] `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json`: accept only the optional bounded seed token while retaining `additionalProperties: false`.
|
||||
- [ ] `scripts/agent_benchmark/manifest_test.py`: add exact regression tests named `test_execution_order_seed_is_canonical_and_permutation_stable`, `test_omitted_execution_order_seed_preserves_legacy_order_and_digest_contract`, and schema/loader invalid-seed parity coverage.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
Write tests. Use at least three cells whose SHA-256 seeded order differs from lexical id order. Assert input permutation independence, same-seed repeatability, different-seed digest/order differentiation, invalid empty/oversize/non-token rejection, and omission compatibility. Existing example validation must remain unchanged.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test
|
||||
```
|
||||
|
||||
Expected: all manifest loader/schema/digest tests pass freshly.
|
||||
|
||||
### [API-2] Bind slot allocation to the canonical matrix order
|
||||
|
||||
**Problem**
|
||||
|
||||
`RunStore.slots` already iterates `manifest.matrix` (`attempts.py:679-681`), but the current test only checks repetitions for a one-cell manifest (`attempts_test.py:457`). There is no regression proof that a seeded multi-cell order reaches attempt allocation without being re-sorted.
|
||||
|
||||
**Solution**
|
||||
|
||||
Do not modify `RunStore.slots`. Extend the attempts-test manifest helper to accept an explicit seed and multiple cells, then assert the exact sequence is the loader's seeded `Manifest.matrix` order with repetitions `1..N` adjacent for each cell.
|
||||
|
||||
Before (`scripts/agent_benchmark/attempts.py:679-681`):
|
||||
|
||||
```python
|
||||
@staticmethod
|
||||
def slots(manifest: Manifest) -> tuple[Slot, ...]:
|
||||
return tuple(Slot(cell.id, repetition) for cell in manifest.matrix for repetition in range(1, manifest.repetitions + 1))
|
||||
```
|
||||
|
||||
After: keep this production code unchanged; lock its behavior with `test_slots_follow_seeded_manifest_order_before_repetitions`.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `scripts/agent_benchmark/attempts_test.py`: add the seeded multi-cell slot-order regression without changing production attempts code.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
Write the named test. Assert both `(cell_id, repetition)` tuples and that the first occurrence order equals `manifest.matrix`. This prevents a future alphabetical sort in the store from silently defeating the seed.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.attempts_test.AttemptStoreTest.test_slots_follow_seeded_manifest_order_before_repetitions
|
||||
```
|
||||
|
||||
Expected: the exact seeded slot sequence passes.
|
||||
|
||||
### [API-3] Preserve the complete benchmark baseline
|
||||
|
||||
**Problem**
|
||||
|
||||
The field touches canonical JSON and a public schema; focused tests alone could miss example or integration drift.
|
||||
|
||||
**Solution**
|
||||
|
||||
Run fresh benchmark suites and all shipped validation fixtures after the code and tests are complete. Do not accept cached output.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `CODE_REVIEW-cloud-G08.md`: record actual implementation decisions and full command output.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
No additional test file beyond API-1/API-2. The full existing suite is the regression oracle.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
for manifest in scripts/fixtures/agent-comparison-benchmark-manifest.example.json scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json; do python3 scripts/agent_comparison_benchmark.py validate --manifest "$manifest"; done
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Expected: three `ok: manifest is valid` lines, all tests `OK`, and no whitespace errors.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Item |
|
||||
|------|------|
|
||||
| `scripts/agent_benchmark/manifest.py` | API-1 |
|
||||
| `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json` | API-1 |
|
||||
| `scripts/agent_benchmark/manifest_test.py` | API-1 |
|
||||
| `scripts/agent_benchmark/attempts_test.py` | API-2 |
|
||||
| `agent-task/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/CODE_REVIEW-cloud-G08.md` | API-3 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
Run from `/config/workspace/iop-s0`; fresh output is required and cached test output is not acceptable.
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test
|
||||
python3 -m unittest scripts.agent_benchmark.attempts_test.AttemptStoreTest.test_slots_follow_seeded_manifest_order_before_repetitions
|
||||
for manifest in scripts/fixtures/agent-comparison-benchmark-manifest.example.json scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json; do python3 scripts/agent_comparison_benchmark.py validate --manifest "$manifest"; done
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Expected: every command exits 0, the three manifests validate, all benchmark tests report `OK`, and `git diff --check` is silent.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,144 @@
|
|||
<!-- task=m-iop-one-shot-agent-model-comparison/02_route_preflight_contract plan=1 tag=API milestone-task=route-readiness -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-12
|
||||
task=m-iop-one-shot-agent-model-comparison/02_route_preflight_contract, plan=1, tag=API
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior plan: `agent-task/m-iop-one-shot-agent-model-comparison/02_route_preflight_contract/plan_cloud_G09_0.log`.
|
||||
- Prior review: `agent-task/m-iop-one-shot-agent-model-comparison/02_route_preflight_contract/code_review_cloud_G09_0.log`.
|
||||
- The prior pair contains no implementation evidence or official verdict. Epic self-review found that it changed public preflight semantics but omitted the project benchmark skill and its contract tests, both of which explicitly require the obsolete direct-only behavior.
|
||||
- Replan makes this packet the sole owner of all-cell preflight documentation/tests, adds the concrete agy preflight type source to analysis, and leaves rubric-selection wording for the dependent final-manifest packet. Baseline manifest validation and the 178-test manifest/attempt/connectivity suite were green at `b197e5db70637f87017a024a847e3e53fdc72e8b`.
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files. Run the applicable verification commands directly and record fresh output in `Verification Results`; implementation-owned output is handoff evidence, not a substitute for reviewer verification. If implementation is present, repair missing or stale verification output instead of failing solely for insufficient recorded evidence. When verification exposes a defect, collect the necessary data, determine the exact root cause, and select one concrete fix before generating the follow-up plan; never delegate investigation or remedy selection to the worker.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals.
|
||||
2. Archive `CODE_REVIEW-cloud-G09.md` → `code_review_cloud_G09_1.log` and `PLAN-cloud-G09.md` → `plan_cloud_G09_1.log`.
|
||||
3. If PASS, write `complete.log` and move active task directory to `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/02_route_preflight_contract/`. If WARN/FAIL, fully write the next filesystem state required by the code-review skill.
|
||||
4. If PASS, preserve the first-line `milestone-task` metadata in `complete.log` and report it for runtime aggregation. Roadmap state evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check applicable `Review-Only Checklist` items at the final `.log` location before reporting.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|---------|
|
||||
| API-1 Require exact all-cell preflight evidence | [ ] |
|
||||
| API-2 Isolate live agy preflight state per cell | [ ] |
|
||||
| API-3 Replace obsolete direct-only code and skill assertions | [ ] |
|
||||
| API-4 Preserve append-only and redaction invariants | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] [API-1] Make collection and append-only validation require one canonical observation for every matrix cell, including preset-only and mixed manifests, with no attempt allocation on blockers.
|
||||
- [ ] [API-2] Store and consume agy live preflight state by cell id, clearing stale admission on every re-preflight and rejecting missing or mismatched state.
|
||||
- [ ] [API-3] Replace direct-only code and project-skill contract tests with mixed/preset all-cell success, blocker, persistence, execution, and per-cell agy regressions.
|
||||
- [ ] [API-4] Run the focused and full benchmark regression suites and confirm secret-free append-only evidence remains valid.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Run applicable required verification and record fresh command/output; repair reviewer-reconstructable evidence gaps instead of forwarding them to another plan.
|
||||
- [ ] For every Required/Suggested finding, record reviewer-collected `Evidence`, exact `Root Cause`, and one `Selected Fix` with affected files/symbols/tests and acceptance commands before creating a follow-up plan.
|
||||
- [ ] Archive active `CODE_REVIEW-cloud-G09.md` to `code_review_cloud_G09_1.log`.
|
||||
- [ ] Archive active `PLAN-cloud-G09.md` to `plan_cloud_G09_1.log`.
|
||||
- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write `complete.log` based on `agent-ops/skills/common/code-review/templates/complete-log-template.md` and leave no active `.md` files.
|
||||
- [ ] If PASS, move active task directory `agent-task/m-iop-one-shot-agent-model-comparison/02_route_preflight_contract/` to `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/02_route_preflight_contract/` and update this checklist at the final archive path.
|
||||
- [ ] If PASS, preserve and report `milestone-task` metadata for runtime aggregation, without modifying roadmap or directly calling `update-roadmap`.
|
||||
- [ ] If PASS for split work, remove empty active parent `agent-task/m-iop-one-shot-agent-model-comparison/` or verify it was kept due to remaining siblings/files.
|
||||
- [ ] If WARN/FAIL, write the next filesystem state matching code-review verdict and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record any deviations from the plan and the rationale here._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record key design decisions here._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Confirm the durable result set equals all `manifest.matrix` cells in canonical order on both write and read.
|
||||
- Confirm preset blockers publish exact closed taxonomy and allocate no attempt.
|
||||
- Confirm agy runtime state is keyed by cell id and a blocked re-preflight clears stale ready state.
|
||||
- Confirm the project benchmark skill and contract tests require one observation for every immutable matrix cell and contain no direct-only/preset-local exception.
|
||||
- Confirm no adapter-specific endpoint, token, or raw caller value enters evidence.
|
||||
|
||||
## Verification Results
|
||||
|
||||
Record actual stdout/stderr under each command. If a command changes, document the replacement and reason in `Deviations from Plan`.
|
||||
|
||||
### API-1/API-2 Focused Verification
|
||||
|
||||
```bash
|
||||
python3 -m unittest \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_cli_mixed_manifest_preflights_and_invokes_every_cell \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_preset_only_public_preflight_appends_exact_results \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_preset_blocker_appends_without_attempt_allocation \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_live_agy_multiple_cells_consume_their_own_preflight_state
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
### API-3 Verification
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.connectivity_integration_test scripts.agent_benchmark.skill_contract_test
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
### API-4 and Final Verification
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test scripts.agent_benchmark.skill_contract_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these (archive, complete.log, and task-directory archive move are review-agent only) |
|
||||
| Archive Evidence Snapshot | Fixed at stub creation from plan when present | Implementing agent uses it as default prior-loop context; read only the specific archive files cited there when more detail is required |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results (section headings + commands) | Implementing agent, then review agent | Implementing agent records initial output; review agent reruns applicable commands and may fill, replace, or append fresh verified output before verdict. Implementing-agent command changes require a `Deviations from Plan` entry |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,292 @@
|
|||
<!-- task=m-iop-one-shot-agent-model-comparison/02_route_preflight_contract plan=1 tag=API milestone-task=route-readiness -->
|
||||
|
||||
# Plan - All-Cell Route Preflight Contract
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections in `CODE_REVIEW-cloud-G09.md` is the mandatory last implementation step. Run every verification command, record actual notes and stdout/stderr, keep both active files in place, and report ready for review; only the code-review skill may finalize or archive this task. If blocked, record only the exact blocker, attempted commands/output, and resume condition in the implementation-owned evidence fields. Do not ask the user, call user-input tools, create control-plane stop files, classify the next state, archive logs, or write `complete.log`.
|
||||
|
||||
## Background
|
||||
|
||||
The benchmark currently validates preset capability locally but deliberately omits preset cells from durable preflight. A mixed direct/preset manifest therefore records a misleading ready subset and allocates no attempts, while the live `agy` adapter retains only the last cell's preflight state. Route readiness for C01-C09 requires every declared cell to be observed and bound independently before any attempt starts.
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior plan: `agent-task/m-iop-one-shot-agent-model-comparison/02_route_preflight_contract/plan_cloud_G09_0.log`.
|
||||
- Prior review: `agent-task/m-iop-one-shot-agent-model-comparison/02_route_preflight_contract/code_review_cloud_G09_0.log`.
|
||||
- The prior pair contains no implementation evidence or official verdict. Epic self-review found that it changed public preflight semantics but omitted the project benchmark skill and its contract tests, both of which explicitly require the obsolete direct-only behavior.
|
||||
- Replan makes this packet the sole owner of all-cell preflight documentation/tests, adds the concrete agy preflight type source to analysis, and leaves rubric-selection wording for the dependent final-manifest packet. Baseline manifest validation and the 178-test manifest/attempt/connectivity suite were green at `b197e5db70637f87017a024a847e3e53fdc72e8b`.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `scripts/agent_benchmark/attempts.py`
|
||||
- `scripts/agent_benchmark/attempts_test.py`
|
||||
- `scripts/agent_benchmark/live_iop.py`
|
||||
- `scripts/agent_benchmark/agy_iop.py`
|
||||
- `scripts/agent_benchmark/connectivity_integration_test.py`
|
||||
- `scripts/agent_benchmark/skill_contract_test.py`
|
||||
- `scripts/agent_benchmark/manifest.py`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-manifest.example.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json`
|
||||
- `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/iop-one-shot-agent-model-comparison.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/iop-one-shot-agent-model-comparison/SDD.md`
|
||||
- `agent-spec/testing/agent-comparison-benchmark.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- SDD: `agent-roadmap/sdd/knowledge-tool-optimization-extension/iop-one-shot-agent-model-comparison/SDD.md`, status `[승인됨]`, implementation lock released, no unresolved user review.
|
||||
- First-line scope: `milestone-task=route-readiness`.
|
||||
- Target scenario: S02 requires execution-day auth, model/preset, effort, stream/finish/idle checks for all C01-C09 or an exact blocker.
|
||||
- Evidence Map row: S02 requires a redacted C01-C09 preflight matrix with auth/route/effort/terminal evidence.
|
||||
- The checklist therefore replaces direct-only subset semantics with exact all-cell evidence and adds per-cell caller-state regressions. Live credentials remain the dependent manifest packet's external gate.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- No separate handoff was supplied. Repository-native evidence came from the attempts/live adapter paths and network-free integration tests above.
|
||||
- Fresh baseline at starting HEAD `b197e5db70637f87017a024a847e3e53fdc72e8b`: all three example manifests validated and the 178-test manifest/attempts/connectivity suite passed.
|
||||
- Current tests explicitly lock the old limitation: `test_cli_mixed_manifest_never_invokes_unobserved_preset_cells`, `test_generic_preset_cells_are_local_contract_only`, and `test_generic_preset_only_public_preflight_fails_closed_without_run_state` (`connectivity_integration_test.py:1009-1214`). These must be replaced, not preserved.
|
||||
- `collect_preflight_observations` skips non-direct cells (`attempts.py:386-423`), durable read/write validates against `_direct_cells` (`attempts.py:514-516`, `563-568`, `632-651`), and public preflight rejects an empty direct subset (`attempts.py:1869-1881`).
|
||||
- `_LiveAdapter` stores one `_agy_preflight` value for every cell (`live_iop.py:843-902`) and reuses it at invocation (`live_iop.py:935-950`), so multiple agy routes can cross-bind.
|
||||
- This packet uses only network-free adapters and mocks. Fresh output is required; cached output is not accepted.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
- Existing tests cover direct cells and intentionally assert that preset cells are omitted.
|
||||
- No mixed manifest test expects both direct and preset results in the same append-only record and then executes both.
|
||||
- No preset-only public preflight test expects a durable ready or blocker record.
|
||||
- No live-adapter test preflights two agy cells and proves each invocation receives its own preflight capability/runtime binding.
|
||||
- Existing corruption and secret-redaction coverage is broad and must continue to pass after the result-set invariant changes.
|
||||
- The project benchmark skill and `skill_contract_test.py` explicitly describe/require direct-only preflight, so they would contradict production immediately after this packet unless changed in the same boundary.
|
||||
|
||||
### Symbol References
|
||||
|
||||
- Rename private `RunStore._direct_cells` to `_preflight_cells` or remove the helper. Its only references are `attempts.py:563` and `attempts.py:632`; update both together.
|
||||
- Rename private `_LiveAdapter._agy_preflight` to `_agy_preflights`; all references are `live_iop.py:858`, `885`, `893-894`, and `936-949`.
|
||||
- No public import or CLI symbol is renamed.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
- This is split child `02_route_preflight_contract`. Its stable invariant is: the preflight result set equals the entire immutable matrix, and every admitted live invocation consumes the preflight state for its own cell id.
|
||||
- PASS evidence is all-cell append/read corruption coverage, mixed/preset-only CLI coverage, and a two-cell agy state test.
|
||||
- It is independent of `01_execution_order_contract`; it consumes `manifest.matrix` order without owning that order. The final `04+02,03_locked_benchmark_manifest` packet depends on this packet and index 03.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
- Do not add the final nine-cell manifest or an execution-order field here.
|
||||
- Do not change caller wire formats, IOP contracts, provider config, credentials, retry policy, attempt allocation semantics, or scoring.
|
||||
- Do not synthesize readiness for missing routes. Every preset must pass the same live config/catalog/capability checks and retain exact `registration_required` or `implementation_gap` taxonomy.
|
||||
- Update only the skill's preflight/run/resume/validation/safety wording here. Do not add the new rubric-selection wording; the dependent final-manifest packet owns that operational transition after the rubric implementation lands.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- `evaluation_mode=isolated-reassessment`; `finalizer=finalize-task-policy.sh`, `finalizer_mode=pair`.
|
||||
- Build closures: scope/context/verification/evidence/ownership/decision all closed. Scores `2+2+2+1+2=G09`; base/final route `grade-boundary`, lane `cloud`, catalog `worker/cloud/G09`, filename `PLAN-cloud-G09.md`.
|
||||
- Review closures: all closed. Scores `2+2+2+1+2=G09`; route `official-review`, lane `cloud`, catalog `review/cloud/G09`, filename `CODE_REVIEW-cloud-G09.md`.
|
||||
- `large_indivisible_context=false`; positive loop risks: `temporal_state`, `boundary_contract`, `structured_interpretation`, `variant_product` (`count=4`, risk boundary matched but does not replace grade basis); `review_rework_count=0`; `evidence_integrity_failure=false`; no capability gap.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] [API-1] Make collection and append-only validation require one canonical observation for every matrix cell, including preset-only and mixed manifests, with no attempt allocation on blockers.
|
||||
- [ ] [API-2] Store and consume agy live preflight state by cell id, clearing stale admission on every re-preflight and rejecting missing or mismatched state.
|
||||
- [ ] [API-3] Replace direct-only code and project-skill contract tests with mixed/preset all-cell success, blocker, persistence, execution, and per-cell agy regressions.
|
||||
- [ ] [API-4] Run the focused and full benchmark regression suites and confirm secret-free append-only evidence remains valid.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Require exact all-cell preflight evidence
|
||||
|
||||
**Problem**
|
||||
|
||||
The collector skips every `execution_preset` (`attempts.py:390-410`), and both record validation and publication compare against `_direct_cells` (`attempts.py:514-516`, `563-568`, `632-651`). A mixed manifest can publish `status=ready` for only its direct subset, after which `run_slots` correctly refuses to execute because the record does not cover the matrix.
|
||||
|
||||
**Solution**
|
||||
|
||||
Validate capabilities and call `adapter.preflight(cell)` for every cell in canonical `manifest.matrix` order. Rename the private helper to `_preflight_cells` returning the full matrix, and use it for durable result cardinality, identity, order, and publication. Remove the direct-cell-empty special case from `preflight_manifest`; a valid manifest is already non-empty. Preserve canonical evidence validation and closed aggregate taxonomy.
|
||||
|
||||
Before (`scripts/agent_benchmark/attempts.py:386-423`):
|
||||
|
||||
```python
|
||||
"""Validate the full registry, then probe direct cells in manifest order."""
|
||||
...
|
||||
for cell in manifest.matrix:
|
||||
if cell.iop.route_kind != "direct":
|
||||
continue
|
||||
observation = adapters[cell.caller].preflight(cell)
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
"""Validate the full registry, then probe every cell in manifest order."""
|
||||
...
|
||||
for cell in manifest.matrix:
|
||||
observation = adapters[cell.caller].preflight(cell)
|
||||
...
|
||||
observations[cell.id] = observation
|
||||
```
|
||||
|
||||
Use the exact same full tuple on write and read so missing, extra, reordered, or foreign results fail closed. If any cell is blocked, publish the all-cell record, allocate zero attempts, and return the existing closed CLI status.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `scripts/agent_benchmark/attempts.py`: collect, write, read, and aggregate the exact full matrix.
|
||||
- [ ] `scripts/agent_benchmark/connectivity_integration_test.py`: update the public preflight/CLI contract tests for all-cell behavior.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
Write/replace tests. Assert preset-only ready records are durable, preset-only blockers are durable with zero cells directory, mixed records include both cells in matrix order, a ready mixed run invokes both exactly once, and missing/extra/reordered record payloads remain rejected.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
python3 -m unittest \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_cli_mixed_manifest_preflights_and_invokes_every_cell \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_preset_only_public_preflight_appends_exact_results \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_preset_blocker_appends_without_attempt_allocation
|
||||
```
|
||||
|
||||
Expected: all named all-cell tests pass.
|
||||
|
||||
### [API-2] Isolate live agy preflight state per cell
|
||||
|
||||
**Problem**
|
||||
|
||||
`_LiveAdapter._agy_preflight` is overwritten on each preflight (`live_iop.py:858`, `885`) and the last value is used for every agy invocation (`live_iop.py:935-949`). C03 and C07 in the same matrix can therefore use the wrong route observation/capability state.
|
||||
|
||||
**Solution**
|
||||
|
||||
Replace the scalar with `dict[str, AgyPreflightResult]`, importing the concrete exported dataclass from `agy_iop.py`. At the start of every cell preflight, remove that cell's prior agy state and admitted binding. Store both only after the current result is ready. At invocation, retrieve by `cell.id` once and pass that same object to `build_agy_invocation`, the invoker, and `observed_result`; reject absent state before building a spec.
|
||||
|
||||
Before (`scripts/agent_benchmark/live_iop.py:935-950`):
|
||||
|
||||
```python
|
||||
if self.caller == AGY_CALLER:
|
||||
if self._agy_preflight is None:
|
||||
raise LiveIopError("stream_incompatible")
|
||||
...
|
||||
result = self._invokers.agy(spec, parser, self._agy_preflight, ...)
|
||||
observed = parser.observed_result(self._agy_preflight.capability, result)
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
if self.caller == AGY_CALLER:
|
||||
agy_preflight = self._agy_preflights.get(cell.id)
|
||||
if agy_preflight is None:
|
||||
raise LiveIopError("stream_incompatible")
|
||||
...
|
||||
result = self._invokers.agy(spec, parser, agy_preflight, ...)
|
||||
observed = parser.observed_result(agy_preflight.capability, result)
|
||||
```
|
||||
|
||||
Do not persist raw endpoint or credential values; the new map remains process-local runtime state.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `scripts/agent_benchmark/live_iop.py`: clear, store, and retrieve agy preflight/admission by exact cell id.
|
||||
- [ ] `scripts/agent_benchmark/connectivity_integration_test.py`: add `test_live_agy_multiple_cells_consume_their_own_preflight_state` with two distinct route kinds/ids and mocked invokers.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
Write the named network-free regression. Preflight direct and preset agy cells with distinguishable observations, invoke in reverse preflight order, and assert each invocation receives its own object and binding. Add a re-preflight blocker case proving stale ready state cannot be invoked.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_live_agy_multiple_cells_consume_their_own_preflight_state
|
||||
```
|
||||
|
||||
Expected: both reverse-order isolation and stale-state rejection pass.
|
||||
|
||||
### [API-3] Replace the obsolete direct-only code and skill assertions
|
||||
|
||||
**Problem**
|
||||
|
||||
Three integration tests at `connectivity_integration_test.py:1009-1214` intentionally require preset omission. The project benchmark skill repeats that contract in its preflight procedure, validation, and safety sections, and `_assert_preflight_contract` plus its mutation at `skill_contract_test.py:256-271,786-794` enforce it. Leaving either surface unchanged would preserve or publish the exact S02 gap this packet closes.
|
||||
|
||||
**Solution**
|
||||
|
||||
Replace those integration tests with the all-cell cases named in API-1 and update exact adapter call/result assertions everywhere affected. Update the project skill so preflight, run, and resume describe a fresh append-only observation for every immutable matrix cell, while retaining closed blocker taxonomy and zero-attempt behavior. Rewrite `_assert_preflight_contract` and its mutation to reject any return of direct-only/preset-local language. Retain the existing three-caller direct case, registration-vs-implementation taxonomy, append-only corruption, missing-adapter, CLI redaction, and zero-attempt blocker assertions.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `scripts/agent_benchmark/connectivity_integration_test.py`: remove only obsolete local-only expectations and add full-matrix expectations.
|
||||
- [ ] `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md`: replace direct-only preflight wording with the exact all-cell durable contract.
|
||||
- [ ] `scripts/agent_benchmark/skill_contract_test.py`: require all-cell skill wording and reject restoration of preset-local semantics.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
Write tests, do not skip. All expected result ids must be derived from `manifest.matrix` and compared in stable order. No test may invent a ready observation without going through the typed fake or live adapter boundary. Skill tests must assert the affirmative all-cell contract and mutation-fail on the former direct-only statements.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.connectivity_integration_test scripts.agent_benchmark.skill_contract_test
|
||||
```
|
||||
|
||||
Expected: all network-free public connectivity tests pass.
|
||||
|
||||
### [API-4] Preserve append-only and redaction invariants
|
||||
|
||||
**Problem**
|
||||
|
||||
Expanding the record set changes durable cardinality and exercises more caller variants; regressions could appear outside the focused cases.
|
||||
|
||||
**Solution**
|
||||
|
||||
Run the complete manifest/attempt/connectivity/skill-contract suites and `git diff --check`. Record full output in the review stub.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `CODE_REVIEW-cloud-G09.md`: record implementation decisions, deviations, and actual command output.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
No further files. Existing corruption, retry/resume, measurement, web-validation, secret-redaction, and live-adapter tests provide the whole-suite oracle.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test scripts.agent_benchmark.skill_contract_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Expected: all tests report `OK`; whitespace check is silent.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Item |
|
||||
|------|------|
|
||||
| `scripts/agent_benchmark/attempts.py` | API-1 |
|
||||
| `scripts/agent_benchmark/live_iop.py` | API-2 |
|
||||
| `scripts/agent_benchmark/connectivity_integration_test.py` | API-1, API-2, API-3 |
|
||||
| `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md` | API-3 |
|
||||
| `scripts/agent_benchmark/skill_contract_test.py` | API-3 |
|
||||
| `agent-task/m-iop-one-shot-agent-model-comparison/02_route_preflight_contract/CODE_REVIEW-cloud-G09.md` | API-4 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
Run from `/config/workspace/iop-s0`; fresh output is required and cached output is not acceptable.
|
||||
|
||||
```bash
|
||||
python3 -m unittest \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_cli_mixed_manifest_preflights_and_invokes_every_cell \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_preset_only_public_preflight_appends_exact_results \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_preset_blocker_appends_without_attempt_allocation \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_live_agy_multiple_cells_consume_their_own_preflight_state
|
||||
python3 -m unittest scripts.agent_benchmark.connectivity_integration_test scripts.agent_benchmark.skill_contract_test
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test scripts.agent_benchmark.skill_contract_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Expected: every command exits 0, all benchmark suites report `OK`, and no secret/raw runtime values appear in durable-record assertions.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,136 @@
|
|||
<!-- task=m-iop-one-shot-agent-model-comparison/02_route_preflight_contract plan=0 tag=API milestone-task=route-readiness -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-12
|
||||
task=m-iop-one-shot-agent-model-comparison/02_route_preflight_contract, plan=0, tag=API
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files. Run the applicable verification commands directly and record fresh output in `Verification Results`; implementation-owned output is handoff evidence, not a substitute for reviewer verification. If implementation is present, repair missing or stale verification output instead of failing solely for insufficient recorded evidence. When verification exposes a defect, collect the necessary data, determine the exact root cause, and select one concrete fix before generating the follow-up plan; never delegate investigation or remedy selection to the worker.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals.
|
||||
2. Archive `CODE_REVIEW-cloud-G09.md` → `code_review_cloud_G09_0.log` and `PLAN-cloud-G09.md` → `plan_cloud_G09_0.log`.
|
||||
3. If PASS, write `complete.log` and move active task directory to `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/02_route_preflight_contract/`. If WARN/FAIL, fully write the next filesystem state required by the code-review skill.
|
||||
4. If PASS, preserve the first-line `milestone-task` metadata in `complete.log` and report it for runtime aggregation. Roadmap state evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check applicable `Review-Only Checklist` items at the final `.log` location before reporting.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|---------|
|
||||
| API-1 Require exact all-cell preflight evidence | [ ] |
|
||||
| API-2 Isolate live agy preflight state per cell | [ ] |
|
||||
| API-3 Replace obsolete direct-only assertions | [ ] |
|
||||
| API-4 Preserve append-only and redaction invariants | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] [API-1] Make collection and append-only validation require one canonical observation for every matrix cell, including preset-only and mixed manifests, with no attempt allocation on blockers.
|
||||
- [ ] [API-2] Store and consume agy live preflight state by cell id, clearing stale admission on every re-preflight and rejecting missing or mismatched state.
|
||||
- [ ] [API-3] Replace direct-only contract tests with mixed/preset all-cell success, blocker, persistence, execution, and per-cell agy regressions.
|
||||
- [ ] [API-4] Run the focused and full benchmark regression suites and confirm secret-free append-only evidence remains valid.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Run applicable required verification and record fresh command/output; repair reviewer-reconstructable evidence gaps instead of forwarding them to another plan.
|
||||
- [ ] For every Required/Suggested finding, record reviewer-collected `Evidence`, exact `Root Cause`, and one `Selected Fix` with affected files/symbols/tests and acceptance commands before creating a follow-up plan.
|
||||
- [ ] Archive active `CODE_REVIEW-cloud-G09.md` to `code_review_cloud_G09_0.log`.
|
||||
- [ ] Archive active `PLAN-cloud-G09.md` to `plan_cloud_G09_0.log`.
|
||||
- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write `complete.log` based on `agent-ops/skills/common/code-review/templates/complete-log-template.md` and leave no active `.md` files.
|
||||
- [ ] If PASS, move active task directory `agent-task/m-iop-one-shot-agent-model-comparison/02_route_preflight_contract/` to `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/02_route_preflight_contract/` and update this checklist at the final archive path.
|
||||
- [ ] If PASS, preserve and report `milestone-task` metadata for runtime aggregation, without modifying roadmap or directly calling `update-roadmap`.
|
||||
- [ ] If PASS for split work, remove empty active parent `agent-task/m-iop-one-shot-agent-model-comparison/` or verify it was kept due to remaining siblings/files.
|
||||
- [ ] If WARN/FAIL, write the next filesystem state matching code-review verdict and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record any deviations from the plan and the rationale here._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record key design decisions here._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Confirm the durable result set equals all `manifest.matrix` cells in canonical order on both write and read.
|
||||
- Confirm preset blockers publish exact closed taxonomy and allocate no attempt.
|
||||
- Confirm agy runtime state is keyed by cell id and a blocked re-preflight clears stale ready state.
|
||||
- Confirm no adapter-specific endpoint, token, or raw caller value enters evidence.
|
||||
|
||||
## Verification Results
|
||||
|
||||
Record actual stdout/stderr under each command. If a command changes, document the replacement and reason in `Deviations from Plan`.
|
||||
|
||||
### API-1/API-2 Focused Verification
|
||||
|
||||
```bash
|
||||
python3 -m unittest \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_cli_mixed_manifest_preflights_and_invokes_every_cell \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_preset_only_public_preflight_appends_exact_results \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_preset_blocker_appends_without_attempt_allocation \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_live_agy_multiple_cells_consume_their_own_preflight_state
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
### API-3 Verification
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.connectivity_integration_test
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
### API-4 and Final Verification
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these (archive, complete.log, and task-directory archive move are review-agent only) |
|
||||
| Archive Evidence Snapshot | Fixed at stub creation from plan when present | Implementing agent uses it as default prior-loop context; read only the specific archive files cited there when more detail is required |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results (section headings + commands) | Implementing agent, then review agent | Implementing agent records initial output; review agent reruns applicable commands and may fill, replace, or append fresh verified output before verdict. Implementing-agent command changes require a `Deviations from Plan` entry |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,276 @@
|
|||
<!-- task=m-iop-one-shot-agent-model-comparison/02_route_preflight_contract plan=0 tag=API milestone-task=route-readiness -->
|
||||
|
||||
# Plan - All-Cell Route Preflight Contract
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections in `CODE_REVIEW-cloud-G09.md` is the mandatory last implementation step. Run every verification command, record actual notes and stdout/stderr, keep both active files in place, and report ready for review; only the code-review skill may finalize or archive this task. If blocked, record only the exact blocker, attempted commands/output, and resume condition in the implementation-owned evidence fields. Do not ask the user, call user-input tools, create control-plane stop files, classify the next state, archive logs, or write `complete.log`.
|
||||
|
||||
## Background
|
||||
|
||||
The benchmark currently validates preset capability locally but deliberately omits preset cells from durable preflight. A mixed direct/preset manifest therefore records a misleading ready subset and allocates no attempts, while the live `agy` adapter retains only the last cell's preflight state. Route readiness for C01-C09 requires every declared cell to be observed and bound independently before any attempt starts.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `scripts/agent_benchmark/attempts.py`
|
||||
- `scripts/agent_benchmark/attempts_test.py`
|
||||
- `scripts/agent_benchmark/live_iop.py`
|
||||
- `scripts/agent_benchmark/connectivity_integration_test.py`
|
||||
- `scripts/agent_benchmark/manifest.py`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-manifest.example.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/iop-one-shot-agent-model-comparison.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/iop-one-shot-agent-model-comparison/SDD.md`
|
||||
- `agent-spec/testing/agent-comparison-benchmark.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- SDD: `agent-roadmap/sdd/knowledge-tool-optimization-extension/iop-one-shot-agent-model-comparison/SDD.md`, status `[승인됨]`, implementation lock released, no unresolved user review.
|
||||
- First-line scope: `milestone-task=route-readiness`.
|
||||
- Target scenario: S02 requires execution-day auth, model/preset, effort, stream/finish/idle checks for all C01-C09 or an exact blocker.
|
||||
- Evidence Map row: S02 requires a redacted C01-C09 preflight matrix with auth/route/effort/terminal evidence.
|
||||
- The checklist therefore replaces direct-only subset semantics with exact all-cell evidence and adds per-cell caller-state regressions. Live credentials remain the dependent manifest packet's external gate.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- No separate handoff was supplied. Repository-native evidence came from the attempts/live adapter paths and network-free integration tests above.
|
||||
- Fresh baseline at starting HEAD `b197e5db70637f87017a024a847e3e53fdc72e8b`: all three example manifests validated and the 178-test manifest/attempts/connectivity suite passed.
|
||||
- Current tests explicitly lock the old limitation: `test_cli_mixed_manifest_never_invokes_unobserved_preset_cells`, `test_generic_preset_cells_are_local_contract_only`, and `test_generic_preset_only_public_preflight_fails_closed_without_run_state` (`connectivity_integration_test.py:1009-1214`). These must be replaced, not preserved.
|
||||
- `collect_preflight_observations` skips non-direct cells (`attempts.py:386-423`), durable read/write validates against `_direct_cells` (`attempts.py:514-516`, `563-568`, `632-651`), and public preflight rejects an empty direct subset (`attempts.py:1869-1881`).
|
||||
- `_LiveAdapter` stores one `_agy_preflight` value for every cell (`live_iop.py:843-902`) and reuses it at invocation (`live_iop.py:935-950`), so multiple agy routes can cross-bind.
|
||||
- This packet uses only network-free adapters and mocks. Fresh output is required; cached output is not accepted.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
- Existing tests cover direct cells and intentionally assert that preset cells are omitted.
|
||||
- No mixed manifest test expects both direct and preset results in the same append-only record and then executes both.
|
||||
- No preset-only public preflight test expects a durable ready or blocker record.
|
||||
- No live-adapter test preflights two agy cells and proves each invocation receives its own preflight capability/runtime binding.
|
||||
- Existing corruption and secret-redaction coverage is broad and must continue to pass after the result-set invariant changes.
|
||||
|
||||
### Symbol References
|
||||
|
||||
- Rename private `RunStore._direct_cells` to `_preflight_cells` or remove the helper. Its only references are `attempts.py:563` and `attempts.py:632`; update both together.
|
||||
- Rename private `_LiveAdapter._agy_preflight` to `_agy_preflights`; all references are `live_iop.py:858`, `885`, `893-894`, and `936-949`.
|
||||
- No public import or CLI symbol is renamed.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
- This is split child `02_route_preflight_contract`. Its stable invariant is: the preflight result set equals the entire immutable matrix, and every admitted live invocation consumes the preflight state for its own cell id.
|
||||
- PASS evidence is all-cell append/read corruption coverage, mixed/preset-only CLI coverage, and a two-cell agy state test.
|
||||
- It is independent of `01_execution_order_contract`; it consumes `manifest.matrix` order without owning that order. `03+01,02_locked_benchmark_manifest` depends on both.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
- Do not add the final nine-cell manifest or an execution-order field here.
|
||||
- Do not change caller wire formats, IOP contracts, provider config, credentials, retry policy, attempt allocation semantics, or scoring.
|
||||
- Do not synthesize readiness for missing routes. Every preset must pass the same live config/catalog/capability checks and retain exact `registration_required` or `implementation_gap` taxonomy.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- `evaluation_mode=first-pass`; `finalizer=finalize-task-policy.sh`, `finalizer_mode=pair`.
|
||||
- Build closures: scope/context/verification/evidence/ownership/decision all closed. Scores `2+2+2+1+2=G09`; base/final route `grade-boundary`, lane `cloud`, catalog `worker/cloud/G09`, filename `PLAN-cloud-G09.md`.
|
||||
- Review closures: all closed. Scores `2+2+2+1+2=G09`; route `official-review`, lane `cloud`, catalog `review/cloud/G09`, filename `CODE_REVIEW-cloud-G09.md`.
|
||||
- `large_indivisible_context=false`; positive loop risks: `temporal_state`, `boundary_contract`, `structured_interpretation`, `variant_product` (`count=4`, risk boundary matched but does not replace grade basis); `review_rework_count=0`; `evidence_integrity_failure=false`; no capability gap.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] [API-1] Make collection and append-only validation require one canonical observation for every matrix cell, including preset-only and mixed manifests, with no attempt allocation on blockers.
|
||||
- [ ] [API-2] Store and consume agy live preflight state by cell id, clearing stale admission on every re-preflight and rejecting missing or mismatched state.
|
||||
- [ ] [API-3] Replace direct-only contract tests with mixed/preset all-cell success, blocker, persistence, execution, and per-cell agy regressions.
|
||||
- [ ] [API-4] Run the focused and full benchmark regression suites and confirm secret-free append-only evidence remains valid.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Require exact all-cell preflight evidence
|
||||
|
||||
**Problem**
|
||||
|
||||
The collector skips every `execution_preset` (`attempts.py:390-410`), and both record validation and publication compare against `_direct_cells` (`attempts.py:514-516`, `563-568`, `632-651`). A mixed manifest can publish `status=ready` for only its direct subset, after which `run_slots` correctly refuses to execute because the record does not cover the matrix.
|
||||
|
||||
**Solution**
|
||||
|
||||
Validate capabilities and call `adapter.preflight(cell)` for every cell in canonical `manifest.matrix` order. Rename the private helper to `_preflight_cells` returning the full matrix, and use it for durable result cardinality, identity, order, and publication. Remove the direct-cell-empty special case from `preflight_manifest`; a valid manifest is already non-empty. Preserve canonical evidence validation and closed aggregate taxonomy.
|
||||
|
||||
Before (`scripts/agent_benchmark/attempts.py:386-423`):
|
||||
|
||||
```python
|
||||
"""Validate the full registry, then probe direct cells in manifest order."""
|
||||
...
|
||||
for cell in manifest.matrix:
|
||||
if cell.iop.route_kind != "direct":
|
||||
continue
|
||||
observation = adapters[cell.caller].preflight(cell)
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
"""Validate the full registry, then probe every cell in manifest order."""
|
||||
...
|
||||
for cell in manifest.matrix:
|
||||
observation = adapters[cell.caller].preflight(cell)
|
||||
...
|
||||
observations[cell.id] = observation
|
||||
```
|
||||
|
||||
Use the exact same full tuple on write and read so missing, extra, reordered, or foreign results fail closed. If any cell is blocked, publish the all-cell record, allocate zero attempts, and return the existing closed CLI status.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `scripts/agent_benchmark/attempts.py`: collect, write, read, and aggregate the exact full matrix.
|
||||
- [ ] `scripts/agent_benchmark/connectivity_integration_test.py`: update the public preflight/CLI contract tests for all-cell behavior.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
Write/replace tests. Assert preset-only ready records are durable, preset-only blockers are durable with zero cells directory, mixed records include both cells in matrix order, a ready mixed run invokes both exactly once, and missing/extra/reordered record payloads remain rejected.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
python3 -m unittest \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_cli_mixed_manifest_preflights_and_invokes_every_cell \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_preset_only_public_preflight_appends_exact_results \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_preset_blocker_appends_without_attempt_allocation
|
||||
```
|
||||
|
||||
Expected: all named all-cell tests pass.
|
||||
|
||||
### [API-2] Isolate live agy preflight state per cell
|
||||
|
||||
**Problem**
|
||||
|
||||
`_LiveAdapter._agy_preflight` is overwritten on each preflight (`live_iop.py:858`, `885`) and the last value is used for every agy invocation (`live_iop.py:935-949`). C03 and C07 in the same matrix can therefore use the wrong route observation/capability state.
|
||||
|
||||
**Solution**
|
||||
|
||||
Replace the scalar with `dict[str, AgyIopPreflight]` (use the concrete existing return type if exported; otherwise retain `Any` only at the value boundary). At the start of every cell preflight, remove that cell's prior agy state and admitted binding. Store both only after the current result is ready. At invocation, retrieve by `cell.id` once and pass that same object to `build_agy_invocation`, the invoker, and `observed_result`; reject absent state before building a spec.
|
||||
|
||||
Before (`scripts/agent_benchmark/live_iop.py:935-950`):
|
||||
|
||||
```python
|
||||
if self.caller == AGY_CALLER:
|
||||
if self._agy_preflight is None:
|
||||
raise LiveIopError("stream_incompatible")
|
||||
...
|
||||
result = self._invokers.agy(spec, parser, self._agy_preflight, ...)
|
||||
observed = parser.observed_result(self._agy_preflight.capability, result)
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
if self.caller == AGY_CALLER:
|
||||
agy_preflight = self._agy_preflights.get(cell.id)
|
||||
if agy_preflight is None:
|
||||
raise LiveIopError("stream_incompatible")
|
||||
...
|
||||
result = self._invokers.agy(spec, parser, agy_preflight, ...)
|
||||
observed = parser.observed_result(agy_preflight.capability, result)
|
||||
```
|
||||
|
||||
Do not persist raw endpoint or credential values; the new map remains process-local runtime state.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `scripts/agent_benchmark/live_iop.py`: clear, store, and retrieve agy preflight/admission by exact cell id.
|
||||
- [ ] `scripts/agent_benchmark/connectivity_integration_test.py`: add `test_live_agy_multiple_cells_consume_their_own_preflight_state` with two distinct route kinds/ids and mocked invokers.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
Write the named network-free regression. Preflight direct and preset agy cells with distinguishable observations, invoke in reverse preflight order, and assert each invocation receives its own object and binding. Add a re-preflight blocker case proving stale ready state cannot be invoked.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_live_agy_multiple_cells_consume_their_own_preflight_state
|
||||
```
|
||||
|
||||
Expected: both reverse-order isolation and stale-state rejection pass.
|
||||
|
||||
### [API-3] Replace the obsolete direct-only assertions
|
||||
|
||||
**Problem**
|
||||
|
||||
Three integration tests at `connectivity_integration_test.py:1009-1214` intentionally require preset omission. Leaving them unchanged would preserve the exact S02 gap this packet closes.
|
||||
|
||||
**Solution**
|
||||
|
||||
Replace those tests with the all-cell cases named in API-1 and update exact adapter call/result assertions everywhere affected. Retain the existing three-caller direct case, registration-vs-implementation taxonomy, append-only corruption, missing-adapter, CLI redaction, and zero-attempt blocker assertions.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `scripts/agent_benchmark/connectivity_integration_test.py`: remove only obsolete local-only expectations and add full-matrix expectations.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
Write tests, do not skip. All expected result ids must be derived from `manifest.matrix` and compared in stable order. No test may invent a ready observation without going through the typed fake or live adapter boundary.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.connectivity_integration_test
|
||||
```
|
||||
|
||||
Expected: all network-free public connectivity tests pass.
|
||||
|
||||
### [API-4] Preserve append-only and redaction invariants
|
||||
|
||||
**Problem**
|
||||
|
||||
Expanding the record set changes durable cardinality and exercises more caller variants; regressions could appear outside the focused cases.
|
||||
|
||||
**Solution**
|
||||
|
||||
Run the complete manifest/attempt/connectivity suites and `git diff --check`. Record full output in the review stub.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `CODE_REVIEW-cloud-G09.md`: record implementation decisions, deviations, and actual command output.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
No further files. Existing corruption, retry/resume, measurement, web-validation, secret-redaction, and live-adapter tests provide the whole-suite oracle.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Expected: all tests report `OK`; whitespace check is silent.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Item |
|
||||
|------|------|
|
||||
| `scripts/agent_benchmark/attempts.py` | API-1 |
|
||||
| `scripts/agent_benchmark/live_iop.py` | API-2 |
|
||||
| `scripts/agent_benchmark/connectivity_integration_test.py` | API-1, API-2, API-3 |
|
||||
| `agent-task/m-iop-one-shot-agent-model-comparison/02_route_preflight_contract/CODE_REVIEW-cloud-G09.md` | API-4 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
Run from `/config/workspace/iop-s0`; fresh output is required and cached output is not acceptable.
|
||||
|
||||
```bash
|
||||
python3 -m unittest \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_cli_mixed_manifest_preflights_and_invokes_every_cell \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_preset_only_public_preflight_appends_exact_results \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_preset_blocker_appends_without_attempt_allocation \
|
||||
scripts.agent_benchmark.connectivity_integration_test.ConnectivityIntegrationTest.test_live_agy_multiple_cells_consume_their_own_preflight_state
|
||||
python3 -m unittest scripts.agent_benchmark.connectivity_integration_test
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Expected: every command exits 0, all benchmark suites report `OK`, and no secret/raw runtime values appear in durable-record assertions.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,163 @@
|
|||
<!-- task=m-iop-one-shot-agent-model-comparison/03+01_rubric_version_contract plan=1 tag=API milestone-task=fixture-lock -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-12
|
||||
task=m-iop-one-shot-agent-model-comparison/03+01_rubric_version_contract, plan=1, tag=API
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior plan: `agent-task/m-iop-one-shot-agent-model-comparison/03+01_rubric_version_contract/plan_local_G08_0.log`.
|
||||
- Prior review: `agent-task/m-iop-one-shot-agent-model-comparison/03+01_rubric_version_contract/code_review_cloud_G08_0.log`.
|
||||
- The prior pair contains no implementation evidence or official verdict. Epic self-review found a semantic contradiction: Scope Rationale prohibited a project-skill edit while API-3 and the file summary required that same edit, creating ownership overlap with the route-preflight packet.
|
||||
- Replan keeps this packet on the dormant additive rubric implementation and dual-version code tests. The dependent final-manifest packet owns the operational skill transition when it selects the new rubric. Baseline manifest validation and the 178-test manifest/attempt/connectivity suite were green at `b197e5db70637f87017a024a847e3e53fdc72e8b`.
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files. Run the applicable verification commands directly and record fresh output in `Verification Results`; implementation-owned output is handoff evidence, not a substitute for reviewer verification. If implementation is present, repair missing or stale verification output instead of failing solely for insufficient recorded evidence. When verification exposes a defect, collect the necessary data, determine the exact root cause, and select one concrete fix before generating the follow-up plan; never delegate investigation or remedy selection to the worker.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals.
|
||||
2. Archive `CODE_REVIEW-cloud-G08.md` → `code_review_cloud_G08_1.log` and `PLAN-local-G08.md` → `plan_local_G08_1.log`.
|
||||
3. If PASS, write `complete.log` and move active task directory to `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/03+01_rubric_version_contract/`. If WARN/FAIL, fully write the next filesystem state required by the code-review skill.
|
||||
4. If PASS, preserve the first-line `milestone-task` metadata in `complete.log` and report it for runtime aggregation. Roadmap state evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check applicable `Review-Only Checklist` items at the final `.log` location before reporting.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|---------|
|
||||
| API-1 Add a closed dual-version rubric catalog | [ ] |
|
||||
| API-2 Bind prompt and durable validation to manifest version | [ ] |
|
||||
| API-3 Preserve historical compatibility and exact new semantics | [ ] |
|
||||
| API-4 Run the complete benchmark regression | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] [API-1] Add an immutable `one-shot-agent-comparison-v1` seven-category rubric while preserving exact `landing-quality-v1` loader/schema/worksheet behavior.
|
||||
- [ ] [API-2] Make evaluator prompt generation and worksheet loading use the manifest-selected rubric version and reject cross-version output.
|
||||
- [ ] [API-3] Add dual-version manifest, rubric, and scoring regressions for exact category order, maxima, total 100, prompt text, and historical compatibility without changing operational skill wording before a tracked manifest selects the new rubric.
|
||||
- [ ] [API-4] Run all rubric/scoring and benchmark regression suites plus shipped manifest validation.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Run applicable required verification and record fresh command/output; repair reviewer-reconstructable evidence gaps instead of forwarding them to another plan.
|
||||
- [ ] For every Required/Suggested finding, record reviewer-collected `Evidence`, exact `Root Cause`, and one `Selected Fix` with affected files/symbols/tests and acceptance commands before creating a follow-up plan.
|
||||
- [ ] Archive active `CODE_REVIEW-cloud-G08.md` to `code_review_cloud_G08_1.log`.
|
||||
- [ ] Archive active `PLAN-local-G08.md` to `plan_local_G08_1.log`.
|
||||
- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write `complete.log` based on `agent-ops/skills/common/code-review/templates/complete-log-template.md` and leave no active `.md` files.
|
||||
- [ ] If PASS, move active task directory `agent-task/m-iop-one-shot-agent-model-comparison/03+01_rubric_version_contract/` to `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/03+01_rubric_version_contract/` and update this checklist at the final archive path.
|
||||
- [ ] If PASS, preserve and report `milestone-task` metadata for runtime aggregation, without modifying roadmap or directly calling `update-roadmap`.
|
||||
- [ ] If PASS for split work, remove empty active parent `agent-task/m-iop-one-shot-agent-model-comparison/` or verify it was kept due to remaining siblings/files.
|
||||
- [ ] If WARN/FAIL, write the next filesystem state matching code-review verdict and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record any deviations from the plan and the rationale here._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record key design decisions here._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Confirm the legacy version/category tuple remains exact and reloadable.
|
||||
- Confirm the new seven categories and maxima exactly match the approved SDD and total 100.
|
||||
- Confirm prompt and both worksheet load paths use `manifest.rubric_version`.
|
||||
- Confirm unknown and cross-version worksheets fail before scored result publication.
|
||||
|
||||
## Verification Results
|
||||
|
||||
Record actual stdout/stderr under each command. If a command changes, document the replacement and reason in `Deviations from Plan`.
|
||||
|
||||
### Dependency Verification
|
||||
|
||||
```bash
|
||||
python3 - <<'PY'
|
||||
from pathlib import Path
|
||||
active = Path("agent-task/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/complete.log")
|
||||
archived = sorted(Path("agent-task/archive").glob("*/*/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/complete.log"))
|
||||
candidates = [path for path in (active, *archived) if path.is_file()]
|
||||
assert len(candidates) == 1, candidates
|
||||
print(f"ok: predecessor complete {candidates[0]}")
|
||||
PY
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
### API-1 Verification
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.manifest_test
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
### API-2 Verification
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.scoring_test.ScoringTest.test_manifest_selected_rubric_drives_prompt_and_worksheet_validation
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
### API-3 Verification
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.scoring_test scripts.agent_benchmark.manifest_test
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
### API-4 and Final Verification
|
||||
|
||||
```bash
|
||||
for manifest in scripts/fixtures/agent-comparison-benchmark-manifest.example.json scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json; do python3 scripts/agent_comparison_benchmark.py validate --manifest "$manifest"; done
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.scoring_test scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these (archive, complete.log, and task-directory archive move are review-agent only) |
|
||||
| Archive Evidence Snapshot | Fixed at stub creation from plan when present | Implementing agent uses it as default prior-loop context; read only the specific archive files cited there when more detail is required |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results (section headings + commands) | Implementing agent, then review agent | Implementing agent records initial output; review agent reruns applicable commands and may fill, replace, or append fresh verified output before verdict. Implementing-agent command changes require a `Deviations from Plan` entry |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,307 @@
|
|||
<!-- task=m-iop-one-shot-agent-model-comparison/03+01_rubric_version_contract plan=1 tag=API milestone-task=fixture-lock -->
|
||||
|
||||
# Plan - Versioned One-Shot Benchmark Rubric Contract
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Do not start until exactly one active or archived `01_execution_order_contract/complete.log` is resolved by the dependency command below. Filling the implementation-owned sections in `CODE_REVIEW-cloud-G08.md` is the mandatory last implementation step. Run every verification command, record actual notes and stdout/stderr, keep both active files in place, and report ready for review; only the code-review skill may finalize or archive this task. If blocked, record only the exact blocker, attempted commands/output, and resume condition in the implementation-owned evidence fields. Do not ask the user, call user-input tools, create control-plane stop files, classify the next state, archive logs, or write `complete.log`.
|
||||
|
||||
## Background
|
||||
|
||||
The approved one-shot comparison SDD fixes seven rubric categories totaling 100, while the existing `landing-quality-v1` worksheet contains five different categories. Reusing that version name would mutate historical meaning, so this packet preserves legacy runs and adds a manifest-selected version for the approved benchmark rubric.
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior plan: `agent-task/m-iop-one-shot-agent-model-comparison/03+01_rubric_version_contract/plan_local_G08_0.log`.
|
||||
- Prior review: `agent-task/m-iop-one-shot-agent-model-comparison/03+01_rubric_version_contract/code_review_cloud_G08_0.log`.
|
||||
- The prior pair contains no implementation evidence or official verdict. Epic self-review found a semantic contradiction: Scope Rationale prohibited a project-skill edit while API-3 and the file summary required that same edit, creating ownership overlap with the route-preflight packet.
|
||||
- Replan keeps this packet on the dormant additive rubric implementation and dual-version code tests. The dependent final-manifest packet owns the operational skill transition when it selects the new rubric. Baseline manifest validation and the 178-test manifest/attempt/connectivity suite were green at `b197e5db70637f87017a024a847e3e53fdc72e8b`.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `scripts/agent_benchmark/manifest.py`
|
||||
- `scripts/agent_benchmark/manifest_test.py`
|
||||
- `scripts/agent_benchmark/rubric.py`
|
||||
- `scripts/agent_benchmark/rubric_test.py`
|
||||
- `scripts/agent_benchmark/scoring.py`
|
||||
- `scripts/agent_benchmark/scoring_test.py`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-manifest.example.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json`
|
||||
- `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/iop-one-shot-agent-model-comparison.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/iop-one-shot-agent-model-comparison/SDD.md`
|
||||
- `agent-spec/testing/agent-comparison-benchmark.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- SDD: `agent-roadmap/sdd/knowledge-tool-optimization-extension/iop-one-shot-agent-model-comparison/SDD.md`, status `[승인됨]`, implementation lock released, no unresolved user review.
|
||||
- First-line scope: `milestone-task=fixture-lock`.
|
||||
- Target scenario: S01 requires the same fixture, viewport, rubric checksum/version for every cell. The approved rubric at SDD lines 90-91 is 요구사항 충족 25, 시각 완성도 25, 반응형·접근성 15, 이미지 활용·디테일 10, 동작 안정성 10, 코드 품질 10, 자체 검증 완결성 5.
|
||||
- Evidence Map row: S01 requires fixture prompt/assets/workspace/rubric digest evidence.
|
||||
- The checklist therefore creates a distinct immutable rubric version, retains legacy validation, and binds prompt/output validation to the manifest-selected version. The final manifest packet supplies the concrete version/digest evidence.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- No separate handoff was supplied. Repository-native evidence came from the complete manifest, rubric, scoring, and test files above.
|
||||
- Fresh baseline at starting HEAD `b197e5db70637f87017a024a847e3e53fdc72e8b`: all example manifests validated and the 178-test manifest/attempt/connectivity suite passed. `rubric_test` and `scoring_test` are additional required fresh suites for this packet.
|
||||
- Current `RUBRIC_CATEGORIES` is five entries totaling 100 (`rubric.py:21-27`). `validate_worksheet` accepts only global `RUBRIC_VERSION` (`rubric.py:82-126`), and `_prompt` hardcodes those categories and `landing-quality-v1` (`scoring.py:1207-1222`).
|
||||
- The safe compatibility rule is additive: keep `landing-quality-v1` and its five-category table readable; add `one-shot-agent-comparison-v1` with the approved seven-category table; reject unknown and cross-version worksheets.
|
||||
- Fresh output is required; cached output is not accepted. No external runner is required.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
- Existing rubric tests validate only the five-category legacy table.
|
||||
- Existing scoring tests generate only legacy worksheets and cannot prove prompt/worksheet selection from the manifest.
|
||||
- Existing manifest/schema tests accept only one rubric constant.
|
||||
- No test rejects a valid legacy worksheet when the allocation requires the new rubric, or vice versa.
|
||||
|
||||
### Symbol References
|
||||
|
||||
- `manifest.RUBRIC_VERSION` is referenced only by `manifest.py:711` and `rubric.py:16,83,126`. Replace internal single-version validation with a closed version catalog while preserving a legacy alias if needed for compatibility.
|
||||
- `rubric.RUBRIC_CATEGORIES` is imported by `scoring.py:44-49`, `scoring_test.py:56`, and `rubric_test.py:10`. Preserve it as the legacy tuple for existing callers, add a version lookup, and migrate production scoring to the lookup.
|
||||
- `_prompt` is private and called at `scoring.py:1879`; change both definition and call together.
|
||||
- `load_worksheet` call sites are `scoring.py:1641` and `scoring.py:1972`; both must pass the manifest-selected expected version.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
- This is dependent child `03+01_rubric_version_contract`. Its stable invariant is: rubric version selects exactly one immutable ordered category table from manifest load through evaluator prompt, worksheet validation, durable result, and historical reload.
|
||||
- PASS evidence is dual-version loader/rubric/scoring tests plus full regression.
|
||||
- Predecessor index 01 is currently active at `agent-task/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/` but has no `complete.log`; implementation is dependency-blocked until that exact file exists. The dependency is required because both packets modify `manifest.py`, `manifest_test.py`, and the schema.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
- Do not delete, rename, or reinterpret `landing-quality-v1`; old manifests and immutable run snapshots must still load and score under the old table.
|
||||
- Do not change automatic web gate eligibility, evaluator identity blinding, score persistence, caller bindings, fixture bytes, or report layout.
|
||||
- Do not update the project benchmark skill or living spec in this packet. The currently tracked manifests still select the legacy rubric; dependent packet 04 owns the atomic skill transition when it adds the first tracked new-version manifest.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- `evaluation_mode=isolated-reassessment`; `finalizer=finalize-task-policy.sh`, `finalizer_mode=pair`.
|
||||
- Build closures: scope/context/verification/evidence/ownership/decision all closed. Scores `2+1+2+1+2=G08`; base/final route `local-fit`, lane `local`, catalog `worker/local/G08`, filename `PLAN-local-G08.md`.
|
||||
- Review closures: all closed. Scores `2+1+2+1+2=G08`; route `official-review`, lane `cloud`, catalog `review/cloud/G08`, filename `CODE_REVIEW-cloud-G08.md`.
|
||||
- `large_indivisible_context=false`; positive loop risks: `boundary_contract`, `structured_interpretation`, `variant_product` (`count=3`); `review_rework_count=0`; `evidence_integrity_failure=false`; no capability gap.
|
||||
|
||||
## Dependencies and Execution Order
|
||||
|
||||
1. Resolve exactly one predecessor `complete.log`: the active sibling path or one matching `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/complete.log`.
|
||||
2. Implement this packet against the reviewed 01 source. Do not infer completion from an active PLAN or review stub.
|
||||
3. On PASS, `04+02,03_locked_benchmark_manifest` may proceed only after both index 02 and this index 03 have `complete.log`.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] [API-1] Add an immutable `one-shot-agent-comparison-v1` seven-category rubric while preserving exact `landing-quality-v1` loader/schema/worksheet behavior.
|
||||
- [ ] [API-2] Make evaluator prompt generation and worksheet loading use the manifest-selected rubric version and reject cross-version output.
|
||||
- [ ] [API-3] Add dual-version manifest, rubric, and scoring regressions for exact category order, maxima, total 100, prompt text, and historical compatibility without changing operational skill wording before a tracked manifest selects the new rubric.
|
||||
- [ ] [API-4] Run all rubric/scoring and benchmark regression suites plus shipped manifest validation.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Add a closed dual-version rubric catalog
|
||||
|
||||
**Problem**
|
||||
|
||||
`manifest.py:30` exposes one allowed rubric version, and `rubric.py:21-27` gives that version a five-category table that differs from the approved SDD. Changing the existing tuple in place would silently rewrite historical `landing-quality-v1` meaning.
|
||||
|
||||
**Solution**
|
||||
|
||||
Keep the legacy string and tuple unchanged. Add `ONE_SHOT_RUBRIC_VERSION = "one-shot-agent-comparison-v1"`, a closed allowed-version tuple in `manifest.py`, and an ordered category mapping in `rubric.py`. The new table is exactly:
|
||||
|
||||
```python
|
||||
ONE_SHOT_RUBRIC_CATEGORIES = (
|
||||
("requirements_fidelity", 25),
|
||||
("visual_completeness", 25),
|
||||
("responsive_accessibility", 15),
|
||||
("image_detail_usage", 10),
|
||||
("behavior_stability", 10),
|
||||
("code_quality", 10),
|
||||
("self_verification", 5),
|
||||
)
|
||||
```
|
||||
|
||||
Add `rubric_categories(version)` that fails closed for unknown versions. Update the manifest schema from a single `const` to an exact two-value enum. `validate_worksheet` must read the worksheet version, optionally require an `expected_version`, and validate against only that version's ordered table.
|
||||
|
||||
Before (`scripts/agent_benchmark/rubric.py:21-27`):
|
||||
|
||||
```python
|
||||
RUBRIC_CATEGORIES = (
|
||||
("task_fidelity", 25),
|
||||
("visual_hierarchy", 25),
|
||||
("responsive_composition", 20),
|
||||
("typography_readability", 15),
|
||||
("polish_consistency", 15),
|
||||
)
|
||||
```
|
||||
|
||||
After: retain that exact tuple as the legacy value and add the new tuple/mapping without mutation.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `scripts/agent_benchmark/manifest.py`: accept the closed legacy/new rubric version catalog.
|
||||
- [ ] `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json`: mirror the exact two-version enum.
|
||||
- [ ] `scripts/agent_benchmark/rubric.py`: provide versioned category lookup and expected-version worksheet validation.
|
||||
- [ ] `scripts/agent_benchmark/rubric_test.py`: verify exact tables, totals, ordering, unknown versions, and cross-version rejection.
|
||||
- [ ] `scripts/agent_benchmark/manifest_test.py`: verify loader/schema parity for both allowed versions and an unknown version.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
Write tests. Preserve every existing legacy assertion. Add `test_one_shot_rubric_exact_categories_and_total_are_accepted`, `test_cross_version_worksheet_is_rejected`, and manifest/schema two-version parity. Assert both category sums equal exactly 100.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.manifest_test
|
||||
```
|
||||
|
||||
Expected: both versions pass their exact table; malformed, unknown, reordered, and cross-version cases fail closed.
|
||||
|
||||
### [API-2] Bind prompt and durable score validation to the manifest version
|
||||
|
||||
**Problem**
|
||||
|
||||
`_prompt` uses global categories and a literal version (`scoring.py:1207-1222`), while both `load_worksheet` calls (`scoring.py:1641`, `1972`) validate only the global legacy version. A new manifest version could therefore ask for or accept the wrong worksheet.
|
||||
|
||||
**Solution**
|
||||
|
||||
Change `_prompt` to accept `rubric_version`, look up its exact table, render that version string, and call it with `manifest.rubric_version`. Extend `load_worksheet`/`validate_worksheet` with `expected_version` and pass `manifest.rubric_version` at both scoring call sites. Keep allocation/result binding to the manifest version and reject a worksheet whose internal version differs before publication.
|
||||
|
||||
Before (`scripts/agent_benchmark/scoring.py:1207-1222`):
|
||||
|
||||
```python
|
||||
def _prompt(blind: BlindWorkspace) -> bytes:
|
||||
categories = ", ".join(
|
||||
f"{ident} ({maximum})" for ident, maximum in RUBRIC_CATEGORIES
|
||||
)
|
||||
...
|
||||
"landing-quality-v1, integer scores within each maximum ..."
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
def _prompt(blind: BlindWorkspace, rubric_version: str) -> bytes:
|
||||
categories = ", ".join(
|
||||
f"{ident} ({maximum})"
|
||||
for ident, maximum in rubric_categories(rubric_version)
|
||||
)
|
||||
...
|
||||
f"{rubric_version}, integer scores within each maximum ..."
|
||||
```
|
||||
|
||||
The `blind` parameter may remain for interface stability even though it is not interpolated. Do not expose producer identity.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `scripts/agent_benchmark/scoring.py`: select prompt categories/version and worksheet validation from `manifest.rubric_version`.
|
||||
- [ ] `scripts/agent_benchmark/scoring_test.py`: add a new-version scoring success and cross-version failure test.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
Write `test_manifest_selected_rubric_drives_prompt_and_worksheet_validation`. Build a new-version manifest, emit the exact seven-category worksheet, and assert success plus prompt category/version content. On a fresh score attempt, emit a structurally valid legacy worksheet and assert `invalid_worksheet` with no scored worksheet publication.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.scoring_test.ScoringTest.test_manifest_selected_rubric_drives_prompt_and_worksheet_validation
|
||||
```
|
||||
|
||||
Expected: selected-version success and cross-version rejection both pass.
|
||||
|
||||
### [API-3] Preserve historical compatibility and exact new semantics
|
||||
|
||||
**Problem**
|
||||
|
||||
An additive version is safe only if old manifests, prompts, worksheets, and immutable score reloads retain byte-level meaning.
|
||||
|
||||
**Solution**
|
||||
|
||||
Keep existing example manifests on `landing-quality-v1`. Make test helpers accept a version instead of replacing their legacy default. Add explicit assertions that legacy canonical worksheet bytes and prompts remain unchanged, while the new version has seven ordered categories and total 100. Do not change the project benchmark skill in this dependency packet: all currently tracked manifests still select the legacy version, and the dependent final-manifest packet owns the atomic operational transition to manifest-selected wording.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `scripts/agent_benchmark/rubric_test.py`: dual-version canonical and boundary assertions.
|
||||
- [ ] `scripts/agent_benchmark/scoring_test.py`: legacy default plus new-version selected behavior.
|
||||
- [ ] `scripts/agent_benchmark/manifest_test.py`: both versions and shipped examples.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
Write tests; do not mass-replace legacy version literals in unrelated fixtures. The new final manifest is added only by the dependent packet.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.scoring_test scripts.agent_benchmark.manifest_test
|
||||
```
|
||||
|
||||
Expected: all legacy and new rubric tests report `OK`.
|
||||
|
||||
### [API-4] Run the complete benchmark regression
|
||||
|
||||
**Problem**
|
||||
|
||||
Rubric validation is used by scoring recovery and durable result reads, so focused tests are not enough.
|
||||
|
||||
**Solution**
|
||||
|
||||
Run all rubric/scoring and the existing benchmark suites, validate shipped legacy examples, and check whitespace.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `CODE_REVIEW-cloud-G08.md`: record exact decisions, deviations, and all actual command output.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
No further files. Existing recovery, redaction, scoring retry, manifest digest, and connectivity suites are the full oracle.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
for manifest in scripts/fixtures/agent-comparison-benchmark-manifest.example.json scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json; do python3 scripts/agent_comparison_benchmark.py validate --manifest "$manifest"; done
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.scoring_test scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Expected: three legacy manifests validate, all tests report `OK`, and whitespace check is silent.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Item |
|
||||
|------|------|
|
||||
| `scripts/agent_benchmark/manifest.py` | API-1 |
|
||||
| `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json` | API-1 |
|
||||
| `scripts/agent_benchmark/rubric.py` | API-1 |
|
||||
| `scripts/agent_benchmark/rubric_test.py` | API-1, API-3 |
|
||||
| `scripts/agent_benchmark/manifest_test.py` | API-1, API-3 |
|
||||
| `scripts/agent_benchmark/scoring.py` | API-2 |
|
||||
| `scripts/agent_benchmark/scoring_test.py` | API-2, API-3 |
|
||||
| `agent-task/m-iop-one-shot-agent-model-comparison/03+01_rubric_version_contract/CODE_REVIEW-cloud-G08.md` | API-4 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
Run from `/config/workspace/iop-s0` after predecessor 01 is complete; fresh output is required and cached output is not acceptable.
|
||||
|
||||
```bash
|
||||
python3 - <<'PY'
|
||||
from pathlib import Path
|
||||
active = Path("agent-task/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/complete.log")
|
||||
archived = sorted(Path("agent-task/archive").glob("*/*/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/complete.log"))
|
||||
candidates = [path for path in (active, *archived) if path.is_file()]
|
||||
assert len(candidates) == 1, candidates
|
||||
print(f"ok: predecessor complete {candidates[0]}")
|
||||
PY
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.manifest_test
|
||||
python3 -m unittest scripts.agent_benchmark.scoring_test.ScoringTest.test_manifest_selected_rubric_drives_prompt_and_worksheet_validation
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.scoring_test scripts.agent_benchmark.manifest_test
|
||||
for manifest in scripts/fixtures/agent-comparison-benchmark-manifest.example.json scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json; do python3 scripts/agent_comparison_benchmark.py validate --manifest "$manifest"; done
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.scoring_test scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Expected: dependency check and every command exit 0, legacy examples validate, all suites report `OK`, and `git diff --check` is silent.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,157 @@
|
|||
<!-- task=m-iop-one-shot-agent-model-comparison/03+01_rubric_version_contract plan=0 tag=API milestone-task=fixture-lock -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-12
|
||||
task=m-iop-one-shot-agent-model-comparison/03+01_rubric_version_contract, plan=0, tag=API
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files. Run the applicable verification commands directly and record fresh output in `Verification Results`; implementation-owned output is handoff evidence, not a substitute for reviewer verification. If implementation is present, repair missing or stale verification output instead of failing solely for insufficient recorded evidence. When verification exposes a defect, collect the necessary data, determine the exact root cause, and select one concrete fix before generating the follow-up plan; never delegate investigation or remedy selection to the worker.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals.
|
||||
2. Archive `CODE_REVIEW-cloud-G08.md` → `code_review_cloud_G08_0.log` and `PLAN-local-G08.md` → `plan_local_G08_0.log`.
|
||||
3. If PASS, write `complete.log` and move active task directory to `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/03+01_rubric_version_contract/`. If WARN/FAIL, fully write the next filesystem state required by the code-review skill.
|
||||
4. If PASS, preserve the first-line `milestone-task` metadata in `complete.log` and report it for runtime aggregation. Roadmap state evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check applicable `Review-Only Checklist` items at the final `.log` location before reporting.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|---------|
|
||||
| API-1 Add a closed dual-version rubric catalog | [ ] |
|
||||
| API-2 Bind prompt and durable validation to manifest version | [ ] |
|
||||
| API-3 Preserve historical compatibility and exact new semantics | [ ] |
|
||||
| API-4 Run the complete benchmark regression | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] [API-1] Add an immutable `one-shot-agent-comparison-v1` seven-category rubric while preserving exact `landing-quality-v1` loader/schema/worksheet behavior.
|
||||
- [ ] [API-2] Make evaluator prompt generation and worksheet loading use the manifest-selected rubric version and reject cross-version output.
|
||||
- [ ] [API-3] Add dual-version manifest, rubric, and scoring regressions for exact category order, maxima, total 100, prompt text, and historical compatibility, and update the project benchmark skill to the manifest-selected contract.
|
||||
- [ ] [API-4] Run all rubric/scoring and benchmark regression suites plus shipped manifest validation.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Run applicable required verification and record fresh command/output; repair reviewer-reconstructable evidence gaps instead of forwarding them to another plan.
|
||||
- [ ] For every Required/Suggested finding, record reviewer-collected `Evidence`, exact `Root Cause`, and one `Selected Fix` with affected files/symbols/tests and acceptance commands before creating a follow-up plan.
|
||||
- [ ] Archive active `CODE_REVIEW-cloud-G08.md` to `code_review_cloud_G08_0.log`.
|
||||
- [ ] Archive active `PLAN-local-G08.md` to `plan_local_G08_0.log`.
|
||||
- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write `complete.log` based on `agent-ops/skills/common/code-review/templates/complete-log-template.md` and leave no active `.md` files.
|
||||
- [ ] If PASS, move active task directory `agent-task/m-iop-one-shot-agent-model-comparison/03+01_rubric_version_contract/` to `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/03+01_rubric_version_contract/` and update this checklist at the final archive path.
|
||||
- [ ] If PASS, preserve and report `milestone-task` metadata for runtime aggregation, without modifying roadmap or directly calling `update-roadmap`.
|
||||
- [ ] If PASS for split work, remove empty active parent `agent-task/m-iop-one-shot-agent-model-comparison/` or verify it was kept due to remaining siblings/files.
|
||||
- [ ] If WARN/FAIL, write the next filesystem state matching code-review verdict and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record any deviations from the plan and the rationale here._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record key design decisions here._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Confirm the legacy version/category tuple remains exact and reloadable.
|
||||
- Confirm the new seven categories and maxima exactly match the approved SDD and total 100.
|
||||
- Confirm prompt and both worksheet load paths use `manifest.rubric_version`.
|
||||
- Confirm unknown and cross-version worksheets fail before scored result publication.
|
||||
|
||||
## Verification Results
|
||||
|
||||
Record actual stdout/stderr under each command. If a command changes, document the replacement and reason in `Deviations from Plan`.
|
||||
|
||||
### Dependency Verification
|
||||
|
||||
```bash
|
||||
python3 - <<'PY'
|
||||
from pathlib import Path
|
||||
active = Path("agent-task/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/complete.log")
|
||||
archived = sorted(Path("agent-task/archive").glob("*/*/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/complete.log"))
|
||||
candidates = [path for path in (active, *archived) if path.is_file()]
|
||||
assert len(candidates) == 1, candidates
|
||||
print(f"ok: predecessor complete {candidates[0]}")
|
||||
PY
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
### API-1 Verification
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.manifest_test
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
### API-2 Verification
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.scoring_test.ScoringTest.test_manifest_selected_rubric_drives_prompt_and_worksheet_validation
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
### API-3 Verification
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.scoring_test scripts.agent_benchmark.manifest_test
|
||||
rg --sort path -n --fixed-strings 'one-shot-agent-comparison-v1' agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
### API-4 and Final Verification
|
||||
|
||||
```bash
|
||||
for manifest in scripts/fixtures/agent-comparison-benchmark-manifest.example.json scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json; do python3 scripts/agent_comparison_benchmark.py validate --manifest "$manifest"; done
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.scoring_test scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these (archive, complete.log, and task-directory archive move are review-agent only) |
|
||||
| Archive Evidence Snapshot | Fixed at stub creation from plan when present | Implementing agent uses it as default prior-loop context; read only the specific archive files cited there when more detail is required |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results (section headings + commands) | Implementing agent, then review agent | Implementing agent records initial output; review agent reruns applicable commands and may fill, replace, or append fresh verified output before verdict. Implementing-agent command changes require a `Deviations from Plan` entry |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,304 @@
|
|||
<!-- task=m-iop-one-shot-agent-model-comparison/03+01_rubric_version_contract plan=0 tag=API milestone-task=fixture-lock -->
|
||||
|
||||
# Plan - Versioned One-Shot Benchmark Rubric Contract
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Do not start until exactly one active or archived `01_execution_order_contract/complete.log` is resolved by the dependency command below. Filling the implementation-owned sections in `CODE_REVIEW-cloud-G08.md` is the mandatory last implementation step. Run every verification command, record actual notes and stdout/stderr, keep both active files in place, and report ready for review; only the code-review skill may finalize or archive this task. If blocked, record only the exact blocker, attempted commands/output, and resume condition in the implementation-owned evidence fields. Do not ask the user, call user-input tools, create control-plane stop files, classify the next state, archive logs, or write `complete.log`.
|
||||
|
||||
## Background
|
||||
|
||||
The approved one-shot comparison SDD fixes seven rubric categories totaling 100, while the existing `landing-quality-v1` worksheet contains five different categories. Reusing that version name would mutate historical meaning, so this packet preserves legacy runs and adds a manifest-selected version for the approved benchmark rubric.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `scripts/agent_benchmark/manifest.py`
|
||||
- `scripts/agent_benchmark/manifest_test.py`
|
||||
- `scripts/agent_benchmark/rubric.py`
|
||||
- `scripts/agent_benchmark/rubric_test.py`
|
||||
- `scripts/agent_benchmark/scoring.py`
|
||||
- `scripts/agent_benchmark/scoring_test.py`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-manifest.example.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json`
|
||||
- `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/iop-one-shot-agent-model-comparison.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/iop-one-shot-agent-model-comparison/SDD.md`
|
||||
- `agent-spec/testing/agent-comparison-benchmark.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- SDD: `agent-roadmap/sdd/knowledge-tool-optimization-extension/iop-one-shot-agent-model-comparison/SDD.md`, status `[승인됨]`, implementation lock released, no unresolved user review.
|
||||
- First-line scope: `milestone-task=fixture-lock`.
|
||||
- Target scenario: S01 requires the same fixture, viewport, rubric checksum/version for every cell. The approved rubric at SDD lines 90-91 is 요구사항 충족 25, 시각 완성도 25, 반응형·접근성 15, 이미지 활용·디테일 10, 동작 안정성 10, 코드 품질 10, 자체 검증 완결성 5.
|
||||
- Evidence Map row: S01 requires fixture prompt/assets/workspace/rubric digest evidence.
|
||||
- The checklist therefore creates a distinct immutable rubric version, retains legacy validation, and binds prompt/output validation to the manifest-selected version. The final manifest packet supplies the concrete version/digest evidence.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- No separate handoff was supplied. Repository-native evidence came from the complete manifest, rubric, scoring, and test files above.
|
||||
- Fresh baseline at starting HEAD `b197e5db70637f87017a024a847e3e53fdc72e8b`: all example manifests validated and the 178-test manifest/attempt/connectivity suite passed. `rubric_test` and `scoring_test` are additional required fresh suites for this packet.
|
||||
- Current `RUBRIC_CATEGORIES` is five entries totaling 100 (`rubric.py:21-27`). `validate_worksheet` accepts only global `RUBRIC_VERSION` (`rubric.py:82-126`), and `_prompt` hardcodes those categories and `landing-quality-v1` (`scoring.py:1207-1222`).
|
||||
- The safe compatibility rule is additive: keep `landing-quality-v1` and its five-category table readable; add `one-shot-agent-comparison-v1` with the approved seven-category table; reject unknown and cross-version worksheets.
|
||||
- Fresh output is required; cached output is not accepted. No external runner is required.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
- Existing rubric tests validate only the five-category legacy table.
|
||||
- Existing scoring tests generate only legacy worksheets and cannot prove prompt/worksheet selection from the manifest.
|
||||
- Existing manifest/schema tests accept only one rubric constant.
|
||||
- No test rejects a valid legacy worksheet when the allocation requires the new rubric, or vice versa.
|
||||
|
||||
### Symbol References
|
||||
|
||||
- `manifest.RUBRIC_VERSION` is referenced only by `manifest.py:711` and `rubric.py:16,83,126`. Replace internal single-version validation with a closed version catalog while preserving a legacy alias if needed for compatibility.
|
||||
- `rubric.RUBRIC_CATEGORIES` is imported by `scoring.py:44-49`, `scoring_test.py:56`, and `rubric_test.py:10`. Preserve it as the legacy tuple for existing callers, add a version lookup, and migrate production scoring to the lookup.
|
||||
- `_prompt` is private and called at `scoring.py:1879`; change both definition and call together.
|
||||
- `load_worksheet` call sites are `scoring.py:1641` and `scoring.py:1972`; both must pass the manifest-selected expected version.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
- This is dependent child `03+01_rubric_version_contract`. Its stable invariant is: rubric version selects exactly one immutable ordered category table from manifest load through evaluator prompt, worksheet validation, durable result, and historical reload.
|
||||
- PASS evidence is dual-version loader/rubric/scoring tests plus full regression.
|
||||
- Predecessor index 01 is currently active at `agent-task/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/` but has no `complete.log`; implementation is dependency-blocked until that exact file exists. The dependency is required because both packets modify `manifest.py`, `manifest_test.py`, and the schema.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
- Do not delete, rename, or reinterpret `landing-quality-v1`; old manifests and immutable run snapshots must still load and score under the old table.
|
||||
- Do not change automatic web gate eligibility, evaluator identity blinding, score persistence, caller bindings, fixture bytes, or report layout.
|
||||
- Do not update the project benchmark skill or living spec in this packet; their current general 100-point description remains true, and operational selection remains manifest-owned.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- `evaluation_mode=first-pass`; `finalizer=finalize-task-policy.sh`, `finalizer_mode=pair`.
|
||||
- Build closures: scope/context/verification/evidence/ownership/decision all closed. Scores `2+1+2+1+2=G08`; base/final route `local-fit`, lane `local`, catalog `worker/local/G08`, filename `PLAN-local-G08.md`.
|
||||
- Review closures: all closed. Scores `2+1+2+1+2=G08`; route `official-review`, lane `cloud`, catalog `review/cloud/G08`, filename `CODE_REVIEW-cloud-G08.md`.
|
||||
- `large_indivisible_context=false`; positive loop risks: `boundary_contract`, `structured_interpretation`, `variant_product` (`count=3`); `review_rework_count=0`; `evidence_integrity_failure=false`; no capability gap.
|
||||
|
||||
## Dependencies and Execution Order
|
||||
|
||||
1. Resolve exactly one predecessor `complete.log`: the active sibling path or one matching `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/complete.log`.
|
||||
2. Implement this packet against the reviewed 01 source. Do not infer completion from an active PLAN or review stub.
|
||||
3. On PASS, `04+02,03_locked_benchmark_manifest` may proceed only after both index 02 and this index 03 have `complete.log`.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] [API-1] Add an immutable `one-shot-agent-comparison-v1` seven-category rubric while preserving exact `landing-quality-v1` loader/schema/worksheet behavior.
|
||||
- [ ] [API-2] Make evaluator prompt generation and worksheet loading use the manifest-selected rubric version and reject cross-version output.
|
||||
- [ ] [API-3] Add dual-version manifest, rubric, and scoring regressions for exact category order, maxima, total 100, prompt text, and historical compatibility, and update the project benchmark skill to the manifest-selected contract.
|
||||
- [ ] [API-4] Run all rubric/scoring and benchmark regression suites plus shipped manifest validation.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Add a closed dual-version rubric catalog
|
||||
|
||||
**Problem**
|
||||
|
||||
`manifest.py:30` exposes one allowed rubric version, and `rubric.py:21-27` gives that version a five-category table that differs from the approved SDD. Changing the existing tuple in place would silently rewrite historical `landing-quality-v1` meaning.
|
||||
|
||||
**Solution**
|
||||
|
||||
Keep the legacy string and tuple unchanged. Add `ONE_SHOT_RUBRIC_VERSION = "one-shot-agent-comparison-v1"`, a closed allowed-version tuple in `manifest.py`, and an ordered category mapping in `rubric.py`. The new table is exactly:
|
||||
|
||||
```python
|
||||
ONE_SHOT_RUBRIC_CATEGORIES = (
|
||||
("requirements_fidelity", 25),
|
||||
("visual_completeness", 25),
|
||||
("responsive_accessibility", 15),
|
||||
("image_detail_usage", 10),
|
||||
("behavior_stability", 10),
|
||||
("code_quality", 10),
|
||||
("self_verification", 5),
|
||||
)
|
||||
```
|
||||
|
||||
Add `rubric_categories(version)` that fails closed for unknown versions. Update the manifest schema from a single `const` to an exact two-value enum. `validate_worksheet` must read the worksheet version, optionally require an `expected_version`, and validate against only that version's ordered table.
|
||||
|
||||
Before (`scripts/agent_benchmark/rubric.py:21-27`):
|
||||
|
||||
```python
|
||||
RUBRIC_CATEGORIES = (
|
||||
("task_fidelity", 25),
|
||||
("visual_hierarchy", 25),
|
||||
("responsive_composition", 20),
|
||||
("typography_readability", 15),
|
||||
("polish_consistency", 15),
|
||||
)
|
||||
```
|
||||
|
||||
After: retain that exact tuple as the legacy value and add the new tuple/mapping without mutation.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `scripts/agent_benchmark/manifest.py`: accept the closed legacy/new rubric version catalog.
|
||||
- [ ] `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json`: mirror the exact two-version enum.
|
||||
- [ ] `scripts/agent_benchmark/rubric.py`: provide versioned category lookup and expected-version worksheet validation.
|
||||
- [ ] `scripts/agent_benchmark/rubric_test.py`: verify exact tables, totals, ordering, unknown versions, and cross-version rejection.
|
||||
- [ ] `scripts/agent_benchmark/manifest_test.py`: verify loader/schema parity for both allowed versions and an unknown version.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
Write tests. Preserve every existing legacy assertion. Add `test_one_shot_rubric_exact_categories_and_total_are_accepted`, `test_cross_version_worksheet_is_rejected`, and manifest/schema two-version parity. Assert both category sums equal exactly 100.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.manifest_test
|
||||
```
|
||||
|
||||
Expected: both versions pass their exact table; malformed, unknown, reordered, and cross-version cases fail closed.
|
||||
|
||||
### [API-2] Bind prompt and durable score validation to the manifest version
|
||||
|
||||
**Problem**
|
||||
|
||||
`_prompt` uses global categories and a literal version (`scoring.py:1207-1222`), while both `load_worksheet` calls (`scoring.py:1641`, `1972`) validate only the global legacy version. A new manifest version could therefore ask for or accept the wrong worksheet.
|
||||
|
||||
**Solution**
|
||||
|
||||
Change `_prompt` to accept `rubric_version`, look up its exact table, render that version string, and call it with `manifest.rubric_version`. Extend `load_worksheet`/`validate_worksheet` with `expected_version` and pass `manifest.rubric_version` at both scoring call sites. Keep allocation/result binding to the manifest version and reject a worksheet whose internal version differs before publication.
|
||||
|
||||
Before (`scripts/agent_benchmark/scoring.py:1207-1222`):
|
||||
|
||||
```python
|
||||
def _prompt(blind: BlindWorkspace) -> bytes:
|
||||
categories = ", ".join(
|
||||
f"{ident} ({maximum})" for ident, maximum in RUBRIC_CATEGORIES
|
||||
)
|
||||
...
|
||||
"landing-quality-v1, integer scores within each maximum ..."
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
def _prompt(blind: BlindWorkspace, rubric_version: str) -> bytes:
|
||||
categories = ", ".join(
|
||||
f"{ident} ({maximum})"
|
||||
for ident, maximum in rubric_categories(rubric_version)
|
||||
)
|
||||
...
|
||||
f"{rubric_version}, integer scores within each maximum ..."
|
||||
```
|
||||
|
||||
The `blind` parameter may remain for interface stability even though it is not interpolated. Do not expose producer identity.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `scripts/agent_benchmark/scoring.py`: select prompt categories/version and worksheet validation from `manifest.rubric_version`.
|
||||
- [ ] `scripts/agent_benchmark/scoring_test.py`: add a new-version scoring success and cross-version failure test.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
Write `test_manifest_selected_rubric_drives_prompt_and_worksheet_validation`. Build a new-version manifest, emit the exact seven-category worksheet, and assert success plus prompt category/version content. On a fresh score attempt, emit a structurally valid legacy worksheet and assert `invalid_worksheet` with no scored worksheet publication.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.scoring_test.ScoringTest.test_manifest_selected_rubric_drives_prompt_and_worksheet_validation
|
||||
```
|
||||
|
||||
Expected: selected-version success and cross-version rejection both pass.
|
||||
|
||||
### [API-3] Preserve historical compatibility and exact new semantics
|
||||
|
||||
**Problem**
|
||||
|
||||
An additive version is safe only if old manifests, prompts, worksheets, and immutable score reloads retain byte-level meaning.
|
||||
|
||||
**Solution**
|
||||
|
||||
Keep existing example manifests on `landing-quality-v1`. Make test helpers accept a version instead of replacing their legacy default. Add explicit assertions that legacy canonical worksheet bytes and prompts remain unchanged, while the new version has seven ordered categories and total 100. Update the project benchmark skill's scoring phase so it requires the exact rubric selected by the immutable manifest and names both supported versions; do not leave its current line 87 fixed to the legacy version.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `scripts/agent_benchmark/rubric_test.py`: dual-version canonical and boundary assertions.
|
||||
- [ ] `scripts/agent_benchmark/scoring_test.py`: legacy default plus new-version selected behavior.
|
||||
- [ ] `scripts/agent_benchmark/manifest_test.py`: both versions and shipped examples.
|
||||
- [ ] `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md`: document manifest-selected exact rubric behavior and the closed supported version set.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
Write tests; do not mass-replace legacy version literals in unrelated fixtures. The new final manifest is added only by the dependent packet.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.scoring_test scripts.agent_benchmark.manifest_test
|
||||
rg --sort path -n --fixed-strings 'one-shot-agent-comparison-v1' agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md
|
||||
```
|
||||
|
||||
Expected: all legacy and new rubric tests report `OK`.
|
||||
|
||||
### [API-4] Run the complete benchmark regression
|
||||
|
||||
**Problem**
|
||||
|
||||
Rubric validation is used by scoring recovery and durable result reads, so focused tests are not enough.
|
||||
|
||||
**Solution**
|
||||
|
||||
Run all rubric/scoring and the existing benchmark suites, validate shipped legacy examples, and check whitespace.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `CODE_REVIEW-cloud-G08.md`: record exact decisions, deviations, and all actual command output.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
No further files. Existing recovery, redaction, scoring retry, manifest digest, and connectivity suites are the full oracle.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
for manifest in scripts/fixtures/agent-comparison-benchmark-manifest.example.json scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json; do python3 scripts/agent_comparison_benchmark.py validate --manifest "$manifest"; done
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.scoring_test scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Expected: three legacy manifests validate, all tests report `OK`, and whitespace check is silent.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Item |
|
||||
|------|------|
|
||||
| `scripts/agent_benchmark/manifest.py` | API-1 |
|
||||
| `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json` | API-1 |
|
||||
| `scripts/agent_benchmark/rubric.py` | API-1 |
|
||||
| `scripts/agent_benchmark/rubric_test.py` | API-1, API-3 |
|
||||
| `scripts/agent_benchmark/manifest_test.py` | API-1, API-3 |
|
||||
| `scripts/agent_benchmark/scoring.py` | API-2 |
|
||||
| `scripts/agent_benchmark/scoring_test.py` | API-2, API-3 |
|
||||
| `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md` | API-3 |
|
||||
| `agent-task/m-iop-one-shot-agent-model-comparison/03+01_rubric_version_contract/CODE_REVIEW-cloud-G08.md` | API-4 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
Run from `/config/workspace/iop-s0` after predecessor 01 is complete; fresh output is required and cached output is not acceptable.
|
||||
|
||||
```bash
|
||||
python3 - <<'PY'
|
||||
from pathlib import Path
|
||||
active = Path("agent-task/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/complete.log")
|
||||
archived = sorted(Path("agent-task/archive").glob("*/*/m-iop-one-shot-agent-model-comparison/01_execution_order_contract/complete.log"))
|
||||
candidates = [path for path in (active, *archived) if path.is_file()]
|
||||
assert len(candidates) == 1, candidates
|
||||
print(f"ok: predecessor complete {candidates[0]}")
|
||||
PY
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.manifest_test
|
||||
python3 -m unittest scripts.agent_benchmark.scoring_test.ScoringTest.test_manifest_selected_rubric_drives_prompt_and_worksheet_validation
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.scoring_test scripts.agent_benchmark.manifest_test
|
||||
rg --sort path -n --fixed-strings 'one-shot-agent-comparison-v1' agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md
|
||||
for manifest in scripts/fixtures/agent-comparison-benchmark-manifest.example.json scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json; do python3 scripts/agent_comparison_benchmark.py validate --manifest "$manifest"; done
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.scoring_test scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Expected: dependency check and every command exit 0, legacy examples validate, all suites report `OK`, and `git diff --check` is silent.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,243 @@
|
|||
<!-- task=m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest plan=1 tag=TEST milestone-task=fixture-lock,route-readiness,matrix-lock -->
|
||||
|
||||
# Code Review Reference - TEST
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-12
|
||||
task=m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest, plan=1, tag=TEST
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior plan: `agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/plan_local_G07_0.log`.
|
||||
- Prior review: `agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/code_review_cloud_G07_0.log`.
|
||||
- The prior pair contains no implementation evidence or official verdict. Epic self-review found that the final manifest would select the new rubric while the public project skill remained fixed to `landing-quality-v1`; it also found an incomplete agy resume gate and ambiguous testbed source provenance.
|
||||
- Replan atomically assigns rubric skill/contract-test synchronization to this dependent packet, requires the adapter-known agy version plus its documented transport surface, and records the independent read-only testbed's exact clean commit without requiring equality to the orchestration repository. A fresh read-only probe also found the current Edge artifact is Mach-O and both runtime artifacts predate that commit, so operator rebuild is an explicit resume condition. Baseline fixture checksum/shape and the 178-test manifest/attempt/connectivity suite were green at `b197e5db70637f87017a024a847e3e53fdc72e8b`.
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files. Run the applicable verification commands directly and record fresh output in `Verification Results`; implementation-owned output is handoff evidence, not a substitute for reviewer verification. If implementation is present, repair missing or stale verification output instead of failing solely for insufficient recorded evidence. When verification exposes a defect, collect the necessary data, determine the exact root cause, and select one concrete fix before generating the follow-up plan; never delegate investigation or remedy selection to the worker.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals.
|
||||
2. Archive `CODE_REVIEW-cloud-G09.md` → `code_review_cloud_G09_1.log` and `PLAN-cloud-G09.md` → `plan_cloud_G09_1.log`.
|
||||
3. If PASS, write `complete.log` and move active task directory to `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/`. If WARN/FAIL, fully write the next filesystem state required by the code-review skill.
|
||||
4. If PASS, preserve the first-line `milestone-task` metadata in `complete.log` and report it for runtime aggregation. Roadmap state evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check applicable `Review-Only Checklist` items at the final `.log` location before reporting.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|---------|
|
||||
| TEST-1 Create one exact immutable readiness manifest | [ ] |
|
||||
| TEST-2 Produce one redacted nine-cell readiness record | [ ] |
|
||||
| TEST-3 Preserve the full benchmark baseline | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] [TEST-1] Add the exact tracked bench-02 manifest and a static regression that locks fixture checksum/version, viewports, rubric, policies, C01-C09 bindings, seed, and seeded order; atomically update project-skill scoring language and its contract test to the manifest-selected rubric catalog.
|
||||
- [ ] [TEST-2] Run the secret-safe caller/testbed/environment gate and public all-cell preflight; require nine ready results or record the exact blocker and resume condition without substitution.
|
||||
- [ ] [TEST-3] Run final schema, focused regression, full benchmark suite, and whitespace verification while leaving scored execution untouched.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Run applicable required verification and record fresh command/output; repair reviewer-reconstructable evidence gaps instead of forwarding them to another plan.
|
||||
- [ ] For every Required/Suggested finding, record reviewer-collected `Evidence`, exact `Root Cause`, and one `Selected Fix` with affected files/symbols/tests and acceptance commands before creating a follow-up plan.
|
||||
- [ ] Archive active `CODE_REVIEW-cloud-G09.md` to `code_review_cloud_G09_1.log`.
|
||||
- [ ] Archive active `PLAN-cloud-G09.md` to `plan_cloud_G09_1.log`.
|
||||
- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write `complete.log` based on `agent-ops/skills/common/code-review/templates/complete-log-template.md` and leave no active `.md` files.
|
||||
- [ ] If PASS, move active task directory `agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/` to `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/` and update this checklist at the final archive path.
|
||||
- [ ] If PASS, preserve and report `milestone-task` metadata for runtime aggregation, without modifying roadmap or directly calling `update-roadmap`.
|
||||
- [ ] If PASS for split work, remove empty active parent `agent-task/m-iop-one-shot-agent-model-comparison/` or verify it was kept due to remaining siblings/files.
|
||||
- [ ] If WARN/FAIL, write the next filesystem state matching code-review verdict and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record any deviations from the plan and the rationale here._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record key design decisions here._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Confirm both predecessor completion files exist before reviewing implementation.
|
||||
- Confirm the complete C01-C09 map, seed-derived order, fixture checksum, two images, policies, and new rubric version are exact.
|
||||
- Confirm hybrid routes include `repair`, cloud stages use high, and work uses `ornith-fast` with no invented effort.
|
||||
- Confirm project-skill scoring uses the exact manifest-selected rubric, names the closed two-version set, and preserves predecessor 02's all-cell preflight contract.
|
||||
- Confirm the external gate captures the clean independent `../iop-s2` dev HEAD and accepts only the production adapter's exact known agy version/help contract.
|
||||
- Confirm both iop-s2 artifacts are current-HEAD AArch64 ELF executables before accepting endpoint readiness; `test -x` alone is insufficient.
|
||||
- Confirm live readiness came only from public preflight and durable evidence contains no endpoint, secret, or raw config value.
|
||||
- Confirm no scored run/score/report command was executed in this Epic.
|
||||
|
||||
## Verification Results
|
||||
|
||||
Record actual stdout/stderr under each command. If a command changes, document the replacement and reason in `Deviations from Plan`. TEST-2's full command block is fixed in the plan and must be copied with its actual output or exact blocker here.
|
||||
|
||||
### Dependency Verification
|
||||
|
||||
```bash
|
||||
python3 - <<'PY'
|
||||
from pathlib import Path
|
||||
root = Path("agent-task")
|
||||
group = "m-iop-one-shot-agent-model-comparison"
|
||||
for subtask in ("02_route_preflight_contract", "03+01_rubric_version_contract"):
|
||||
active = root / group / subtask / "complete.log"
|
||||
archived = sorted((root / "archive").glob(f"*/*/{group}/{subtask}/complete.log"))
|
||||
candidates = [path for path in (active, *archived) if path.is_file()]
|
||||
assert len(candidates) == 1, (subtask, candidates)
|
||||
print(f"ok: predecessor complete {candidates[0]}")
|
||||
PY
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
### TEST-1 Verification
|
||||
|
||||
```bash
|
||||
python3 scripts/agent_comparison_benchmark.py validate --manifest scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test.ManifestValidationTest.test_iop_one_shot_manifest_locks_benchmark_readiness
|
||||
python3 -m unittest scripts.agent_benchmark.skill_contract_test
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
### TEST-2 External Preflight
|
||||
|
||||
```bash
|
||||
set -euo pipefail
|
||||
python3 - <<'PY'
|
||||
from pathlib import Path
|
||||
root = Path("agent-task")
|
||||
group = "m-iop-one-shot-agent-model-comparison"
|
||||
for subtask in ("02_route_preflight_contract", "03+01_rubric_version_contract"):
|
||||
active = root / group / subtask / "complete.log"
|
||||
archived = sorted((root / "archive").glob(f"*/*/{group}/{subtask}/complete.log"))
|
||||
candidates = [path for path in (active, *archived) if path.is_file()]
|
||||
assert len(candidates) == 1, (subtask, candidates)
|
||||
print(f"ok: predecessor complete {candidates[0]}")
|
||||
PY
|
||||
test "$(git -C ../iop-s2 branch --show-current)" = "dev"
|
||||
test -z "$(git -C ../iop-s2 status --short)"
|
||||
testbed_head="$(git -C ../iop-s2 rev-parse HEAD)"
|
||||
test -n "$testbed_head"
|
||||
printf 'ok: testbed branch=dev head=%s clean=true\n' "$testbed_head"
|
||||
command -v readelf >/dev/null
|
||||
test -x ../iop-s2/build/bin/iop-edge
|
||||
test -x ../iop-s2/build/dev/iop-node
|
||||
test -f ../iop-s2/configs/edge.yaml
|
||||
commit_epoch="$(git -C ../iop-s2 show -s --format=%ct HEAD)"
|
||||
for binary in ../iop-s2/build/bin/iop-edge ../iop-s2/build/dev/iop-node; do
|
||||
test "$(stat -c %Y "$binary")" -ge "$commit_epoch"
|
||||
readelf -h "$binary" | rg 'Machine:\s+AArch64' >/dev/null
|
||||
done
|
||||
../iop-s2/build/bin/iop-edge --help >/dev/null
|
||||
test -n "$(../iop-s2/build/dev/iop-node version)"
|
||||
printf 'ok: current-HEAD Linux AArch64 Edge/Node artifacts are executable\n'
|
||||
claude --version
|
||||
agy --version
|
||||
codex --version
|
||||
python3 - <<'PY'
|
||||
import subprocess
|
||||
from scripts.agent_benchmark.agy_iop import AGY_KNOWN_VERSION, inspect_agy_iop_capability
|
||||
version_run = subprocess.run(["agy", "--version"], check=True, capture_output=True, text=True)
|
||||
help_run = subprocess.run(["agy", "--help"], check=True, capture_output=True, text=True)
|
||||
version = (version_run.stdout + version_run.stderr).strip()
|
||||
help_text = help_run.stdout + help_run.stderr
|
||||
capability = inspect_agy_iop_capability(version, help_text)
|
||||
assert capability.version == AGY_KNOWN_VERSION, (capability.version, AGY_KNOWN_VERSION)
|
||||
assert capability.iop_transport_supported, capability
|
||||
print(f"ok: adapter-known agy {AGY_KNOWN_VERSION} transport is documented")
|
||||
PY
|
||||
python3 - <<'PY'
|
||||
import json, os, re
|
||||
name = re.compile(r"^[A-Za-z_][A-Za-z0-9_]*$")
|
||||
for caller in ("CLAUDE", "AGY", "CODEX"):
|
||||
assert os.environ.get(f"IOP_BENCH_{caller}_BASE_URL")
|
||||
ref = os.environ.get(f"IOP_BENCH_{caller}_SECRET_ENV", "")
|
||||
assert name.fullmatch(ref) and os.environ.get(ref)
|
||||
config_ref = os.environ.get("IOP_BENCH_CONFIG_OBSERVATION_ENV", "")
|
||||
assert name.fullmatch(config_ref) and os.environ.get(config_ref)
|
||||
value = json.loads(os.environ[config_ref])
|
||||
assert value.get("schema_version") == "1" and isinstance(value.get("routes"), list)
|
||||
print("ok: benchmark environment references present")
|
||||
PY
|
||||
preflight_output="$(python3 scripts/agent_comparison_benchmark.py preflight --manifest scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json)"
|
||||
printf '%s\n' "$preflight_output"
|
||||
case "$preflight_output" in *"status=ready ready=9 registration_required=0 implementation_gap=0"*) ;; *) exit 1 ;; esac
|
||||
run_id="${preflight_output#*run_id=}"; run_id="${run_id%% *}"
|
||||
python3 - "$run_id" <<'PY'
|
||||
import json, os, sys
|
||||
from pathlib import Path
|
||||
from scripts.agent_benchmark.manifest import load_manifest
|
||||
manifest_path = Path("scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json")
|
||||
manifest = load_manifest(manifest_path, repo_root=Path.cwd())
|
||||
run_root = Path(manifest.output_root) / sys.argv[1]
|
||||
record = json.loads((run_root / "preflight/preflight-000001.json").read_text(encoding="ascii"))
|
||||
assert record["status"] == "ready"
|
||||
assert [item["cell"]["id"] for item in record["results"]] == [cell.id for cell in manifest.matrix]
|
||||
assert len(record["results"]) == 9 and all(item["status"] == "ready" for item in record["results"])
|
||||
sensitive = []
|
||||
for caller in ("CLAUDE", "AGY", "CODEX"):
|
||||
sensitive.append(os.environ[f"IOP_BENCH_{caller}_BASE_URL"].encode())
|
||||
sensitive.append(os.environ[os.environ[f"IOP_BENCH_{caller}_SECRET_ENV"]].encode())
|
||||
config_ref = os.environ["IOP_BENCH_CONFIG_OBSERVATION_ENV"]
|
||||
sensitive.append(os.environ[config_ref].encode())
|
||||
durable = b"".join(path.read_bytes() for path in run_root.rglob("*") if path.is_file())
|
||||
assert all(value and value not in durable for value in sensitive)
|
||||
print("ok: nine ready results are manifest-bound and runtime values are absent")
|
||||
PY
|
||||
```
|
||||
|
||||
_Actual output or exact blocker and resume condition:_
|
||||
|
||||
### TEST-3 and Final Verification
|
||||
|
||||
```bash
|
||||
for manifest in scripts/fixtures/agent-comparison-benchmark-manifest.example.json scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json; do python3 scripts/agent_comparison_benchmark.py validate --manifest "$manifest"; done
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.scoring_test scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test scripts.agent_benchmark.skill_contract_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these (archive, complete.log, and task-directory archive move are review-agent only) |
|
||||
| Archive Evidence Snapshot | Fixed at stub creation from plan when present | Implementing agent uses it as default prior-loop context; read only the specific archive files cited there when more detail is required |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results (section headings + commands) | Implementing agent, then review agent | Implementing agent records initial output; review agent reruns applicable commands and may fill, replace, or append fresh verified output before verdict. Implementing-agent command changes require a `Deviations from Plan` entry |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,363 @@
|
|||
<!-- task=m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest plan=1 tag=TEST milestone-task=fixture-lock,route-readiness,matrix-lock -->
|
||||
|
||||
# Plan - Locked C01-C09 Benchmark Readiness Manifest
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Do not start until the dependency command resolves exactly one active or archived `complete.log` for both predecessor indices 02 and 03. Filling the implementation-owned sections in `CODE_REVIEW-cloud-G09.md` is the mandatory last implementation step. Run every verification command, record actual notes and stdout/stderr, keep both active files in place, and report ready for review; only the code-review skill may finalize or archive this task. If blocked, record only the exact blocker, attempted commands/output, and resume condition in the implementation-owned evidence fields. Do not ask the user, call user-input tools, create control-plane stop files, classify the next state, archive logs, or write `complete.log`.
|
||||
|
||||
## Background
|
||||
|
||||
The repository already contains the reusable two-image vanilla landing-page fixture, but no immutable manifest combines it with the approved C01-C09 matrix, explicit seed, new rubric version, and exact live route evidence. This packet creates that final tracked input and runs the public preflight without substituting missing callers, credentials, models, or presets.
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior plan: `agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/plan_local_G07_0.log`.
|
||||
- Prior review: `agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/code_review_cloud_G07_0.log`.
|
||||
- The prior pair contains no implementation evidence or official verdict. Epic self-review found that the final manifest would select the new rubric while the public project skill remained fixed to `landing-quality-v1`; it also found an incomplete agy resume gate and ambiguous testbed source provenance.
|
||||
- Replan atomically assigns rubric skill/contract-test synchronization to this dependent packet, requires the adapter-known agy version plus its documented transport surface, and records the independent read-only testbed's exact clean commit without requiring equality to the orchestration repository. A fresh read-only probe also found the current Edge artifact is Mach-O and both runtime artifacts predate that commit, so operator rebuild is an explicit resume condition. Baseline fixture checksum/shape and the 178-test manifest/attempt/connectivity suite were green at `b197e5db70637f87017a024a847e3e53fdc72e8b`.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `scripts/agent_benchmark/manifest.py`
|
||||
- `scripts/agent_benchmark/manifest_test.py`
|
||||
- `scripts/agent_benchmark/attempts.py`
|
||||
- `scripts/agent_benchmark/attempts_test.py`
|
||||
- `scripts/agent_benchmark/live_iop.py`
|
||||
- `scripts/agent_benchmark/agy_iop.py`
|
||||
- `scripts/agent_benchmark/connectivity_integration_test.py`
|
||||
- `scripts/agent_benchmark/skill_contract_test.py`
|
||||
- `scripts/agent_benchmark/rubric.py`
|
||||
- `scripts/agent_benchmark/rubric_test.py`
|
||||
- `scripts/agent_benchmark/scoring.py`
|
||||
- `scripts/agent_benchmark/scoring_test.py`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-manifest.example.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark/prompt.md`
|
||||
- `scripts/fixtures/agent-comparison-benchmark/reference.txt`
|
||||
- `scripts/fixtures/agent-comparison-benchmark/aurora-grid.svg`
|
||||
- `scripts/fixtures/agent-comparison-benchmark/orbit-rings.svg`
|
||||
- `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/iop-one-shot-agent-model-comparison.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/iop-one-shot-agent-model-comparison/SDD.md`
|
||||
- `agent-spec/testing/agent-comparison-benchmark.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- SDD: `agent-roadmap/sdd/knowledge-tool-optimization-extension/iop-one-shot-agent-model-comparison/SDD.md`, status `[승인됨]`, implementation lock released, no unresolved user review.
|
||||
- First-line scope: `milestone-task=fixture-lock,route-readiness,matrix-lock`.
|
||||
- S01 requires identical prompt/assets/workspace/viewports/rubric checksum/version; S02 requires a redacted all-cell preflight or exact blocker; S03 requires the immutable nine-cell manifest, repetitions 1, explicit seed, fresh/isolated policy, timeout, and bindings.
|
||||
- Evidence Map rows S01-S03 drive the static manifest assertions, secret-safe external preflight, and durable nine-result inspection below.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- No separate handoff was supplied. Repository-native evidence came from the complete fixture, manifest/schema, public CLI, adapter/store tests, SDD, spec, and contracts listed above.
|
||||
- Fresh baseline at starting HEAD `b197e5db70637f87017a024a847e3e53fdc72e8b`: all three shipped manifests validated; `python3 -m unittest scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test` ran 178 tests and passed.
|
||||
- Fixture evidence: prompt requires exactly root `index.html`, `styles.css`, `script.js`, vanilla HTML/CSS/JS, both local images, responsive desktop/mobile behavior, accessibility, and no external dependency. The assets resolve to exactly two images plus one reference file; checksum is `sha256:7dc1be6ed4a9f2f873016b708b99d827b0249c74f2ac287e1fcf8deade8dcd98`. Viewports are 1920x1080 and 375x812.
|
||||
- The direct example already fixes C01-C05 model/effort pairs. The approved SDD names the hybrid scenarios `gemini-hybrid` and `gpt-hybrid`; this packet uses those exact strings as both virtual request model and preset route id. Missing registration must remain `registration_required`, never fallback.
|
||||
- Planned seeded order for `bench-02-c01-c09-v1` under the reviewed domain-separated algorithm is: C02, C05, C03, C06, C08, C09, C01, C07, C04.
|
||||
- Fresh output is required; cached output is not accepted.
|
||||
|
||||
#### External Verification Preflight
|
||||
|
||||
- Runner/workdir: local Linux `aarch64`, `/config/workspace/iop-s0`.
|
||||
- Testbed: `/config/workspace/iop-s2`, branch `dev`, exact observed HEAD `1f2f7f1066fcf165a9e469bae77203b569b6f772`, clean at planning time. It is the manifest-selected independent read-only runtime, so equality with the iop-s0 orchestration HEAD is neither expected nor a valid sync test; execution must capture its exact current `dev` HEAD and clean state as provenance before using existing artifacts.
|
||||
- Artifacts: `../iop-s2/build/bin/iop-edge`, `../iop-s2/build/dev/iop-node`, and `../iop-s2/configs/edge.yaml` exist, but existence is insufficient. The Edge magic bytes are Mach-O (`cf fa ed fe`) and it exits 126 with `Exec format error` on Linux; Node is AArch64 ELF. Both artifact mtimes predate current testbed HEAD (`2026-08-10T00:09:38+09:00`), so they are stale and Edge is host-incompatible. The tracked example config exposes only local example models and no benchmark route/preset catalog; do not treat it as live registration.
|
||||
- Callers: Claude Code `2.1.227`, agy `1.1.12`, Codex CLI `0.147.0` were installed. Claude/Codex help probes completed. The production adapter accepts only `agy` `1.1.11`; current `1.1.12` also lacks the required `AGY_PROVIDER`, `AGY_OPENAI_BASE_URL`, and `AGY_OPENAI_API_KEY` help tokens, so it is definitively incompatible rather than merely missing a few tokens.
|
||||
- Runtime/config: every `IOP_BENCH_{CLAUDE,AGY,CODEX}_{BASE_URL,SECRET_ENV}` and `IOP_BENCH_CONFIG_OBSERVATION_ENV` was missing. Consequently runtime identity, endpoint hosts/ports, live catalog, auth, and route registration could not be observed. `ss` is unavailable; the public preflight's endpoint probe is the authoritative reachability check once URLs exist.
|
||||
- Exact resume condition: an operator must rebuild both iop-s2 artifacts from its exact clean `dev` HEAD for Linux AArch64, provide all benchmark env references and their non-empty referenced secret/config values, register all direct/hybrid routes with exact bindings, run those rebuilt Edge/Node endpoints, and install adapter-known `agy` `1.1.11` whose help satisfies every documented transport token. Execution proves both artifacts are AArch64 ELF, are not older than the captured commit, and expose their help/version commands. If adapter version support changes in a reviewed predecessor, use that code's exact `AGY_KNOWN_VERSION` and capability parser instead. No raw value is written to the repository or review file.
|
||||
- If the public preflight returns `registration_required` or `implementation_gap`, preserve its run record, record only the closed summary and run id in the review evidence, and stop. Do not edit `../iop-s2`, install tools, substitute models, or invoke callers outside the benchmark CLI.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
- No shipped manifest contains exactly C01-C09 with one explicit seed and the approved new rubric version.
|
||||
- Existing generic preset examples do not express Gemini/GPT plan→ornith-fast work→review/repair stage bindings.
|
||||
- No static regression asserts the fixture checksum, exact two-image paths, policy fields, cell map, seeded order, and hybrid binding table together.
|
||||
- The project benchmark skill still hardcodes the legacy worksheet, and its contract test does not require the closed version catalog or manifest-selected scoring language.
|
||||
- Live readiness is currently blocked by stale/host-incompatible testbed artifacts, missing environment/config, and incompatible agy version/help evidence; the plan contains the exact secret-safe resume/preflight command.
|
||||
|
||||
### Symbol References
|
||||
|
||||
- No symbol is renamed or removed.
|
||||
- The new fixture path is consumed by `load_manifest`, public CLI commands, `RunStore`, and the manifest regression test; no new API is introduced.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
- Direct-small classification produced zero code changes. The existing fixture bytes already meet the prompt/image/workspace requirements, but their final version/checksum evidence is inseparable from the planned seeded/rubric manifest and all-cell preflight. Editing a partial manifest before those contracts land would collide with planned work.
|
||||
- This is dependent child `04+02,03_locked_benchmark_manifest`. Its stable invariant is one tracked manifest whose fixture, policies, nine exact cells, order, rubric, and durable preflight all agree.
|
||||
- Predecessor index 02 is active at `agent-task/m-iop-one-shot-agent-model-comparison/02_route_preflight_contract/` without `complete.log`; predecessor index 03 is active at `agent-task/m-iop-one-shot-agent-model-comparison/03+01_rubric_version_contract/` without `complete.log`. Both are currently unsatisfied; do not implement until both exact completion files exist.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
- Do not modify fixture prompt/reference/image bytes, benchmark runtime code, scoring/reporting logic, caller adapters, `../iop-s2`, credentials, or external registrations.
|
||||
- Update only the project skill's scoring wording and matching contract assertions here; preserve the all-cell preflight wording delivered by predecessor 02.
|
||||
- Do not run scored `run`, `resume`, `score`, or `report`; this Epic ends at immutable input plus readiness preflight.
|
||||
- Do not use the generic example aliases, omit `repair`, lower efforts, or replace unavailable models/presets.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- `evaluation_mode=isolated-reassessment`; `finalizer=finalize-task-policy.sh`, `finalizer_mode=pair`.
|
||||
- Build closures: scope/context/verification/evidence/ownership/decision all closed. Scores `2+1+2+2+2=G09`; base/final route `grade-boundary`, lane `cloud`, catalog `worker/cloud/G09`, filename `PLAN-cloud-G09.md`.
|
||||
- Review closures: all closed. Scores `2+1+2+2+2=G09`; route `official-review`, lane `cloud`, catalog `review/cloud/G09`, filename `CODE_REVIEW-cloud-G09.md`.
|
||||
- `large_indivisible_context=false`; positive loop risks: `temporal_state`, `boundary_contract`, `structured_interpretation`, `variant_product` (`count=4`, risk boundary matched but grade remains the basis); `review_rework_count=0`; `evidence_integrity_failure=false`; no capability gap. External rebuild/registration is a named execution precondition, not a cloud-resolvable planning gap.
|
||||
|
||||
## Dependencies and Execution Order
|
||||
|
||||
1. Resolve exactly one active-sibling or matching archived `complete.log` for `02_route_preflight_contract`.
|
||||
2. Resolve exactly one active-sibling or matching archived `complete.log` for `03+01_rubric_version_contract`; its own index-01 dependency is already encoded there.
|
||||
3. Create and statically validate the locked manifest and regression test.
|
||||
4. Run the secret-safe external preflight once. Stop on its exact closed blocker; do not proceed to scored execution.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] [TEST-1] Add the exact tracked bench-02 manifest and a static regression that locks fixture checksum/version, viewports, rubric, policies, C01-C09 bindings, seed, and seeded order; atomically update project-skill scoring language and its contract test to the manifest-selected rubric catalog.
|
||||
- [ ] [TEST-2] Run the secret-safe caller/testbed/environment gate and public all-cell preflight; require nine ready results or record the exact blocker and resume condition without substitution.
|
||||
- [ ] [TEST-3] Run final schema, focused regression, full benchmark suite, and whitespace verification while leaving scored execution untouched.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [TEST-1] Create one exact immutable readiness manifest
|
||||
|
||||
**Problem**
|
||||
|
||||
The generic fixture manifest has only three homogeneous presets (`agent-comparison-benchmark-manifest.example.json:42-91`), while the direct preflight example has only C01-C05 equivalents (`agent-comparison-benchmark-direct-preflight.example.json:42-108`). Neither can satisfy S01-S03.
|
||||
|
||||
**Solution**
|
||||
|
||||
Create `scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json` with these immutable common fields:
|
||||
|
||||
```json
|
||||
{
|
||||
"pipeline_version": "2",
|
||||
"environment": "dev",
|
||||
"testbed": "../iop-s2",
|
||||
"execution_order_seed": "bench-02-c01-c09-v1",
|
||||
"repetitions": 1,
|
||||
"session_policy": "fresh",
|
||||
"setup_cache_policy": "isolated",
|
||||
"timeout": {"run_seconds": 300, "idle_seconds": 30, "quiet_seconds": 10, "cleanup_grace_seconds": 5},
|
||||
"rubric_version": "one-shot-agent-comparison-v1",
|
||||
"output_root": "agent-test/runs/bench-02"
|
||||
}
|
||||
```
|
||||
|
||||
Reuse the existing evaluator, fixture paths/checksum, and viewports byte-for-byte: `desktop_1080` at 1920x1080 and `mobile_375` at 375x812. Use the exact matrix below; stage order is selector, plan, work, review, repair. Cloud selector/plan/review/repair effort is `high`; `ornith-fast` work omits effort.
|
||||
|
||||
| ID | Cell id | Caller | Route kind/id | Request model/effort | Expected bindings |
|
||||
|----|---------|--------|---------------|----------------------|-------------------|
|
||||
| C01 | `c01-claude-sonnet-direct` | claude | direct / `claude-sonnet-5` | `claude-sonnet-5` / max | request Sonnet/max |
|
||||
| C02 | `c02-claude-gemini-direct` | claude | direct / `gemini-3.6-flash` | Gemini / high | request Gemini/high |
|
||||
| C03 | `c03-agy-gemini-direct` | agy | direct / `gemini-3.6-flash` | Gemini / high | request Gemini/high |
|
||||
| C04 | `c04-claude-gpt-direct` | claude | direct / `gpt-5.6-luna` | GPT / xhigh | request GPT/xhigh |
|
||||
| C05 | `c05-codex-gpt-direct` | codex | direct / `gpt-5.6-luna` | GPT / xhigh | request GPT/xhigh |
|
||||
| C06 | `c06-claude-gemini-hybrid` | claude | preset / `gemini-hybrid` | `gemini-hybrid` / high | Gemini/high, Gemini/high, ornith-fast, Gemini/high, Gemini/high |
|
||||
| C07 | `c07-agy-gemini-hybrid` | agy | preset / `gemini-hybrid` | `gemini-hybrid` / high | same as C06 |
|
||||
| C08 | `c08-claude-gpt-hybrid` | claude | preset / `gpt-hybrid` | `gpt-hybrid` / xhigh | GPT/high, GPT/high, ornith-fast, GPT/high, GPT/high |
|
||||
| C09 | `c09-codex-gpt-hybrid` | codex | preset / `gpt-hybrid` | `gpt-hybrid` / xhigh | same as C08 |
|
||||
|
||||
Add `test_iop_one_shot_manifest_locks_benchmark_readiness` to `manifest_test.py`. Assert exact fixture version/checksum/assets, exactly two image workspace paths, viewports, rubric, timeout/session/cache/repetition/output root, seed, complete cell payloads, digest shape, and loaded seeded order:
|
||||
|
||||
```text
|
||||
c02-claude-gemini-direct
|
||||
c05-codex-gpt-direct
|
||||
c03-agy-gemini-direct
|
||||
c06-claude-gemini-hybrid
|
||||
c08-claude-gpt-hybrid
|
||||
c09-codex-gpt-hybrid
|
||||
c01-claude-sonnet-direct
|
||||
c07-agy-gemini-hybrid
|
||||
c04-claude-gpt-direct
|
||||
```
|
||||
|
||||
In the same dependent change, replace the project skill's fixed `landing-quality-v1` scoring sentence with a requirement to use the exact immutable manifest-selected rubric. Name the closed supported set (`landing-quality-v1`, `one-shot-agent-comparison-v1`) and forbid fallback or reinterpretation. Extend `skill_contract_test.py` so its base contract requires that wording and both version literals, while a mutation back to fixed-legacy wording fails. Preserve predecessor 02's all-cell preflight assertions unchanged.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json`: add the exact locked readiness manifest.
|
||||
- [ ] `scripts/agent_benchmark/manifest_test.py`: add the exact static regression and recompute fixture checksum from resolved inputs.
|
||||
- [ ] `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md`: document manifest-selected exact scoring and the closed two-version rubric set.
|
||||
- [ ] `scripts/agent_benchmark/skill_contract_test.py`: require that scoring contract without weakening the all-cell preflight contract.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
Write the named manifest test. Compare complete nested cell payloads, not just counts. Recalculate the fixture checksum with production helpers and assert the manifest digest is `sha256:` plus 64 lowercase hex characters. In the skill contract test, assert both exact supported rubric versions and manifest selection, then mutation-test the legacy-only regression.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
python3 scripts/agent_comparison_benchmark.py validate --manifest scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test.ManifestValidationTest.test_iop_one_shot_manifest_locks_benchmark_readiness
|
||||
python3 -m unittest scripts.agent_benchmark.skill_contract_test
|
||||
```
|
||||
|
||||
Expected: the manifest validates, the exact lock test passes, and the skill contract suite accepts only manifest-selected rubric wording.
|
||||
|
||||
### [TEST-2] Produce one redacted nine-cell readiness record
|
||||
|
||||
**Problem**
|
||||
|
||||
Planning-time external preflight is blocked: benchmark environment references are missing, the dev runtime identity is unknown, and agy help lacks three adapter-required tokens. Static fixtures cannot claim live readiness.
|
||||
|
||||
**Solution**
|
||||
|
||||
First run dependency, clean-testbed, artifact, caller-version/help, and secret-reference checks. The environment check dereferences names without printing names or values. Then invoke only the public benchmark `preflight` command once. Parse its run id, require `status=ready ready=9 registration_required=0 implementation_gap=0`, and inspect the durable result list against the loaded manifest. Search durable bytes for exact runtime base URLs, dereferenced secrets, and raw config observation JSON without printing them.
|
||||
|
||||
If any setup check fails or preflight returns 69, record the exact safe output and resume condition in `CODE_REVIEW-cloud-G09.md` and stop. Do not mutate external state.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `CODE_REVIEW-cloud-G09.md`: record safe preflight output, run id, nine-result inspection, or exact blocker/resume evidence.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
No additional unit test; this is required external execution evidence. The public CLI and durable record are the acceptance oracle. Never call a caller/provider directly.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
set -euo pipefail
|
||||
python3 - <<'PY'
|
||||
from pathlib import Path
|
||||
root = Path("agent-task")
|
||||
group = "m-iop-one-shot-agent-model-comparison"
|
||||
for subtask in ("02_route_preflight_contract", "03+01_rubric_version_contract"):
|
||||
active = root / group / subtask / "complete.log"
|
||||
archived = sorted((root / "archive").glob(f"*/*/{group}/{subtask}/complete.log"))
|
||||
candidates = [path for path in (active, *archived) if path.is_file()]
|
||||
assert len(candidates) == 1, (subtask, candidates)
|
||||
print(f"ok: predecessor complete {candidates[0]}")
|
||||
PY
|
||||
test "$(git -C ../iop-s2 branch --show-current)" = "dev"
|
||||
test -z "$(git -C ../iop-s2 status --short)"
|
||||
testbed_head="$(git -C ../iop-s2 rev-parse HEAD)"
|
||||
test -n "$testbed_head"
|
||||
printf 'ok: testbed branch=dev head=%s clean=true\n' "$testbed_head"
|
||||
command -v readelf >/dev/null
|
||||
test -x ../iop-s2/build/bin/iop-edge
|
||||
test -x ../iop-s2/build/dev/iop-node
|
||||
test -f ../iop-s2/configs/edge.yaml
|
||||
commit_epoch="$(git -C ../iop-s2 show -s --format=%ct HEAD)"
|
||||
for binary in ../iop-s2/build/bin/iop-edge ../iop-s2/build/dev/iop-node; do
|
||||
test "$(stat -c %Y "$binary")" -ge "$commit_epoch"
|
||||
readelf -h "$binary" | rg 'Machine:\s+AArch64' >/dev/null
|
||||
done
|
||||
../iop-s2/build/bin/iop-edge --help >/dev/null
|
||||
test -n "$(../iop-s2/build/dev/iop-node version)"
|
||||
printf 'ok: current-HEAD Linux AArch64 Edge/Node artifacts are executable\n'
|
||||
claude --version
|
||||
agy --version
|
||||
codex --version
|
||||
python3 - <<'PY'
|
||||
import subprocess
|
||||
from scripts.agent_benchmark.agy_iop import AGY_KNOWN_VERSION, inspect_agy_iop_capability
|
||||
version_run = subprocess.run(["agy", "--version"], check=True, capture_output=True, text=True)
|
||||
help_run = subprocess.run(["agy", "--help"], check=True, capture_output=True, text=True)
|
||||
version = (version_run.stdout + version_run.stderr).strip()
|
||||
help_text = help_run.stdout + help_run.stderr
|
||||
capability = inspect_agy_iop_capability(version, help_text)
|
||||
assert capability.version == AGY_KNOWN_VERSION, (capability.version, AGY_KNOWN_VERSION)
|
||||
assert capability.iop_transport_supported, capability
|
||||
print(f"ok: adapter-known agy {AGY_KNOWN_VERSION} transport is documented")
|
||||
PY
|
||||
python3 - <<'PY'
|
||||
import json, os, re
|
||||
name = re.compile(r"^[A-Za-z_][A-Za-z0-9_]*$")
|
||||
for caller in ("CLAUDE", "AGY", "CODEX"):
|
||||
assert os.environ.get(f"IOP_BENCH_{caller}_BASE_URL")
|
||||
ref = os.environ.get(f"IOP_BENCH_{caller}_SECRET_ENV", "")
|
||||
assert name.fullmatch(ref) and os.environ.get(ref)
|
||||
config_ref = os.environ.get("IOP_BENCH_CONFIG_OBSERVATION_ENV", "")
|
||||
assert name.fullmatch(config_ref) and os.environ.get(config_ref)
|
||||
value = json.loads(os.environ[config_ref])
|
||||
assert value.get("schema_version") == "1" and isinstance(value.get("routes"), list)
|
||||
print("ok: benchmark environment references present")
|
||||
PY
|
||||
preflight_output="$(python3 scripts/agent_comparison_benchmark.py preflight --manifest scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json)"
|
||||
printf '%s\n' "$preflight_output"
|
||||
case "$preflight_output" in *"status=ready ready=9 registration_required=0 implementation_gap=0"*) ;; *) exit 1 ;; esac
|
||||
run_id="${preflight_output#*run_id=}"; run_id="${run_id%% *}"
|
||||
python3 - "$run_id" <<'PY'
|
||||
import json, os, sys
|
||||
from pathlib import Path
|
||||
from scripts.agent_benchmark.manifest import load_manifest
|
||||
manifest_path = Path("scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json")
|
||||
manifest = load_manifest(manifest_path, repo_root=Path.cwd())
|
||||
run_root = Path(manifest.output_root) / sys.argv[1]
|
||||
record = json.loads((run_root / "preflight/preflight-000001.json").read_text(encoding="ascii"))
|
||||
assert record["status"] == "ready"
|
||||
assert [item["cell"]["id"] for item in record["results"]] == [cell.id for cell in manifest.matrix]
|
||||
assert len(record["results"]) == 9 and all(item["status"] == "ready" for item in record["results"])
|
||||
sensitive = []
|
||||
for caller in ("CLAUDE", "AGY", "CODEX"):
|
||||
sensitive.append(os.environ[f"IOP_BENCH_{caller}_BASE_URL"].encode())
|
||||
sensitive.append(os.environ[os.environ[f"IOP_BENCH_{caller}_SECRET_ENV"]].encode())
|
||||
config_ref = os.environ["IOP_BENCH_CONFIG_OBSERVATION_ENV"]
|
||||
sensitive.append(os.environ[config_ref].encode())
|
||||
durable = b"".join(path.read_bytes() for path in run_root.rglob("*") if path.is_file())
|
||||
assert all(value and value not in durable for value in sensitive)
|
||||
print("ok: nine ready results are manifest-bound and runtime values are absent")
|
||||
PY
|
||||
```
|
||||
|
||||
Expected: every setup command exits 0; preflight prints one ready summary with nine results; durable inspection prints its safe success line. At planning time this block is expected to stop at agy/environment setup until the recorded resume condition is satisfied.
|
||||
|
||||
### [TEST-3] Preserve the full benchmark baseline
|
||||
|
||||
**Problem**
|
||||
|
||||
The final file consumes the new order, rubric, and all-cell contracts. It must not regress legacy fixtures or other benchmark behavior.
|
||||
|
||||
**Solution**
|
||||
|
||||
Validate all four manifests, run the focused static test and complete rubric/scoring/manifest/attempt/connectivity suites, and check the diff.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `CODE_REVIEW-cloud-G09.md`: record final command output after live readiness succeeds; if TEST-2 is blocked, leave this item unchecked and record the resume condition.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
No additional files beyond TEST-1. Existing suites are the regression oracle; external preflight output is not cached.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
for manifest in scripts/fixtures/agent-comparison-benchmark-manifest.example.json scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json; do python3 scripts/agent_comparison_benchmark.py validate --manifest "$manifest"; done
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.scoring_test scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test scripts.agent_benchmark.skill_contract_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Expected: four manifests validate, all suites report `OK`, and whitespace check is silent.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Item |
|
||||
|------|------|
|
||||
| `scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json` | TEST-1 |
|
||||
| `scripts/agent_benchmark/manifest_test.py` | TEST-1 |
|
||||
| `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md` | TEST-1 |
|
||||
| `scripts/agent_benchmark/skill_contract_test.py` | TEST-1 |
|
||||
| `agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/CODE_REVIEW-cloud-G09.md` | TEST-2, TEST-3 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
Run from `/config/workspace/iop-s0` in dependency order. Fresh output is required; cached output is not acceptable. Run the full TEST-2 command block exactly once after its setup checks pass, then run:
|
||||
|
||||
```bash
|
||||
python3 scripts/agent_comparison_benchmark.py validate --manifest scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test.ManifestValidationTest.test_iop_one_shot_manifest_locks_benchmark_readiness
|
||||
python3 -m unittest scripts.agent_benchmark.skill_contract_test
|
||||
for manifest in scripts/fixtures/agent-comparison-benchmark-manifest.example.json scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json; do python3 scripts/agent_comparison_benchmark.py validate --manifest "$manifest"; done
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.scoring_test scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test scripts.agent_benchmark.skill_contract_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Expected: static validation/test pass, TEST-2 has one nine-ready durable preflight record with no runtime values, four manifests validate, all suites report `OK`, and `git diff --check` is silent. No scored benchmark command is run.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,209 @@
|
|||
<!-- task=m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest plan=0 tag=TEST milestone-task=fixture-lock,route-readiness,matrix-lock -->
|
||||
|
||||
# Code Review Reference - TEST
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-12
|
||||
task=m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest, plan=0, tag=TEST
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files. Run the applicable verification commands directly and record fresh output in `Verification Results`; implementation-owned output is handoff evidence, not a substitute for reviewer verification. If implementation is present, repair missing or stale verification output instead of failing solely for insufficient recorded evidence. When verification exposes a defect, collect the necessary data, determine the exact root cause, and select one concrete fix before generating the follow-up plan; never delegate investigation or remedy selection to the worker.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals.
|
||||
2. Archive `CODE_REVIEW-cloud-G07.md` → `code_review_cloud_G07_0.log` and `PLAN-local-G07.md` → `plan_local_G07_0.log`.
|
||||
3. If PASS, write `complete.log` and move active task directory to `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/`. If WARN/FAIL, fully write the next filesystem state required by the code-review skill.
|
||||
4. If PASS, preserve the first-line `milestone-task` metadata in `complete.log` and report it for runtime aggregation. Roadmap state evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check applicable `Review-Only Checklist` items at the final `.log` location before reporting.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|---------|
|
||||
| TEST-1 Create one exact immutable readiness manifest | [ ] |
|
||||
| TEST-2 Produce one redacted nine-cell readiness record | [ ] |
|
||||
| TEST-3 Preserve the full benchmark baseline | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] [TEST-1] Add the exact tracked bench-02 manifest and a static regression that locks fixture checksum/version, viewports, rubric, policies, C01-C09 bindings, seed, and seeded order.
|
||||
- [ ] [TEST-2] Run the secret-safe caller/testbed/environment gate and public all-cell preflight; require nine ready results or record the exact blocker and resume condition without substitution.
|
||||
- [ ] [TEST-3] Run final schema, focused regression, full benchmark suite, and whitespace verification while leaving scored execution untouched.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Run applicable required verification and record fresh command/output; repair reviewer-reconstructable evidence gaps instead of forwarding them to another plan.
|
||||
- [ ] For every Required/Suggested finding, record reviewer-collected `Evidence`, exact `Root Cause`, and one `Selected Fix` with affected files/symbols/tests and acceptance commands before creating a follow-up plan.
|
||||
- [ ] Archive active `CODE_REVIEW-cloud-G07.md` to `code_review_cloud_G07_0.log`.
|
||||
- [ ] Archive active `PLAN-local-G07.md` to `plan_local_G07_0.log`.
|
||||
- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write `complete.log` based on `agent-ops/skills/common/code-review/templates/complete-log-template.md` and leave no active `.md` files.
|
||||
- [ ] If PASS, move active task directory `agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/` to `agent-task/archive/YYYY/MM/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/` and update this checklist at the final archive path.
|
||||
- [ ] If PASS, preserve and report `milestone-task` metadata for runtime aggregation, without modifying roadmap or directly calling `update-roadmap`.
|
||||
- [ ] If PASS for split work, remove empty active parent `agent-task/m-iop-one-shot-agent-model-comparison/` or verify it was kept due to remaining siblings/files.
|
||||
- [ ] If WARN/FAIL, write the next filesystem state matching code-review verdict and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record any deviations from the plan and the rationale here._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record key design decisions here._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Confirm both predecessor completion files exist before reviewing implementation.
|
||||
- Confirm the complete C01-C09 map, seed-derived order, fixture checksum, two images, policies, and new rubric version are exact.
|
||||
- Confirm hybrid routes include `repair`, cloud stages use high, and work uses `ornith-fast` with no invented effort.
|
||||
- Confirm live readiness came only from public preflight and durable evidence contains no endpoint, secret, or raw config value.
|
||||
- Confirm no scored run/score/report command was executed in this Epic.
|
||||
|
||||
## Verification Results
|
||||
|
||||
Record actual stdout/stderr under each command. If a command changes, document the replacement and reason in `Deviations from Plan`. TEST-2's full command block is fixed in the plan and must be copied with its actual output or exact blocker here.
|
||||
|
||||
### Dependency Verification
|
||||
|
||||
```bash
|
||||
python3 - <<'PY'
|
||||
from pathlib import Path
|
||||
root = Path("agent-task")
|
||||
group = "m-iop-one-shot-agent-model-comparison"
|
||||
for subtask in ("02_route_preflight_contract", "03+01_rubric_version_contract"):
|
||||
active = root / group / subtask / "complete.log"
|
||||
archived = sorted((root / "archive").glob(f"*/*/{group}/{subtask}/complete.log"))
|
||||
candidates = [path for path in (active, *archived) if path.is_file()]
|
||||
assert len(candidates) == 1, (subtask, candidates)
|
||||
print(f"ok: predecessor complete {candidates[0]}")
|
||||
PY
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
### TEST-1 Verification
|
||||
|
||||
```bash
|
||||
python3 scripts/agent_comparison_benchmark.py validate --manifest scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test.ManifestValidationTest.test_iop_one_shot_manifest_locks_benchmark_readiness
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
### TEST-2 External Preflight
|
||||
|
||||
```bash
|
||||
set -euo pipefail
|
||||
python3 - <<'PY'
|
||||
from pathlib import Path
|
||||
root = Path("agent-task")
|
||||
group = "m-iop-one-shot-agent-model-comparison"
|
||||
for subtask in ("02_route_preflight_contract", "03+01_rubric_version_contract"):
|
||||
active = root / group / subtask / "complete.log"
|
||||
archived = sorted((root / "archive").glob(f"*/*/{group}/{subtask}/complete.log"))
|
||||
candidates = [path for path in (active, *archived) if path.is_file()]
|
||||
assert len(candidates) == 1, (subtask, candidates)
|
||||
print(f"ok: predecessor complete {candidates[0]}")
|
||||
PY
|
||||
test "$(git -C ../iop-s2 branch --show-current)" = "dev"
|
||||
test -z "$(git -C ../iop-s2 status --short)"
|
||||
test -x ../iop-s2/build/bin/iop-edge
|
||||
test -x ../iop-s2/build/dev/iop-node
|
||||
test -f ../iop-s2/configs/edge.yaml
|
||||
claude --version
|
||||
agy --version
|
||||
codex --version
|
||||
agy_help="$(agy --help 2>&1)"; for token in --print --output-format --sandbox --model --effort AGY_PROVIDER AGY_OPENAI_BASE_URL AGY_OPENAI_API_KEY stream-json; do grep -F -- "$token" <<<"$agy_help" >/dev/null; done
|
||||
python3 - <<'PY'
|
||||
import json, os, re
|
||||
name = re.compile(r"^[A-Za-z_][A-Za-z0-9_]*$")
|
||||
for caller in ("CLAUDE", "AGY", "CODEX"):
|
||||
assert os.environ.get(f"IOP_BENCH_{caller}_BASE_URL")
|
||||
ref = os.environ.get(f"IOP_BENCH_{caller}_SECRET_ENV", "")
|
||||
assert name.fullmatch(ref) and os.environ.get(ref)
|
||||
config_ref = os.environ.get("IOP_BENCH_CONFIG_OBSERVATION_ENV", "")
|
||||
assert name.fullmatch(config_ref) and os.environ.get(config_ref)
|
||||
value = json.loads(os.environ[config_ref])
|
||||
assert value.get("schema_version") == "1" and isinstance(value.get("routes"), list)
|
||||
print("ok: benchmark environment references present")
|
||||
PY
|
||||
preflight_output="$(python3 scripts/agent_comparison_benchmark.py preflight --manifest scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json)"
|
||||
printf '%s\n' "$preflight_output"
|
||||
case "$preflight_output" in *"status=ready ready=9 registration_required=0 implementation_gap=0"*) ;; *) exit 1 ;; esac
|
||||
run_id="${preflight_output#*run_id=}"; run_id="${run_id%% *}"
|
||||
python3 - "$run_id" <<'PY'
|
||||
import json, os, sys
|
||||
from pathlib import Path
|
||||
from scripts.agent_benchmark.manifest import load_manifest
|
||||
manifest_path = Path("scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json")
|
||||
manifest = load_manifest(manifest_path, repo_root=Path.cwd())
|
||||
run_root = Path(manifest.output_root) / sys.argv[1]
|
||||
record = json.loads((run_root / "preflight/preflight-000001.json").read_text(encoding="ascii"))
|
||||
assert record["status"] == "ready"
|
||||
assert [item["cell"]["id"] for item in record["results"]] == [cell.id for cell in manifest.matrix]
|
||||
assert len(record["results"]) == 9 and all(item["status"] == "ready" for item in record["results"])
|
||||
sensitive = []
|
||||
for caller in ("CLAUDE", "AGY", "CODEX"):
|
||||
sensitive.append(os.environ[f"IOP_BENCH_{caller}_BASE_URL"].encode())
|
||||
sensitive.append(os.environ[os.environ[f"IOP_BENCH_{caller}_SECRET_ENV"]].encode())
|
||||
config_ref = os.environ["IOP_BENCH_CONFIG_OBSERVATION_ENV"]
|
||||
sensitive.append(os.environ[config_ref].encode())
|
||||
durable = b"".join(path.read_bytes() for path in run_root.rglob("*") if path.is_file())
|
||||
assert all(value and value not in durable for value in sensitive)
|
||||
print("ok: nine ready results are manifest-bound and runtime values are absent")
|
||||
PY
|
||||
```
|
||||
|
||||
_Actual output or exact blocker and resume condition:_
|
||||
|
||||
### TEST-3 and Final Verification
|
||||
|
||||
```bash
|
||||
for manifest in scripts/fixtures/agent-comparison-benchmark-manifest.example.json scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json; do python3 scripts/agent_comparison_benchmark.py validate --manifest "$manifest"; done
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.scoring_test scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
_Actual output:_
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these (archive, complete.log, and task-directory archive move are review-agent only) |
|
||||
| Archive Evidence Snapshot | Fixed at stub creation from plan when present | Implementing agent uses it as default prior-loop context; read only the specific archive files cited there when more detail is required |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results (section headings + commands) | Implementing agent, then review agent | Implementing agent records initial output; review agent reruns applicable commands and may fill, replace, or append fresh verified output before verdict. Implementing-agent command changes require a `Deviations from Plan` entry |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,320 @@
|
|||
<!-- task=m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest plan=0 tag=TEST milestone-task=fixture-lock,route-readiness,matrix-lock -->
|
||||
|
||||
# Plan - Locked C01-C09 Benchmark Readiness Manifest
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Do not start until the dependency command resolves exactly one active or archived `complete.log` for both predecessor indices 02 and 03. Filling the implementation-owned sections in `CODE_REVIEW-cloud-G07.md` is the mandatory last implementation step. Run every verification command, record actual notes and stdout/stderr, keep both active files in place, and report ready for review; only the code-review skill may finalize or archive this task. If blocked, record only the exact blocker, attempted commands/output, and resume condition in the implementation-owned evidence fields. Do not ask the user, call user-input tools, create control-plane stop files, classify the next state, archive logs, or write `complete.log`.
|
||||
|
||||
## Background
|
||||
|
||||
The repository already contains the reusable two-image vanilla landing-page fixture, but no immutable manifest combines it with the approved C01-C09 matrix, explicit seed, new rubric version, and exact live route evidence. This packet creates that final tracked input and runs the public preflight without substituting missing callers, credentials, models, or presets.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `scripts/agent_benchmark/manifest.py`
|
||||
- `scripts/agent_benchmark/manifest_test.py`
|
||||
- `scripts/agent_benchmark/attempts.py`
|
||||
- `scripts/agent_benchmark/attempts_test.py`
|
||||
- `scripts/agent_benchmark/live_iop.py`
|
||||
- `scripts/agent_benchmark/connectivity_integration_test.py`
|
||||
- `scripts/agent_benchmark/rubric.py`
|
||||
- `scripts/agent_benchmark/rubric_test.py`
|
||||
- `scripts/agent_benchmark/scoring.py`
|
||||
- `scripts/agent_benchmark/scoring_test.py`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-manifest.example.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json`
|
||||
- `scripts/fixtures/agent-comparison-benchmark/prompt.md`
|
||||
- `scripts/fixtures/agent-comparison-benchmark/reference.txt`
|
||||
- `scripts/fixtures/agent-comparison-benchmark/aurora-grid.svg`
|
||||
- `scripts/fixtures/agent-comparison-benchmark/orbit-rings.svg`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/iop-one-shot-agent-model-comparison.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/iop-one-shot-agent-model-comparison/SDD.md`
|
||||
- `agent-spec/testing/agent-comparison-benchmark.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- SDD: `agent-roadmap/sdd/knowledge-tool-optimization-extension/iop-one-shot-agent-model-comparison/SDD.md`, status `[승인됨]`, implementation lock released, no unresolved user review.
|
||||
- First-line scope: `milestone-task=fixture-lock,route-readiness,matrix-lock`.
|
||||
- S01 requires identical prompt/assets/workspace/viewports/rubric checksum/version; S02 requires a redacted all-cell preflight or exact blocker; S03 requires the immutable nine-cell manifest, repetitions 1, explicit seed, fresh/isolated policy, timeout, and bindings.
|
||||
- Evidence Map rows S01-S03 drive the static manifest assertions, secret-safe external preflight, and durable nine-result inspection below.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- No separate handoff was supplied. Repository-native evidence came from the complete fixture, manifest/schema, public CLI, adapter/store tests, SDD, spec, and contracts listed above.
|
||||
- Fresh baseline at starting HEAD `b197e5db70637f87017a024a847e3e53fdc72e8b`: all three shipped manifests validated; `python3 -m unittest scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test` ran 178 tests and passed.
|
||||
- Fixture evidence: prompt requires exactly root `index.html`, `styles.css`, `script.js`, vanilla HTML/CSS/JS, both local images, responsive desktop/mobile behavior, accessibility, and no external dependency. The assets resolve to exactly two images plus one reference file; checksum is `sha256:7dc1be6ed4a9f2f873016b708b99d827b0249c74f2ac287e1fcf8deade8dcd98`. Viewports are 1920x1080 and 375x812.
|
||||
- The direct example already fixes C01-C05 model/effort pairs. The approved SDD names the hybrid scenarios `gemini-hybrid` and `gpt-hybrid`; this packet uses those exact strings as both virtual request model and preset route id. Missing registration must remain `registration_required`, never fallback.
|
||||
- Planned seeded order for `bench-02-c01-c09-v1` under the reviewed domain-separated algorithm is: C02, C05, C03, C06, C08, C09, C01, C07, C04.
|
||||
- Fresh output is required; cached output is not accepted.
|
||||
|
||||
#### External Verification Preflight
|
||||
|
||||
- Runner/workdir: local Linux `aarch64`, `/config/workspace/iop-s0`.
|
||||
- Testbed: `/config/workspace/iop-s2`, branch `dev`, observed HEAD `1f2f7f...`, clean at planning time. Source sync beyond the clean dev checkout was not asserted.
|
||||
- Artifacts: `../iop-s2/build/bin/iop-edge`, `../iop-s2/build/dev/iop-node`, and `../iop-s2/configs/edge.yaml` exist. The tracked example config exposes only local example models and no benchmark route/preset catalog; do not treat it as live registration.
|
||||
- Callers: Claude Code `2.1.227`, agy `1.1.12`, Codex CLI `0.147.0` were installed. Claude/Codex help probes completed. Current agy help contains `--print`, `--output-format`, `--sandbox`, `--model`, `--effort`, and `stream-json`, but lacks the adapter-required `AGY_PROVIDER`, `AGY_OPENAI_BASE_URL`, and `AGY_OPENAI_API_KEY` tokens.
|
||||
- Runtime/config: every `IOP_BENCH_{CLAUDE,AGY,CODEX}_{BASE_URL,SECRET_ENV}` and `IOP_BENCH_CONFIG_OBSERVATION_ENV` was missing. Consequently runtime identity, endpoint hosts/ports, live catalog, auth, and route registration could not be observed. `ss` is unavailable; the public preflight's endpoint probe is the authoritative reachability check once URLs exist.
|
||||
- Exact resume condition: an operator must provide all benchmark env references and their non-empty referenced secret/config values, register all direct/hybrid routes with exact bindings, run the dev Edge/Node endpoints, and install an agy build whose help satisfies the current adapter capability tokens. No raw value is written to the repository or review file.
|
||||
- If the public preflight returns `registration_required` or `implementation_gap`, preserve its run record, record only the closed summary and run id in the review evidence, and stop. Do not edit `../iop-s2`, install tools, substitute models, or invoke callers outside the benchmark CLI.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
- No shipped manifest contains exactly C01-C09 with one explicit seed and the approved new rubric version.
|
||||
- Existing generic preset examples do not express Gemini/GPT plan→ornith-fast work→review/repair stage bindings.
|
||||
- No static regression asserts the fixture checksum, exact two-image paths, policy fields, cell map, seeded order, and hybrid binding table together.
|
||||
- Live readiness is currently blocked by missing environment/config and incompatible agy help evidence; the plan contains the exact secret-safe resume/preflight command.
|
||||
|
||||
### Symbol References
|
||||
|
||||
- No symbol is renamed or removed.
|
||||
- The new fixture path is consumed by `load_manifest`, public CLI commands, `RunStore`, and the manifest regression test; no new API is introduced.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
- Direct-small classification produced zero code changes. The existing fixture bytes already meet the prompt/image/workspace requirements, but their final version/checksum evidence is inseparable from the planned seeded/rubric manifest and all-cell preflight. Editing a partial manifest before those contracts land would collide with planned work.
|
||||
- This is dependent child `04+02,03_locked_benchmark_manifest`. Its stable invariant is one tracked manifest whose fixture, policies, nine exact cells, order, rubric, and durable preflight all agree.
|
||||
- Predecessor index 02 is active at `agent-task/m-iop-one-shot-agent-model-comparison/02_route_preflight_contract/` without `complete.log`; predecessor index 03 is active at `agent-task/m-iop-one-shot-agent-model-comparison/03+01_rubric_version_contract/` without `complete.log`. Both are currently unsatisfied; do not implement until both exact completion files exist.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
- Do not modify fixture prompt/reference/image bytes, benchmark runtime code, scoring/reporting logic, caller adapters, `../iop-s2`, credentials, or external registrations.
|
||||
- Do not run scored `run`, `resume`, `score`, or `report`; this Epic ends at immutable input plus readiness preflight.
|
||||
- Do not use the generic example aliases, omit `repair`, lower efforts, or replace unavailable models/presets.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- `evaluation_mode=first-pass`; `finalizer=finalize-task-policy.sh`, `finalizer_mode=pair`.
|
||||
- Build closures: scope/context/verification/evidence/ownership/decision all closed. Scores `1+1+1+2+2=G07`; base/final route `local-fit`, lane `local`, catalog `worker/local/G07`, filename `PLAN-local-G07.md`.
|
||||
- Review closures: all closed. Scores `1+1+1+2+2=G07`; route `official-review`, lane `cloud`, catalog `review/cloud/G07`, filename `CODE_REVIEW-cloud-G07.md`.
|
||||
- `large_indivisible_context=false`; positive loop risks: `boundary_contract`, `structured_interpretation`, `variant_product` (`count=3`); `review_rework_count=0`; `evidence_integrity_failure=false`; no capability gap. External registration is a named execution precondition, not a cloud-resolvable planning gap.
|
||||
|
||||
## Dependencies and Execution Order
|
||||
|
||||
1. Resolve exactly one active-sibling or matching archived `complete.log` for `02_route_preflight_contract`.
|
||||
2. Resolve exactly one active-sibling or matching archived `complete.log` for `03+01_rubric_version_contract`; its own index-01 dependency is already encoded there.
|
||||
3. Create and statically validate the locked manifest and regression test.
|
||||
4. Run the secret-safe external preflight once. Stop on its exact closed blocker; do not proceed to scored execution.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] [TEST-1] Add the exact tracked bench-02 manifest and a static regression that locks fixture checksum/version, viewports, rubric, policies, C01-C09 bindings, seed, and seeded order.
|
||||
- [ ] [TEST-2] Run the secret-safe caller/testbed/environment gate and public all-cell preflight; require nine ready results or record the exact blocker and resume condition without substitution.
|
||||
- [ ] [TEST-3] Run final schema, focused regression, full benchmark suite, and whitespace verification while leaving scored execution untouched.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [TEST-1] Create one exact immutable readiness manifest
|
||||
|
||||
**Problem**
|
||||
|
||||
The generic fixture manifest has only three homogeneous presets (`agent-comparison-benchmark-manifest.example.json:42-91`), while the direct preflight example has only C01-C05 equivalents (`agent-comparison-benchmark-direct-preflight.example.json:42-108`). Neither can satisfy S01-S03.
|
||||
|
||||
**Solution**
|
||||
|
||||
Create `scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json` with these immutable common fields:
|
||||
|
||||
```json
|
||||
{
|
||||
"pipeline_version": "2",
|
||||
"environment": "dev",
|
||||
"testbed": "../iop-s2",
|
||||
"execution_order_seed": "bench-02-c01-c09-v1",
|
||||
"repetitions": 1,
|
||||
"session_policy": "fresh",
|
||||
"setup_cache_policy": "isolated",
|
||||
"timeout": {"run_seconds": 300, "idle_seconds": 30, "quiet_seconds": 10, "cleanup_grace_seconds": 5},
|
||||
"rubric_version": "one-shot-agent-comparison-v1",
|
||||
"output_root": "agent-test/runs/bench-02"
|
||||
}
|
||||
```
|
||||
|
||||
Reuse the existing evaluator, fixture paths/checksum, and desktop/mobile viewports byte-for-byte. Use the exact matrix below; stage order is selector, plan, work, review, repair. Cloud selector/plan/review/repair effort is `high`; `ornith-fast` work omits effort.
|
||||
|
||||
| ID | Cell id | Caller | Route kind/id | Request model/effort | Expected bindings |
|
||||
|----|---------|--------|---------------|----------------------|-------------------|
|
||||
| C01 | `c01-claude-sonnet-direct` | claude | direct / `claude-sonnet-5` | `claude-sonnet-5` / max | request Sonnet/max |
|
||||
| C02 | `c02-claude-gemini-direct` | claude | direct / `gemini-3.6-flash` | Gemini / high | request Gemini/high |
|
||||
| C03 | `c03-agy-gemini-direct` | agy | direct / `gemini-3.6-flash` | Gemini / high | request Gemini/high |
|
||||
| C04 | `c04-claude-gpt-direct` | claude | direct / `gpt-5.6-luna` | GPT / xhigh | request GPT/xhigh |
|
||||
| C05 | `c05-codex-gpt-direct` | codex | direct / `gpt-5.6-luna` | GPT / xhigh | request GPT/xhigh |
|
||||
| C06 | `c06-claude-gemini-hybrid` | claude | preset / `gemini-hybrid` | `gemini-hybrid` / high | Gemini/high, Gemini/high, ornith-fast, Gemini/high, Gemini/high |
|
||||
| C07 | `c07-agy-gemini-hybrid` | agy | preset / `gemini-hybrid` | `gemini-hybrid` / high | same as C06 |
|
||||
| C08 | `c08-claude-gpt-hybrid` | claude | preset / `gpt-hybrid` | `gpt-hybrid` / xhigh | GPT/high, GPT/high, ornith-fast, GPT/high, GPT/high |
|
||||
| C09 | `c09-codex-gpt-hybrid` | codex | preset / `gpt-hybrid` | `gpt-hybrid` / xhigh | same as C08 |
|
||||
|
||||
Add `test_iop_one_shot_manifest_locks_benchmark_readiness` to `manifest_test.py`. Assert exact fixture version/checksum/assets, exactly two image workspace paths, viewports, rubric, timeout/session/cache/repetition/output root, seed, complete cell payloads, digest shape, and loaded seeded order:
|
||||
|
||||
```text
|
||||
c02-claude-gemini-direct
|
||||
c05-codex-gpt-direct
|
||||
c03-agy-gemini-direct
|
||||
c06-claude-gemini-hybrid
|
||||
c08-claude-gpt-hybrid
|
||||
c09-codex-gpt-hybrid
|
||||
c01-claude-sonnet-direct
|
||||
c07-agy-gemini-hybrid
|
||||
c04-claude-gpt-direct
|
||||
```
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json`: add the exact locked readiness manifest.
|
||||
- [ ] `scripts/agent_benchmark/manifest_test.py`: add the exact static regression and recompute fixture checksum from resolved inputs.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
Write the named test. Compare complete nested cell payloads, not just counts. Recalculate the fixture checksum with production helpers and assert the manifest digest is `sha256:` plus 64 lowercase hex characters.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
python3 scripts/agent_comparison_benchmark.py validate --manifest scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test.ManifestValidationTest.test_iop_one_shot_manifest_locks_benchmark_readiness
|
||||
```
|
||||
|
||||
Expected: the manifest validates and the exact lock test passes.
|
||||
|
||||
### [TEST-2] Produce one redacted nine-cell readiness record
|
||||
|
||||
**Problem**
|
||||
|
||||
Planning-time external preflight is blocked: benchmark environment references are missing, the dev runtime identity is unknown, and agy help lacks three adapter-required tokens. Static fixtures cannot claim live readiness.
|
||||
|
||||
**Solution**
|
||||
|
||||
First run dependency, clean-testbed, artifact, caller-version/help, and secret-reference checks. The environment check dereferences names without printing names or values. Then invoke only the public benchmark `preflight` command once. Parse its run id, require `status=ready ready=9 registration_required=0 implementation_gap=0`, and inspect the durable result list against the loaded manifest. Search durable bytes for exact runtime base URLs, dereferenced secrets, and raw config observation JSON without printing them.
|
||||
|
||||
If any setup check fails or preflight returns 69, record the exact safe output and resume condition in `CODE_REVIEW-cloud-G07.md` and stop. Do not mutate external state.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `CODE_REVIEW-cloud-G07.md`: record safe preflight output, run id, nine-result inspection, or exact blocker/resume evidence.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
No additional unit test; this is required external execution evidence. The public CLI and durable record are the acceptance oracle. Never call a caller/provider directly.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
set -euo pipefail
|
||||
python3 - <<'PY'
|
||||
from pathlib import Path
|
||||
root = Path("agent-task")
|
||||
group = "m-iop-one-shot-agent-model-comparison"
|
||||
for subtask in ("02_route_preflight_contract", "03+01_rubric_version_contract"):
|
||||
active = root / group / subtask / "complete.log"
|
||||
archived = sorted((root / "archive").glob(f"*/*/{group}/{subtask}/complete.log"))
|
||||
candidates = [path for path in (active, *archived) if path.is_file()]
|
||||
assert len(candidates) == 1, (subtask, candidates)
|
||||
print(f"ok: predecessor complete {candidates[0]}")
|
||||
PY
|
||||
test "$(git -C ../iop-s2 branch --show-current)" = "dev"
|
||||
test -z "$(git -C ../iop-s2 status --short)"
|
||||
test -x ../iop-s2/build/bin/iop-edge
|
||||
test -x ../iop-s2/build/dev/iop-node
|
||||
test -f ../iop-s2/configs/edge.yaml
|
||||
claude --version
|
||||
agy --version
|
||||
codex --version
|
||||
agy_help="$(agy --help 2>&1)"; for token in --print --output-format --sandbox --model --effort AGY_PROVIDER AGY_OPENAI_BASE_URL AGY_OPENAI_API_KEY stream-json; do grep -F -- "$token" <<<"$agy_help" >/dev/null; done
|
||||
python3 - <<'PY'
|
||||
import json, os, re
|
||||
name = re.compile(r"^[A-Za-z_][A-Za-z0-9_]*$")
|
||||
for caller in ("CLAUDE", "AGY", "CODEX"):
|
||||
assert os.environ.get(f"IOP_BENCH_{caller}_BASE_URL")
|
||||
ref = os.environ.get(f"IOP_BENCH_{caller}_SECRET_ENV", "")
|
||||
assert name.fullmatch(ref) and os.environ.get(ref)
|
||||
config_ref = os.environ.get("IOP_BENCH_CONFIG_OBSERVATION_ENV", "")
|
||||
assert name.fullmatch(config_ref) and os.environ.get(config_ref)
|
||||
value = json.loads(os.environ[config_ref])
|
||||
assert value.get("schema_version") == "1" and isinstance(value.get("routes"), list)
|
||||
print("ok: benchmark environment references present")
|
||||
PY
|
||||
preflight_output="$(python3 scripts/agent_comparison_benchmark.py preflight --manifest scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json)"
|
||||
printf '%s\n' "$preflight_output"
|
||||
case "$preflight_output" in *"status=ready ready=9 registration_required=0 implementation_gap=0"*) ;; *) exit 1 ;; esac
|
||||
run_id="${preflight_output#*run_id=}"; run_id="${run_id%% *}"
|
||||
python3 - "$run_id" <<'PY'
|
||||
import json, os, sys
|
||||
from pathlib import Path
|
||||
from scripts.agent_benchmark.manifest import load_manifest
|
||||
manifest_path = Path("scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json")
|
||||
manifest = load_manifest(manifest_path, repo_root=Path.cwd())
|
||||
run_root = Path(manifest.output_root) / sys.argv[1]
|
||||
record = json.loads((run_root / "preflight/preflight-000001.json").read_text(encoding="ascii"))
|
||||
assert record["status"] == "ready"
|
||||
assert [item["cell"]["id"] for item in record["results"]] == [cell.id for cell in manifest.matrix]
|
||||
assert len(record["results"]) == 9 and all(item["status"] == "ready" for item in record["results"])
|
||||
sensitive = []
|
||||
for caller in ("CLAUDE", "AGY", "CODEX"):
|
||||
sensitive.append(os.environ[f"IOP_BENCH_{caller}_BASE_URL"].encode())
|
||||
sensitive.append(os.environ[os.environ[f"IOP_BENCH_{caller}_SECRET_ENV"]].encode())
|
||||
config_ref = os.environ["IOP_BENCH_CONFIG_OBSERVATION_ENV"]
|
||||
sensitive.append(os.environ[config_ref].encode())
|
||||
durable = b"".join(path.read_bytes() for path in run_root.rglob("*") if path.is_file())
|
||||
assert all(value and value not in durable for value in sensitive)
|
||||
print("ok: nine ready results are manifest-bound and runtime values are absent")
|
||||
PY
|
||||
```
|
||||
|
||||
Expected: every setup command exits 0; preflight prints one ready summary with nine results; durable inspection prints its safe success line. At planning time this block is expected to stop at agy/environment setup until the recorded resume condition is satisfied.
|
||||
|
||||
### [TEST-3] Preserve the full benchmark baseline
|
||||
|
||||
**Problem**
|
||||
|
||||
The final file consumes the new order, rubric, and all-cell contracts. It must not regress legacy fixtures or other benchmark behavior.
|
||||
|
||||
**Solution**
|
||||
|
||||
Validate all four manifests, run the focused static test and complete rubric/scoring/manifest/attempt/connectivity suites, and check the diff.
|
||||
|
||||
**Modified Files and Checklist**
|
||||
|
||||
- [ ] `CODE_REVIEW-cloud-G07.md`: record final command output after live readiness succeeds; if TEST-2 is blocked, leave this item unchecked and record the resume condition.
|
||||
|
||||
**Test Strategy**
|
||||
|
||||
No additional files beyond TEST-1. Existing suites are the regression oracle; external preflight output is not cached.
|
||||
|
||||
**Verification**
|
||||
|
||||
```bash
|
||||
for manifest in scripts/fixtures/agent-comparison-benchmark-manifest.example.json scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json; do python3 scripts/agent_comparison_benchmark.py validate --manifest "$manifest"; done
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.scoring_test scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Expected: four manifests validate, all suites report `OK`, and whitespace check is silent.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Item |
|
||||
|------|------|
|
||||
| `scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json` | TEST-1 |
|
||||
| `scripts/agent_benchmark/manifest_test.py` | TEST-1 |
|
||||
| `agent-task/m-iop-one-shot-agent-model-comparison/04+02,03_locked_benchmark_manifest/CODE_REVIEW-cloud-G07.md` | TEST-2, TEST-3 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
Run from `/config/workspace/iop-s0` in dependency order. Fresh output is required; cached output is not acceptable. Run the full TEST-2 command block exactly once after its setup checks pass, then run:
|
||||
|
||||
```bash
|
||||
python3 scripts/agent_comparison_benchmark.py validate --manifest scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json
|
||||
python3 -m unittest scripts.agent_benchmark.manifest_test.ManifestValidationTest.test_iop_one_shot_manifest_locks_benchmark_readiness
|
||||
for manifest in scripts/fixtures/agent-comparison-benchmark-manifest.example.json scripts/fixtures/agent-comparison-benchmark-direct-preflight.example.json scripts/fixtures/agent-comparison-benchmark-supported-direct.example.json scripts/fixtures/agent-comparison-benchmark-iop-one-shot.json; do python3 scripts/agent_comparison_benchmark.py validate --manifest "$manifest"; done
|
||||
python3 -m unittest scripts.agent_benchmark.rubric_test scripts.agent_benchmark.scoring_test scripts.agent_benchmark.manifest_test scripts.agent_benchmark.attempts_test scripts.agent_benchmark.connectivity_integration_test
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Expected: static validation/test pass, TEST-2 has one nine-ready durable preflight record with no runtime values, four manifests validate, all suites report `OK`, and `git diff --check` is silent. No scored benchmark command is run.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
Loading…
Reference in a new issue