feat(epic): pipeline-contract 작업을 준비한다
This commit is contained in:
parent
6a7ef5786d
commit
bdb39532ad
38 changed files with 5809 additions and 0 deletions
|
|
@ -0,0 +1,133 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/01_benchmark_manifest plan=2 tag=API milestone-task=benchmark-manifest -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-09
|
||||
task=m-agent-comparison-benchmark-pipeline/01_benchmark_manifest, plan=2, tag=API
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/plan_local_G05_1.log` and `agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/code_review_cloud_G05_1.log` (generation 1 retains generation 0 history).
|
||||
- Review state: the prior pair was unimplemented and had no official verdict; this explicit self-review archived it through plan `write` mode.
|
||||
- Self-review defects: the prior pair still left `repetitions` defaulting, timeout fields, and the caller-request versus expected IOP route/binding shape to implementation judgment. That ambiguity would let downstream workspace/lifecycle code derive incompatible canonical identities from the same intended benchmark cell.
|
||||
- Scope carried forward: `benchmark-manifest`, SDD S01, standard-library validation, the public `validate` command, and credential-free tests remain unchanged.
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files and verify that output in `Verification Results` matches code.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals.
|
||||
2. Archive `CODE_REVIEW-cloud-G05.md` → `code_review_cloud_G05_2.log` and `PLAN-local-G05.md` → `plan_local_G05_2.log`.
|
||||
3. If PASS, write `complete.log` and move active task directory to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/`. If WARN/FAIL, fully write the next filesystem state required by the code-review skill.
|
||||
4. If PASS, preserve `milestone-task=benchmark-manifest` in `complete.log` and report it for runtime aggregation. Roadmap state evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check applicable `Review-Only Checklist` items at the final `.log` location before reporting.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|---------|
|
||||
| API-1 Define and load the closed manifest | [ ] |
|
||||
| API-2 Expose deterministic validation | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement the exact closed manifest/schema with canonical default expansion, caller-request/expected-route separation, repository-root path rules, canonical fixture digest, immutable values, and deterministic matrix ordering.
|
||||
- [ ] Add the public validation CLI, deterministic example fixture, and normal/boundary/redaction/matrix-extension tests wired to the Makefile target.
|
||||
- [ ] Run focused and aggregate manifest verification plus patch-integrity checks.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Archive active `CODE_REVIEW-*-G??.md` to `code_review_cloud_G05_2.log`.
|
||||
- [ ] Archive active `PLAN-*-G??.md` to `plan_local_G05_2.log`.
|
||||
- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write `complete.log` based on `agent-ops/skills/common/code-review/templates/complete-log-template.md` and leave no active `.md` files.
|
||||
- [ ] If PASS, move active task directory to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/` and update this checklist at the final archive path.
|
||||
- [ ] If PASS, preserve and report `milestone-task=benchmark-manifest` for runtime aggregation, without modifying roadmap or directly calling `update-roadmap`.
|
||||
- [ ] If PASS for split work, remove empty active parent or verify it was kept due to remaining siblings/files.
|
||||
- [ ] If WARN/FAIL, write the next filesystem state matching code-review verdict and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record any deviations from the plan and the rationale here._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record key design decisions here._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Schema and loader enforce the same exact top-level/cell/timeout shape, canonical omitted-`repetitions=1` expansion, closed enums, bounds, and unknown-member policy.
|
||||
- Caller-visible `request_model`/`requested_effort` stay separate from direct/preset route and expected binding evidence; no model or effort is translated.
|
||||
- Asset mappings bind contained repository sources to normalized contained workspace destinations without collision.
|
||||
- Declared workspace checksum uses the versioned length-framed destination/content algorithm; canonical manifest digest binds expanded JSON plus prompt and asset source/destination/content; cell/binding ordering is stable.
|
||||
- `output_root` remains under `agent-test/runs`; exact `../iop-s2` remains runtime provenance and is never a fixture root.
|
||||
- Errors never echo prompt, asset, secret, or private-endpoint content; no downstream execution behavior entered this packet.
|
||||
|
||||
## Verification Results
|
||||
|
||||
### `python3 -m unittest scripts.agent_benchmark.manifest_test`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 scripts/agent_comparison_benchmark.py validate --manifest scripts/fixtures/agent-comparison-benchmark-manifest.example.json`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `make test-agent-comparison-benchmark`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `git diff --check`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these (archive, complete.log, and task-directory archive move are review-agent only) |
|
||||
| Archive Evidence Snapshot | Fixed at stub creation from plan when present | Implementing agent uses it as default prior-loop context; read only the specific archive files cited there when more detail is required |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results (section headings + commands) | Fixed at stub creation | Implementing agent fills in command output only; command changes require a `Deviations from Plan` entry |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,192 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/01_benchmark_manifest plan=2 tag=API milestone-task=benchmark-manifest -->
|
||||
|
||||
# Benchmark Manifest Contract
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections of `CODE_REVIEW-cloud-G05.md` is the mandatory final implementation step. Run every verification command, paste actual stdout/stderr, keep both active files in place, and report ready for review; finalization belongs only to the official code-review skill. If blocked, record exact evidence and the resume condition only in implementation-owned fields. Do not ask the user, create stop files, archive logs, or write `complete.log`.
|
||||
|
||||
## Background
|
||||
|
||||
The approved SDD is the only current benchmark-matrix definition. This packet creates the closed manifest, canonical digest/path rules, and deterministic validation entrypoint consumed by every later slice. It does not create workspaces, invoke callers, or create scored attempts.
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/plan_local_G05_1.log` and `agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/code_review_cloud_G05_1.log` (generation 1 retains generation 0 history).
|
||||
- Review state: the prior pair was unimplemented and had no official verdict; this explicit self-review archived it through plan `write` mode.
|
||||
- Self-review defects: the prior pair still left `repetitions` defaulting, timeout fields, and the caller-request versus expected IOP route/binding shape to implementation judgment. That ambiguity would let downstream workspace/lifecycle code derive incompatible canonical identities from the same intended benchmark cell.
|
||||
- Scope carried forward: `benchmark-manifest`, SDD S01, standard-library validation, the public `validate` command, and credential-free tests remain unchanged.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `AGENTS.md`
|
||||
- `agent-ops/rules/project/rules.md`
|
||||
- `agent-ops/rules/project/domain/testing/rules.md`
|
||||
- `agent-ops/rules/common/rules-roadmap.md`
|
||||
- `agent-ops/skills/common/plan/SKILL.md`
|
||||
- `agent-ops/skills/common/code-review/SKILL.md`
|
||||
- `agent-ops/skills/common/finalize-task-routing/SKILL.md`
|
||||
- `agent-ops/skills/common/plan/templates/review-stub-template.md`
|
||||
- `agent-test/local/rules.md`
|
||||
- `agent-test/local/testing-smoke.md`
|
||||
- `agent-roadmap/current.md`
|
||||
- `agent-roadmap/priority-queue.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/PHASE.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/agent-comparison-benchmark-pipeline.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md`
|
||||
- `agent-ops/rules/common/rules-agent-spec.md`
|
||||
- `agent-spec/index.md`
|
||||
- `agent-spec/input/openai-compatible-surface.md`
|
||||
- `agent-spec/runtime/provider-pool-config-refresh.md`
|
||||
- `agent-contract/index.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
- `scripts/fixtures/single-request-claude-smoke-manifest.schema.json`
|
||||
- `Makefile`
|
||||
- `go.mod`
|
||||
- `.gitignore`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- SDD status is `[승인됨]` with lock `해제`; first-line task id is `benchmark-manifest`.
|
||||
- S01 requires code-free matrix extension, closed validation, and canonical ordering.
|
||||
- Evidence Map S01 requires executable schema/fixture validation plus a matrix-extension test. API-1 owns the data contract and API-2 owns the public command/evidence.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- Handoff: none. Repository-native sources are the approved SDD, testing rules/profile, current schema convention, and Makefile.
|
||||
- Current checkout: `/config/workspace/iop-s0`, branch `feature/agent-comparison-benchmark-pipeline`; Python 3.12.3, Git 2.43.0, Linux arm64.
|
||||
- No provider, network, browser, or sibling-checkout write is permitted. Python standard library only; `go.mod` defines no Python package policy.
|
||||
- `bench-01` is the current `bench` lane head with no queue blocker. Milestone/SDD locks are released.
|
||||
- Confidence: high. All verification is deterministic and local.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
- No executable schema covers callers, the caller-visible IOP request, expected direct/preset binding evidence, prompt/assets, the omitted-`repetitions` default, lifecycle timeout components, session/cache policy, viewport/rubric metadata, or evidence root.
|
||||
- Tests must cover every closed enum/bound, duplicate ids/stages, canonical default expansion and ordering, digest drift, path escape/symlink escape, forbidden secret/endpoint members, and data-only matrix extension.
|
||||
|
||||
### Symbol References
|
||||
|
||||
None. The package and CLI are new.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
This predecessor-free packet has a stable independent contract: one immutable validated manifest and one deterministic `validate` command. Workspace/process/attempt behavior consumes it later and is excluded here.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
Excluded: workspace/session materialization, caller processes, retries/resume, caller adapters, metrics, browser validation, scoring, and Markdown reporting. The manifest may describe those later inputs but must not implement their behavior.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- evaluation_mode: `isolated-reassessment`
|
||||
- finalizer: `finalize-task-policy.sh`, mode `pair`
|
||||
- build closures all `true`; scores `2/0/1/1/1`; base/route `local-fit`; grade `G05`; catalog `worker/local/G05`; filename `PLAN-local-G05.md`.
|
||||
- review closures all `true`; scores `2/0/1/1/1`; route `official-review`; grade `G05`; catalog `review/cloud/G05`; filename `CODE_REVIEW-cloud-G05.md`.
|
||||
- large_indivisible_context: `false`; risks: `boundary_contract`, `structured_interpretation`, `variant_product`; rework `0`; evidence integrity failure `false`; capability gap none.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement the exact closed manifest/schema with canonical default expansion, caller-request/expected-route separation, repository-root path rules, canonical fixture digest, immutable values, and deterministic matrix ordering.
|
||||
- [ ] Add the public validation CLI, deterministic example fixture, and normal/boundary/redaction/matrix-extension tests wired to the Makefile target.
|
||||
- [ ] Run focused and aggregate manifest verification plus patch-integrity checks.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Define and load the closed manifest
|
||||
|
||||
**Problem:** `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md:78-83` defines the inputs but not executable path roots, ordering, or digest framing. Without one canonical rule, later workspace and resume code can accept the same JSON while deriving different identities.
|
||||
|
||||
**Solution:** Add a closed Draft 2020-12 schema plus a standard-library loader for the supported schema subset. The exact top-level fields are `pipeline_version="1"`, `environment="dev"`, `testbed="../iop-s2"`, `fixture`, `matrix`, optional `repetitions` (canonical default `1`), `session_policy="fresh"`, `setup_cache_policy="isolated"`, `timeout`, `viewports`, `rubric_version`, and `output_root`; every object rejects unknown members. `timeout` contains integer `run_seconds` (1..86400), `idle_seconds` (1..600), `quiet_seconds` (1..60), and `cleanup_grace_seconds` (1..60). `viewports` is a non-empty list of unique closed `{id,width,height}` records with bounded token ids and dimensions 1..8192; `rubric_version` is a bounded stable token.
|
||||
|
||||
Each matrix cell contains unique path-safe `id` matching `^[a-z0-9][a-z0-9_-]{0,63}$`, closed `caller=claude|agy|codex`, and an `iop` object with bounded `request_model`, bounded `requested_effort`, `route_kind=direct|execution_preset`, stable bounded `route_id`, and unique `expected_bindings[]` records (`stage`, `model`, optional `effort`). Direct cells require exactly one `stage=request` binding. Execution-preset cells require exactly one each of `selector`, `plan`, `work`, and `review`, and may include one `repair`; canonical order is the fixed rank `request,selector,plan,work,review,repair`, not lexical order. `request_model`/`requested_effort` are the only values later adapters send to IOP; route/binding fields are expected preflight evidence and never alternate endpoint/provider inputs. The loader sorts cells by id and bindings by fixed stage rank after applying defaults, but never translates a model or effort value.
|
||||
|
||||
Resolve `fixture.prompt` and each `fixture.assets[]` object (`source`, `workspace_path`) from the repository root, require regular contained source files, require normalized contained workspace destinations, and reject absolute paths, `..` escapes, symlink escapes, and destination collisions. Require stable `fixture.version`. Define declared `fixture.checksum` as `sha256:` over `b"IOP-BENCH-WORKSPACE\0"` followed by assets sorted by `workspace_path`, with each UTF-8 destination and file content framed by an unsigned 64-bit big-endian byte length. Separately expose `manifest.digest` over `b"IOP-BENCH-MANIFEST\0"`, canonical expanded JSON (`sort_keys=True`, compact separators, UTF-8), and length-framed prompt path/content plus every asset source/destination/content. This binds all scored inputs without making the digest circular. Restrict `output_root` to a repository-relative resolved path contained by `agent-test/runs`; the exact testbed is provenance only, never a fixture root. Return frozen dataclasses/tuples, never mutable raw dictionaries.
|
||||
|
||||
Before (`SDD.md:78-83`):
|
||||
|
||||
```text
|
||||
The manifest shape is documented, but executable path and digest identity are undefined.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
manifest = load_manifest(path, repo_root=repo_root)
|
||||
assert manifest.repetitions == 1 # when omitted
|
||||
assert manifest.fixture.checksum == digest_workspace_inputs(manifest.fixture.assets)
|
||||
assert manifest.digest == digest_manifest_and_resolved_inputs(manifest)
|
||||
assert tuple(cell.id for cell in manifest.matrix) == tuple(sorted(cell.id for cell in manifest.matrix))
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json` with the exact closed top-level/cell/timeout/viewports shape, constants/default declaration, fixture version/checksum, asset mappings, enums, and numeric bounds.
|
||||
- [ ] Add `scripts/agent_benchmark/__init__.py` exporting only stable immutable manifest APIs.
|
||||
- [ ] Add `scripts/agent_benchmark/manifest.py` with schema-subset validation, canonical digest/path checks, ordering, and sanitized errors.
|
||||
- [ ] Add `scripts/fixtures/agent-comparison-benchmark/prompt.md` and `scripts/fixtures/agent-comparison-benchmark/reference.txt` as inert deterministic fixture inputs.
|
||||
- [ ] Add `scripts/fixtures/agent-comparison-benchmark-manifest.example.json` with the computed checksum and no credential/private endpoint.
|
||||
|
||||
**Test Strategy:** Write `scripts/agent_benchmark/manifest_test.py`. Cover valid minimum/example, omitted repetitions canonicalizing identically to explicit `1`, data-only matrix extension, deterministic cell/binding sorting, duplicate ids/stages, direct/preset request-versus-evidence shapes, every enum/bound, unknown members, absolute/escaping/symlink source/destination/output paths, destination collisions, workspace checksum framing/drift, prompt/asset/manifest digest drift, non-positive repetitions/timeout components, and secret/private-endpoint fields without echoing their values.
|
||||
|
||||
**Verification:** `python3 -m unittest scripts.agent_benchmark.manifest_test` exits 0.
|
||||
|
||||
### [API-2] Expose deterministic validation
|
||||
|
||||
**Problem:** `Makefile:1-23` has no benchmark target or reusable benchmark CLI, so later slices could duplicate validation and error handling.
|
||||
|
||||
**Solution:** Add `scripts/agent_comparison_benchmark.py validate --manifest PATH` with stable exits (`0` valid, `64` usage, `69` invalid), one sanitized success/error line, and no manifest content echo. Add one Make target using fresh `unittest discover` for `*_test.py` and tracked-example validation.
|
||||
|
||||
Before (`Makefile:1`):
|
||||
|
||||
```make
|
||||
.PHONY: ...
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```text
|
||||
python3 scripts/agent_comparison_benchmark.py validate --manifest PATH
|
||||
make test-agent-comparison-benchmark
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_comparison_benchmark.py` with `validate` and sanitized exit mapping.
|
||||
- [ ] Add `scripts/agent_benchmark/manifest_test.py` with loader and real-CLI tests.
|
||||
- [ ] Update `Makefile` `.PHONY` and add `test-agent-comparison-benchmark` with fresh discovery and example validation.
|
||||
- [ ] Record exact output in `CODE_REVIEW-cloud-G05.md`.
|
||||
|
||||
**Test Strategy:** Invoke the real CLI for valid, missing, malformed, secret-bearing, and checksum-drift manifests. Assert raw sentinels never appear in stdout/stderr.
|
||||
|
||||
**Verification:** Focused tests, public example validation, and the aggregate target exit 0.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Items |
|
||||
|------|-------|
|
||||
| `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json` | API-1 |
|
||||
| `scripts/fixtures/agent-comparison-benchmark-manifest.example.json` | API-1 |
|
||||
| `scripts/fixtures/agent-comparison-benchmark/prompt.md` | API-1 |
|
||||
| `scripts/fixtures/agent-comparison-benchmark/reference.txt` | API-1 |
|
||||
| `scripts/agent_benchmark/__init__.py` | API-1 |
|
||||
| `scripts/agent_benchmark/manifest.py` | API-1 |
|
||||
| `scripts/agent_benchmark/manifest_test.py` | API-1, API-2 |
|
||||
| `scripts/agent_comparison_benchmark.py` | API-2 |
|
||||
| `Makefile` | API-2 |
|
||||
| `agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/CODE_REVIEW-cloud-G05.md` | API-2 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
1. `python3 -m unittest scripts.agent_benchmark.manifest_test`
|
||||
- Expected: all schema, ordering, path, digest, boundary, and redaction cases pass in a fresh process.
|
||||
2. `python3 scripts/agent_comparison_benchmark.py validate --manifest scripts/fixtures/agent-comparison-benchmark-manifest.example.json`
|
||||
- Expected: exit 0 with one sanitized success line and no prompt/asset content.
|
||||
3. `make test-agent-comparison-benchmark`
|
||||
- Expected: all credential-free benchmark tests and example validation pass without provider/network calls.
|
||||
4. `git diff --check`
|
||||
- Expected: exit 0 with no whitespace errors.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,126 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/01_benchmark_manifest plan=0 tag=API milestone-task=benchmark-manifest -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-09
|
||||
task=m-agent-comparison-benchmark-pipeline/01_benchmark_manifest, plan=0, tag=API
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files and verify that output in `Verification Results` matches code.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals.
|
||||
2. Archive `CODE_REVIEW-cloud-G05.md` → `code_review_cloud_G05_0.log` and `PLAN-local-G05.md` → `plan_local_G05_0.log`.
|
||||
3. If PASS, write `complete.log` and move active task directory to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/`. If WARN/FAIL, fully write the next filesystem state required by the code-review skill.
|
||||
4. If PASS, preserve `milestone-task=benchmark-manifest` in `complete.log` and report it for runtime aggregation; roadmap evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check applicable `Review-Only Checklist` items at the final `.log` location before reporting.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|---------|
|
||||
| API-1 Define and load the closed manifest | [ ] |
|
||||
| API-2 Expose deterministic validation | [ ] |
|
||||
| API-3 Verify the packet boundary | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement the closed manifest schema and immutable standard-library loader with canonical matrix/path/checksum validation.
|
||||
- [ ] Add the public validation CLI, deterministic example fixture, and normal/boundary/matrix-extension tests wired to the Makefile target.
|
||||
- [ ] Run the fresh benchmark manifest test target, validate the tracked example through the CLI, and run `git diff --check`.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Archive active `CODE_REVIEW-*-G??.md` to `code_review_cloud_G05_0.log`.
|
||||
- [ ] Archive active `PLAN-*-G??.md` to `plan_local_G05_0.log`.
|
||||
- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write `complete.log` based on `agent-ops/skills/common/code-review/templates/complete-log-template.md` and leave no active `.md` files.
|
||||
- [ ] If PASS, move active task directory `agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/` to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/` and update this checklist at the final archive path.
|
||||
- [ ] If PASS, preserve and report `milestone-task=benchmark-manifest` for runtime aggregation, without modifying roadmap or directly calling `update-roadmap`.
|
||||
- [ ] If PASS for split work, remove the empty active parent only when no sibling remains.
|
||||
- [ ] If WARN/FAIL, write the next filesystem state matching code-review verdict and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record any deviations from the plan and the rationale here._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record key design decisions here._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- The schema and executable loader reject the same closed set and bounds.
|
||||
- Matrix extension changes only manifest data; canonical ordering and duplicate rejection are deterministic.
|
||||
- Paths are relative/contained, checksums bind fixture content, and validation output never echoes prompt or secret material.
|
||||
- No workspace, process, provider, network, retry, or report behavior entered this packet.
|
||||
|
||||
## Verification Results
|
||||
|
||||
Paste actual stdout/stderr below each command. Do not summarize or reconstruct output. If a command changes, record the replacement and reason in `Deviations from Plan` first.
|
||||
|
||||
### `python3 -m unittest scripts.agent_benchmark.manifest_test`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 scripts/agent_comparison_benchmark.py validate --manifest scripts/fixtures/agent-comparison-benchmark-manifest.example.json`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `make test-agent-comparison-benchmark`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `git diff --check`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results | Fixed headings/commands | Implementing agent fills actual output only |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,131 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/01_benchmark_manifest plan=1 tag=API milestone-task=benchmark-manifest -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-09
|
||||
task=m-agent-comparison-benchmark-pipeline/01_benchmark_manifest, plan=1, tag=API
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/plan_local_G05_0.log` and `agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/code_review_cloud_G05_0.log`.
|
||||
- Review state: unimplemented, no official verdict, replaced through explicit plan `write` mode.
|
||||
- Corrected scope: checkout identity plus closed source/destination path roots and collision-safe fixture digest framing.
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files and verify that output in `Verification Results` matches code.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals.
|
||||
2. Archive `CODE_REVIEW-cloud-G05.md` → `code_review_cloud_G05_1.log` and `PLAN-local-G05.md` → `plan_local_G05_1.log`.
|
||||
3. If PASS, write `complete.log` and move active task directory to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/`. If WARN/FAIL, fully write the next filesystem state required by the code-review skill.
|
||||
4. If PASS, preserve `milestone-task=benchmark-manifest` in `complete.log` and report it for runtime aggregation. Roadmap state evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check applicable `Review-Only Checklist` items at the final `.log` location before reporting.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|---------|
|
||||
| API-1 Define and load the closed manifest | [ ] |
|
||||
| API-2 Expose deterministic validation | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement the closed manifest/schema with repository-root path rules, canonical fixture digest, immutable values, and deterministic matrix ordering.
|
||||
- [ ] Add the public validation CLI, deterministic example fixture, and normal/boundary/redaction/matrix-extension tests wired to the Makefile target.
|
||||
- [ ] Run focused and aggregate manifest verification plus patch-integrity checks.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Archive active `CODE_REVIEW-*-G??.md` to `code_review_cloud_G05_1.log`.
|
||||
- [ ] Archive active `PLAN-*-G??.md` to `plan_local_G05_1.log`.
|
||||
- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write `complete.log` based on `agent-ops/skills/common/code-review/templates/complete-log-template.md` and leave no active `.md` files.
|
||||
- [ ] If PASS, move active task directory to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/` and update this checklist at the final archive path.
|
||||
- [ ] If PASS, preserve and report `milestone-task=benchmark-manifest` for runtime aggregation, without modifying roadmap or directly calling `update-roadmap`.
|
||||
- [ ] If PASS for split work, remove empty active parent or verify it was kept due to remaining siblings/files.
|
||||
- [ ] If WARN/FAIL, write the next filesystem state matching code-review verdict and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record any deviations from the plan and the rationale here._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record key design decisions here._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Schema and loader enforce the same closed enums, bounds, and unknown-member policy.
|
||||
- Asset mappings bind contained repository sources to normalized contained workspace destinations without collision.
|
||||
- Declared workspace checksum frames sorted destination/content records; canonical manifest digest also binds prompt and asset source/destination/content; matrix ordering is stable.
|
||||
- `output_root` remains under `agent-test/runs`; exact `../iop-s2` remains runtime provenance and is never a fixture root.
|
||||
- Errors never echo prompt, asset, secret, or private-endpoint content; no downstream execution behavior entered this packet.
|
||||
|
||||
## Verification Results
|
||||
|
||||
### `python3 -m unittest scripts.agent_benchmark.manifest_test`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 scripts/agent_comparison_benchmark.py validate --manifest scripts/fixtures/agent-comparison-benchmark-manifest.example.json`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `make test-agent-comparison-benchmark`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `git diff --check`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these (archive, complete.log, and task-directory archive move are review-agent only) |
|
||||
| Archive Evidence Snapshot | Fixed at stub creation from plan when present | Implementing agent uses it as default prior-loop context; read only the specific archive files cited there when more detail is required |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results (section headings + commands) | Fixed at stub creation | Implementing agent fills in command output only; command changes require a `Deviations from Plan` entry |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,195 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/01_benchmark_manifest plan=0 tag=API milestone-task=benchmark-manifest -->
|
||||
|
||||
# Benchmark Manifest Contract
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections of `CODE_REVIEW-cloud-G05.md` is the mandatory final implementation step. Run every verification command, paste actual stdout/stderr, keep both active files in place, and report ready for review; finalization belongs only to the official code-review skill. If blocked, record the exact blocker, attempted commands/output, and resume condition only in the implementation-owned evidence fields. Do not ask the user, call user-input tools, create control-plane stop files, classify the next state, archive logs, or write `complete.log`.
|
||||
|
||||
## Background
|
||||
|
||||
The benchmark matrix is currently encoded only in the approved SDD, so adding callers, routes, fixtures, or repetitions would otherwise require editing execution code. This packet establishes the closed manifest contract and deterministic validation entrypoint that all later pipeline slices consume. It does not execute an external caller or create scored attempts.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `AGENTS.md`
|
||||
- `agent-ops/rules/project/rules.md`
|
||||
- `agent-ops/rules/project/domain/testing/rules.md`
|
||||
- `agent-test/local/rules.md`
|
||||
- `agent-test/local/testing-smoke.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/agent-comparison-benchmark-pipeline.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md`
|
||||
- `agent-spec/input/openai-compatible-surface.md`
|
||||
- `agent-spec/runtime/provider-pool-config-refresh.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
- `scripts/e2e-single-request-claude.sh`
|
||||
- `scripts/fixtures/single-request-claude-smoke-manifest.schema.json`
|
||||
- `Makefile`
|
||||
- `go.mod`
|
||||
- `.gitignore`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- SDD: `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md`; status `[승인됨]`, lock `해제`.
|
||||
- First-line milestone task: `benchmark-manifest`.
|
||||
- Acceptance Scenario: S01. A new caller/model/prompt/repetition matrix must validate and canonicalize without code changes.
|
||||
- Evidence Map: schema/fixture validation plus a matrix-extension test must prove config-driven behavior. Those requirements become API-1 and API-2, and the final verification runs the credential-free test target and example-manifest validation.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- Handoff: no separate verification handoff was supplied.
|
||||
- Sources: `agent-test/local/rules.md`, `agent-test/local/testing-smoke.md`, `Makefile`, and the existing single-request harness/schema.
|
||||
- Environment: local checkout `/config/workspace/iop-s0-bench-01`; Python 3.12.3 and Git 2.43.0 are available on Linux arm64.
|
||||
- Commands/criteria: fresh Python unit tests, example validation through the real CLI entrypoint, and `git diff --check`; no provider or network call is permitted.
|
||||
- Preconditions: predecessor-free packet; Python standard library only. `go.mod` contains no Python dependency policy, so this packet must not introduce an undeclared `jsonschema` dependency.
|
||||
- Constraints: closed objects, relative contained fixture/evidence paths, positive repetitions, stable unique cell ids, deterministic ordering, and no raw secret/private endpoint fields.
|
||||
- Gap: the existing `./scripts/e2e-single-request-claude.sh --self-test` baseline currently fails at `run-early-signal: supervisor cleanup was not bounded`; it is not a pass criterion for this manifest-only packet and must not be reported as caused or fixed here.
|
||||
- Repository-native fallback evidence: `scripts/e2e-single-request-claude.sh:261-417` validates a closed schema and manifest using the Python standard library; its schema is closed at every object.
|
||||
- Confidence: high for manifest validation; no external environment is involved.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
- No current test covers a general caller/route/preset/fixture/repetition manifest.
|
||||
- New tests must cover a valid minimum manifest, unknown fields, malformed/absolute/escaping paths, invalid caller/route/effort values, duplicate ids, non-positive repetitions, checksum mismatch, secret-shaped keys, and matrix extension/canonical ordering.
|
||||
|
||||
### Symbol References
|
||||
|
||||
None. All Python modules and CLI commands are new.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
This is the dependency root. Its stable contract is an immutable validated manifest value plus a deterministic `validate` command; it can PASS without workspace creation, child processes, or attempt persistence. Later packets consume this API and are encoded as dependent sibling directories.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
Excluded: workspace copying, caller process execution, retries/resume, caller-specific adapters, timing/usage normalization, web validation, scoring, and final Markdown comparison reports. Those belong to later pipeline packets or other Milestone Epics and must not be introduced here.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- evaluation_mode: `first-pass`
|
||||
- finalizer: `finalize-task-policy.sh`, mode `pair`
|
||||
- build closures: scope/context/verification/evidence/ownership/decision all `true`; scores `2/0/1/1/1`; base `local-fit`; route `local-fit`; grade `G05`; catalog `worker/local/G05`; filename `PLAN-local-G05.md`.
|
||||
- review closures: all `true`; scores `2/0/1/1/1`; route `official-review`; grade `G05`; catalog `review/cloud/G05`; filename `CODE_REVIEW-cloud-G05.md`.
|
||||
- large_indivisible_context: `false`.
|
||||
- matched loop risks: `boundary_contract`, `structured_interpretation`, `variant_product` (3); risk boundary `false`.
|
||||
- recovery signals: `review_rework_count=0`, `evidence_integrity_failure=false`.
|
||||
- capability gap: none.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement the closed manifest schema and immutable standard-library loader with canonical matrix/path/checksum validation.
|
||||
- [ ] Add the public validation CLI, deterministic example fixture, and normal/boundary/matrix-extension tests wired to the Makefile target.
|
||||
- [ ] Run the fresh benchmark manifest test target, validate the tracked example through the CLI, and run `git diff --check`.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Define and load the closed manifest
|
||||
|
||||
**Problem:** `SDD.md:78-94` defines the input and secrecy contract, but no project-owned manifest/schema or immutable loader exists. The existing S12 schema at `scripts/fixtures/single-request-claude-smoke-manifest.schema.json:1` is fixed to one Claude qualification and cannot express the benchmark matrix.
|
||||
|
||||
**Solution:** Add a Draft 2020-12 closed JSON schema and a Python standard-library loader. Require exactly `pipeline_version`, `environment=dev`, `testbed=../iop-s2`, `fixture`, `matrix`, `repetitions`, `session_policy=fresh`, `setup_cache_policy`, `timeout`, `viewports`, `rubric_version`, and `output_root`. Validate relative contained paths, fixture checksum/version, unique stable cell ids, caller enum `claude|agy|codex`, explicit direct/preset route identity, model/effort, positive repetitions, bounded positive timeouts, deterministic ordering, and secret-safe keys/values. Return frozen values rather than mutable raw dictionaries.
|
||||
|
||||
Before (`agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md:78-83`):
|
||||
|
||||
```text
|
||||
manifest input is documented, but no executable schema or loader exists.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
manifest = load_manifest(path, repo_root=repo_root)
|
||||
for cell in manifest.matrix: # stable id order
|
||||
assert cell.caller in {"claude", "agy", "codex"}
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json` with closed objects and explicit enums/bounds.
|
||||
- [ ] Add `scripts/agent_benchmark/__init__.py` exporting only the stable manifest types/loader.
|
||||
- [ ] Add `scripts/agent_benchmark/manifest.py` with schema-contract checks, path/checksum validation, canonical ordering, and frozen types.
|
||||
- [ ] Add `scripts/fixtures/agent-comparison-benchmark/prompt.md` and one inert asset used only by the deterministic example.
|
||||
- [ ] Add `scripts/fixtures/agent-comparison-benchmark-manifest.example.json` with the computed fixture checksum and no secret/private endpoint.
|
||||
|
||||
**Test Strategy:** Write `scripts/agent_benchmark/manifest_test.py`. Test the minimum valid example, adding another matrix cell without code changes, canonical ordering, duplicate ids, every enum/boundary, unknown members, absolute/escaping/symlink fixture paths, checksum drift, non-positive repetitions/timeouts, and forbidden secret-shaped inputs.
|
||||
|
||||
**Verification:** `python3 -m unittest scripts.agent_benchmark.manifest_test` exits 0 with all cases passing.
|
||||
|
||||
### [API-2] Expose deterministic validation
|
||||
|
||||
**Problem:** `scripts/e2e-single-request-claude.sh:2014-2035` has a one-off `--validate-manifest` command, while the benchmark needs one reusable command that later run/resume flows can call without duplicating validation logic.
|
||||
|
||||
**Solution:** Add `scripts/agent_comparison_benchmark.py validate --manifest <path>` with stable exit classes (`0` valid, `64` usage, `69` validation), sanitized diagnostics, and no manifest content echo. Add a Make target that runs the package tests and validates the tracked example; later modules are automatically included by test discovery.
|
||||
|
||||
Before (`scripts/e2e-single-request-claude.sh:2014-2035`):
|
||||
|
||||
```bash
|
||||
validate-manifest)
|
||||
# S12-only validator
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```text
|
||||
python3 scripts/agent_comparison_benchmark.py validate --manifest PATH
|
||||
make test-agent-comparison-benchmark
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_comparison_benchmark.py` with the `validate` subcommand and sanitized error mapping.
|
||||
- [ ] Update `Makefile` `.PHONY` and add `test-agent-comparison-benchmark` using fresh unittest discovery plus example validation.
|
||||
- [ ] Extend `scripts/agent_benchmark/manifest_test.py` with CLI exit/output and no-secret-echo assertions.
|
||||
|
||||
**Test Strategy:** Write subprocess tests against the real CLI for valid, malformed, secret-bearing, and missing manifests. Assert the raw offending value never appears in stdout/stderr.
|
||||
|
||||
**Verification:** `make test-agent-comparison-benchmark` exits 0 and the output contains no fixture prompt or secret sentinel.
|
||||
|
||||
### [API-3] Verify the packet boundary
|
||||
|
||||
**Problem:** A schema can appear correct while the executable loader, example checksum, and CLI disagree.
|
||||
|
||||
**Solution:** Run the exact public validation path after fresh tests, then check whitespace/patch integrity.
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Record exact test and validation output in `CODE_REVIEW-cloud-G05.md`.
|
||||
- [ ] Confirm no file outside the Modified Files Summary changed.
|
||||
|
||||
**Test Strategy:** No additional test file; this item executes the public entrypoint and repository patch check.
|
||||
|
||||
**Verification:** All Final Verification commands exit 0.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Items |
|
||||
|------|-------|
|
||||
| `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json` | API-1 |
|
||||
| `scripts/fixtures/agent-comparison-benchmark-manifest.example.json` | API-1 |
|
||||
| `scripts/fixtures/agent-comparison-benchmark/prompt.md` | API-1 |
|
||||
| `scripts/fixtures/agent-comparison-benchmark/reference.txt` | API-1 |
|
||||
| `scripts/agent_benchmark/__init__.py` | API-1 |
|
||||
| `scripts/agent_benchmark/manifest.py` | API-1 |
|
||||
| `scripts/agent_benchmark/manifest_test.py` | API-1, API-2 |
|
||||
| `scripts/agent_comparison_benchmark.py` | API-2 |
|
||||
| `Makefile` | API-2 |
|
||||
| `agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/CODE_REVIEW-cloud-G05.md` | API-3 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
1. `python3 -m unittest scripts.agent_benchmark.manifest_test`
|
||||
- Expected: fresh test process exits 0; valid, boundary, canonical-order, checksum, containment, and redaction cases pass.
|
||||
- Cache: not applicable; Python unittest runs fresh.
|
||||
2. `python3 scripts/agent_comparison_benchmark.py validate --manifest scripts/fixtures/agent-comparison-benchmark-manifest.example.json`
|
||||
- Expected: exit 0 with one sanitized validation-success line and no raw prompt content.
|
||||
- Cache: not applicable.
|
||||
3. `make test-agent-comparison-benchmark`
|
||||
- Expected: all discovered benchmark tests and example validation pass without any provider/network call.
|
||||
- Cache: not applicable.
|
||||
4. `git diff --check`
|
||||
- Expected: exit 0 with no whitespace errors.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,185 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/01_benchmark_manifest plan=1 tag=API milestone-task=benchmark-manifest -->
|
||||
|
||||
# Benchmark Manifest Contract
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections of `CODE_REVIEW-cloud-G05.md` is the mandatory final implementation step. Run every verification command, paste actual stdout/stderr, keep both active files in place, and report ready for review; finalization belongs only to the official code-review skill. If blocked, record exact evidence and the resume condition only in implementation-owned fields. Do not ask the user, create stop files, archive logs, or write `complete.log`.
|
||||
|
||||
## Background
|
||||
|
||||
The approved SDD is the only current benchmark-matrix definition. This packet creates the closed manifest, canonical digest/path rules, and deterministic validation entrypoint consumed by every later slice. It does not create workspaces, invoke callers, or create scored attempts.
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/plan_local_G05_0.log` and `agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/code_review_cloud_G05_0.log`.
|
||||
- Review state: the prior pair was unimplemented and had no official verdict; this explicit self-review archived it through plan `write` mode.
|
||||
- Self-review defects: the prior verification context named a different checkout, and fixture/evidence path roots plus the fixture checksum algorithm were not closed enough for downstream workspace identity.
|
||||
- Scope carried forward: `benchmark-manifest`, SDD S01, standard-library validation, the public `validate` command, and credential-free tests remain unchanged.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `AGENTS.md`
|
||||
- `agent-ops/rules/project/rules.md`
|
||||
- `agent-ops/rules/project/domain/testing/rules.md`
|
||||
- `agent-ops/rules/common/rules-roadmap.md`
|
||||
- `agent-ops/skills/common/plan/SKILL.md`
|
||||
- `agent-ops/skills/common/finalize-task-routing/SKILL.md`
|
||||
- `agent-test/local/rules.md`
|
||||
- `agent-test/local/testing-smoke.md`
|
||||
- `agent-roadmap/current.md`
|
||||
- `agent-roadmap/priority-queue.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/PHASE.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/agent-comparison-benchmark-pipeline.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md`
|
||||
- `agent-ops/rules/common/rules-agent-spec.md`
|
||||
- `agent-spec/index.md`
|
||||
- `agent-spec/input/openai-compatible-surface.md`
|
||||
- `agent-spec/runtime/provider-pool-config-refresh.md`
|
||||
- `agent-contract/index.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
- `scripts/fixtures/single-request-claude-smoke-manifest.schema.json`
|
||||
- `Makefile`
|
||||
- `go.mod`
|
||||
- `.gitignore`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- SDD status is `[승인됨]` with lock `해제`; first-line task id is `benchmark-manifest`.
|
||||
- S01 requires code-free matrix extension, closed validation, and canonical ordering.
|
||||
- Evidence Map S01 requires executable schema/fixture validation plus a matrix-extension test. API-1 owns the data contract and API-2 owns the public command/evidence.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- Handoff: none. Repository-native sources are the approved SDD, testing rules/profile, current schema convention, and Makefile.
|
||||
- Current checkout: `/config/workspace/iop-s0`, branch `feature/agent-comparison-benchmark-pipeline`; Python 3.12.3, Git 2.43.0, Linux arm64.
|
||||
- No provider, network, browser, or sibling-checkout write is permitted. Python standard library only; `go.mod` defines no Python package policy.
|
||||
- `bench-01` is the current `bench` lane head with no queue blocker. Milestone/SDD locks are released.
|
||||
- Confidence: high. All verification is deterministic and local.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
- No executable schema covers callers, route kind, model/effort, prompt/assets, repetitions, timeout, session/cache policy, or evidence root.
|
||||
- Tests must cover every closed enum/bound, duplicate ids, canonical ordering, digest drift, path escape/symlink escape, forbidden secret/endpoint members, and data-only matrix extension.
|
||||
|
||||
### Symbol References
|
||||
|
||||
None. The package and CLI are new.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
This predecessor-free packet has a stable independent contract: one immutable validated manifest and one deterministic `validate` command. Workspace/process/attempt behavior consumes it later and is excluded here.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
Excluded: workspace/session materialization, caller processes, retries/resume, caller adapters, metrics, browser validation, scoring, and Markdown reporting. The manifest may describe those later inputs but must not implement their behavior.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- evaluation_mode: `isolated-reassessment`
|
||||
- finalizer: `finalize-task-policy.sh`, mode `pair`
|
||||
- build closures all `true`; scores `2/0/1/1/1`; base/route `local-fit`; grade `G05`; catalog `worker/local/G05`; filename `PLAN-local-G05.md`.
|
||||
- review closures all `true`; scores `2/0/1/1/1`; route `official-review`; grade `G05`; catalog `review/cloud/G05`; filename `CODE_REVIEW-cloud-G05.md`.
|
||||
- large_indivisible_context: `false`; risks: `boundary_contract`, `structured_interpretation`, `variant_product`; rework `0`; evidence integrity failure `false`; capability gap none.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement the closed manifest/schema with repository-root path rules, canonical fixture digest, immutable values, and deterministic matrix ordering.
|
||||
- [ ] Add the public validation CLI, deterministic example fixture, and normal/boundary/redaction/matrix-extension tests wired to the Makefile target.
|
||||
- [ ] Run focused and aggregate manifest verification plus patch-integrity checks.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Define and load the closed manifest
|
||||
|
||||
**Problem:** `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md:78-83` defines the inputs but not executable path roots, ordering, or digest framing. Without one canonical rule, later workspace and resume code can accept the same JSON while deriving different identities.
|
||||
|
||||
**Solution:** Add a closed Draft 2020-12 schema plus a standard-library loader for the supported schema subset. Resolve `fixture.prompt` and each `fixture.assets[]` object (`source`, `workspace_path`) from the repository root, require regular contained source files, require normalized contained workspace destinations, and reject absolute paths, `..` escapes, symlink escapes, and destination collisions. Require a stable `fixture.version`. Define declared `fixture.checksum` as the initial workspace checksum: `sha256:` over assets sorted by `workspace_path`, framing each record with length-delimited UTF-8 workspace path and file bytes. Separately expose a canonical `manifest.digest` that frames canonical JSON plus the prompt source/content and every asset source/destination/content; this binds all scored inputs while keeping the workspace checksum reproducible from materialized workspace bytes. Require the exact testbed value `../iop-s2`; it is provenance/runtime input, not a fixture root. Restrict `output_root` to a repository-relative path contained by `agent-test/runs`. Canonicalize matrix cells by unique stable `id`. Require `session_policy="fresh"`; use a closed initial `setup_cache_policy` vocabulary and record it without hidden defaults. Return frozen dataclasses/tuples, never mutable raw dictionaries.
|
||||
|
||||
Before (`SDD.md:78-83`):
|
||||
|
||||
```text
|
||||
The manifest shape is documented, but executable path and digest identity are undefined.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
manifest = load_manifest(path, repo_root=repo_root)
|
||||
assert manifest.fixture.checksum == digest_workspace_inputs(manifest.fixture.assets)
|
||||
assert manifest.digest == digest_manifest_and_resolved_inputs(manifest)
|
||||
assert tuple(cell.id for cell in manifest.matrix) == tuple(sorted(cell.id for cell in manifest.matrix))
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json` with closed objects, exact constants, fixture version/checksum, asset source/destination mappings, enums, and numeric bounds.
|
||||
- [ ] Add `scripts/agent_benchmark/__init__.py` exporting only stable immutable manifest APIs.
|
||||
- [ ] Add `scripts/agent_benchmark/manifest.py` with schema-subset validation, canonical digest/path checks, ordering, and sanitized errors.
|
||||
- [ ] Add `scripts/fixtures/agent-comparison-benchmark/prompt.md` and `scripts/fixtures/agent-comparison-benchmark/reference.txt` as inert deterministic fixture inputs.
|
||||
- [ ] Add `scripts/fixtures/agent-comparison-benchmark-manifest.example.json` with the computed checksum and no credential/private endpoint.
|
||||
|
||||
**Test Strategy:** Write `scripts/agent_benchmark/manifest_test.py`. Cover valid minimum/example, data-only matrix extension, deterministic cell sorting, duplicate ids, every enum/bound, unknown members, absolute/escaping/symlink source and destination paths, destination collisions, output-root policy, workspace checksum drift, prompt/asset/manifest digest drift, non-positive repetitions/timeouts, and secret/private-endpoint fields without echoing their values.
|
||||
|
||||
**Verification:** `python3 -m unittest scripts.agent_benchmark.manifest_test` exits 0.
|
||||
|
||||
### [API-2] Expose deterministic validation
|
||||
|
||||
**Problem:** `Makefile:1-23` has no benchmark target or reusable benchmark CLI, so later slices could duplicate validation and error handling.
|
||||
|
||||
**Solution:** Add `scripts/agent_comparison_benchmark.py validate --manifest PATH` with stable exits (`0` valid, `64` usage, `69` invalid), one sanitized success/error line, and no manifest content echo. Add one Make target using fresh `unittest discover` for `*_test.py` and tracked-example validation.
|
||||
|
||||
Before (`Makefile:1`):
|
||||
|
||||
```make
|
||||
.PHONY: ...
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```text
|
||||
python3 scripts/agent_comparison_benchmark.py validate --manifest PATH
|
||||
make test-agent-comparison-benchmark
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_comparison_benchmark.py` with `validate` and sanitized exit mapping.
|
||||
- [ ] Add `scripts/agent_benchmark/manifest_test.py` with loader and real-CLI tests.
|
||||
- [ ] Update `Makefile` `.PHONY` and add `test-agent-comparison-benchmark` with fresh discovery and example validation.
|
||||
- [ ] Record exact output in `CODE_REVIEW-cloud-G05.md`.
|
||||
|
||||
**Test Strategy:** Invoke the real CLI for valid, missing, malformed, secret-bearing, and checksum-drift manifests. Assert raw sentinels never appear in stdout/stderr.
|
||||
|
||||
**Verification:** Focused tests, public example validation, and the aggregate target exit 0.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Items |
|
||||
|------|-------|
|
||||
| `scripts/fixtures/agent-comparison-benchmark-manifest.schema.json` | API-1 |
|
||||
| `scripts/fixtures/agent-comparison-benchmark-manifest.example.json` | API-1 |
|
||||
| `scripts/fixtures/agent-comparison-benchmark/prompt.md` | API-1 |
|
||||
| `scripts/fixtures/agent-comparison-benchmark/reference.txt` | API-1 |
|
||||
| `scripts/agent_benchmark/__init__.py` | API-1 |
|
||||
| `scripts/agent_benchmark/manifest.py` | API-1 |
|
||||
| `scripts/agent_benchmark/manifest_test.py` | API-1, API-2 |
|
||||
| `scripts/agent_comparison_benchmark.py` | API-2 |
|
||||
| `Makefile` | API-2 |
|
||||
| `agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/CODE_REVIEW-cloud-G05.md` | API-2 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
1. `python3 -m unittest scripts.agent_benchmark.manifest_test`
|
||||
- Expected: all schema, ordering, path, digest, boundary, and redaction cases pass in a fresh process.
|
||||
2. `python3 scripts/agent_comparison_benchmark.py validate --manifest scripts/fixtures/agent-comparison-benchmark-manifest.example.json`
|
||||
- Expected: exit 0 with one sanitized success line and no prompt/asset content.
|
||||
3. `make test-agent-comparison-benchmark`
|
||||
- Expected: all credential-free benchmark tests and example validation pass without provider/network calls.
|
||||
4. `git diff --check`
|
||||
- Expected: exit 0 with no whitespace errors.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,133 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace plan=3 tag=API milestone-task=isolated-workspace -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-09
|
||||
task=m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace, plan=3, tag=API
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/plan_cloud_G05_2.log` and `agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/code_review_cloud_G05_2.log` (generation 2 retains earlier history).
|
||||
- Review state: unimplemented, no official verdict, replaced through explicit plan `write` mode.
|
||||
- Self-review defects: the prior pair made `prepare_workspace` exclusively create the attempt root, while dependent `repeat-attempt` also owns exclusive attempt allocation. The two valid plans therefore could not compose, and opaque identity strings still lacked one shared safe path grammar.
|
||||
- Scope carried forward: contained exclusive allocation, integrity evidence, source non-mutation, predecessor resolution, and credential-free tests.
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files and verify that output in `Verification Results` matches code.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals.
|
||||
2. Archive `CODE_REVIEW-cloud-G05.md` → `code_review_cloud_G05_3.log` and `PLAN-cloud-G05.md` → `plan_cloud_G05_3.log`.
|
||||
3. If PASS, write `complete.log` and move the active task directory to its monthly group archive. If WARN/FAIL, write the next state required by the code-review skill.
|
||||
4. If PASS, preserve `milestone-task=isolated-workspace` and report it for runtime aggregation; roadmap evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check applicable `Review-Only Checklist` items at the final `.log` location.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|---------|
|
||||
| API-1 Materialize one clean workspace and session | [ ] |
|
||||
| API-2 Prove cross-attempt isolation and source integrity | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement fixture-seeded workspace and empty caller-session materialization beneath an already allocated empty attempt root, with shared identity/path validation and no copy/write of the `../iop-s2` runtime testbed.
|
||||
- [ ] Add deterministic containment, checksum, collision, session-freshness, matrix-isolation, and testbed-nonmutation tests.
|
||||
- [ ] Resolve predecessor index `01`, then run focused, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Archive active `CODE_REVIEW-*-G??.md` to `code_review_cloud_G05_3.log`.
|
||||
- [ ] Archive active `PLAN-*-G??.md` to `plan_cloud_G05_3.log`.
|
||||
- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write `complete.log` from the canonical template and leave no active `.md` files.
|
||||
- [ ] If PASS, move the active task directory to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/` and update this checklist there.
|
||||
- [ ] If PASS, preserve/report `milestone-task=isolated-workspace` without modifying roadmap or calling `update-roadmap` directly.
|
||||
- [ ] If PASS for split work, remove empty active parent or verify it remains for sibling files.
|
||||
- [ ] If WARN/FAIL, write the next filesystem state matching the verdict and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record any deviations from the plan and the rationale here._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record key design decisions here._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Exactly one predecessor completion is resolved before implementation and final manifest APIs are reused.
|
||||
- The attempt store owns only exclusive attempt-root allocation; workspace preparation validates the shared identity grammar and exclusively creates only `workspace/`, `session/`, and `prepared.json` below an empty root.
|
||||
- Only declared source/destination assets seed workspaces; prompt/testbed content is not copied implicitly.
|
||||
- Every cell/repetition/attempt gets a distinct empty session locator with no host history/resume reuse.
|
||||
- Testbed Git provenance is checked before/after without traversal or mutation; unsupported dirty/cache states fail closed.
|
||||
- Containment, symlink, collision, checksum, isolation, and atomic prepared-evidence boundaries are exercised.
|
||||
|
||||
## Verification Results
|
||||
|
||||
### Predecessor completion check from `PLAN-cloud-G05.md`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 -m unittest scripts.agent_benchmark.workspace_test`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `make test-agent-comparison-benchmark`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `git diff --check`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these |
|
||||
| Archive Evidence Snapshot | Fixed at stub creation from plan when present | Read only cited archive evidence when more detail is required |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results (section headings + commands) | Fixed at stub creation | Fill output only; changed commands require a deviation entry |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,181 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace plan=3 tag=API milestone-task=isolated-workspace -->
|
||||
|
||||
# Isolated Benchmark Workspace
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections of `CODE_REVIEW-cloud-G05.md` is mandatory. Resolve the encoded predecessor before editing, run every verification command, paste actual output, and leave both active files in place for official review. Do not finalize, archive, write `complete.log`, ask the user, or change predecessor ownership.
|
||||
|
||||
## Background
|
||||
|
||||
Every cell/repetition needs a clean workspace and fresh caller-session locator derived from the same manifest fixture. `../iop-s2` is the dev IOP runtime testbed whose provenance and non-mutation are recorded; it is not copied into the scored workspace. Process execution and attempt retry remain downstream.
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/plan_cloud_G05_2.log` and `agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/code_review_cloud_G05_2.log` (generation 2 retains earlier history).
|
||||
- Review state: unimplemented, no official verdict, replaced through explicit plan `write` mode.
|
||||
- Self-review defects: the prior pair made `prepare_workspace` exclusively create the attempt root, while dependent `repeat-attempt` also owns exclusive attempt allocation. The two valid plans therefore could not compose, and opaque identity strings still lacked one shared safe path grammar.
|
||||
- Scope carried forward: contained exclusive allocation, integrity evidence, source non-mutation, predecessor resolution, and credential-free tests.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `AGENTS.md`
|
||||
- `agent-ops/rules/project/rules.md`
|
||||
- `agent-ops/rules/project/domain/testing/rules.md`
|
||||
- `agent-ops/rules/common/rules-roadmap.md`
|
||||
- `agent-ops/skills/common/plan/SKILL.md`
|
||||
- `agent-ops/skills/common/code-review/SKILL.md`
|
||||
- `agent-ops/skills/common/finalize-task-routing/SKILL.md`
|
||||
- `agent-ops/skills/common/plan/templates/review-stub-template.md`
|
||||
- `agent-test/local/rules.md`
|
||||
- `agent-test/local/testing-smoke.md`
|
||||
- `agent-roadmap/current.md`
|
||||
- `agent-roadmap/priority-queue.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/PHASE.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/agent-comparison-benchmark-pipeline.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md`
|
||||
- `agent-spec/index.md`
|
||||
- `agent-spec/input/openai-compatible-surface.md`
|
||||
- `agent-spec/runtime/provider-pool-config-refresh.md`
|
||||
- `agent-contract/index.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
- `Makefile`
|
||||
- `.gitignore`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- S03 requires identical fixture checksum, clean isolated workspace, fresh caller session, no history/resume reuse, and no `../iop-s2` mutation.
|
||||
- Evidence Map S03 requires checksum/containment/non-mutation tests. D05 and Milestone lines 39/51 explicitly separate runtime testbed source from result workspace.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- Local Python 3.12.3/Linux arm64; standard library and temporary Git fixtures only.
|
||||
- Current fixed testbed is `/config/workspace/iop-s2`, branch `feature/single-request-plan-review-templates`, HEAD `6b5ee6c366a26a30cec9180fe30ce442d4df27fd`, clean at review time. Branch name is provenance, not an acceptance constant.
|
||||
- No provider/network call and no write to the real sibling checkout. `bench-01` remains the unblocked `bench` lane head.
|
||||
- Predecessor index `01` has no active/archive `complete.log`; implementation must wait.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
No current test composes an already allocated attempt root with exclusive workspace/session children, validates the shared identity grammar, materializes declared fixture assets, creates fresh caller-session state, rejects destination escape/collision, or proves the runtime testbed is untouched.
|
||||
|
||||
### Symbol References
|
||||
|
||||
Resolve the completed predecessor's immutable manifest, asset mapping, digest, and CLI dispatcher symbols before editing. Do not duplicate manifest parsing.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
Workspace/session creation is one filesystem containment and provenance boundary. It independently returns immutable prepared-workspace metadata consumed by lifecycle/attempt code. Predecessor `01` is currently unsatisfied.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
Excluded: copying the IOP repository, a standalone public `prepare` command, caller execution, attempt allocation/retry, scoring, provider adapters, browser validation, and reports. The dependent attempt runner owns the public stateful command and supplies the empty attempt root.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- evaluation_mode: `isolated-reassessment`
|
||||
- finalizer: `finalize-task-policy.sh`, mode `pair`
|
||||
- build closures all `true`; scores `1/1/1/1/1`; base `local-fit`, route `risk-boundary`; grade `G05`; catalog `worker/cloud/G05`; filename `PLAN-cloud-G05.md`.
|
||||
- review closures all `true`; scores `1/1/1/1/1`; route `official-review`; grade `G05`; catalog `review/cloud/G05`; filename `CODE_REVIEW-cloud-G05.md`.
|
||||
- large_indivisible_context `false`; risks `temporal_state`, `boundary_contract`, `structured_interpretation`, `variant_product`; rework `0`; evidence integrity failure `false`; capability gap none.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement fixture-seeded workspace and empty caller-session materialization beneath an already allocated empty attempt root, with shared identity/path validation and no copy/write of the `../iop-s2` runtime testbed.
|
||||
- [ ] Add deterministic containment, checksum, collision, session-freshness, matrix-isolation, and testbed-nonmutation tests.
|
||||
- [ ] Resolve predecessor index `01`, then run focused, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Materialize one clean workspace and session
|
||||
|
||||
**Problem:** `SDD.md:69` and `SDD.md:102` require fixture identity plus a fresh caller session, while Milestone lines 39/51 state that `../iop-s2` is separate from result workspaces. The prior plan instead proposed copying the testbed and did not allocate isolated session state.
|
||||
|
||||
**Solution:** Add a standard-library `prepare_workspace` API accepting a validated manifest, an already exclusively allocated empty `attempt_root`, and a frozen identity. The identity grammar is shared with the dependent attempt store: `run-YYYYMMDDTHHMMSSZ-<12 lowercase hex>`, manifest cell id `^[a-z0-9][a-z0-9_-]{0,63}$`, positive repetition rendered as `repetition-%04d`, and `attempt-%06d`. Resolve and verify the non-symlink root is exactly `<output_root>/<run_id>/cells/<cell_id>/repetition-NNNN/attempt-NNNNNN`; reject absolute/traversal/symlink/foreign/non-empty roots. The caller owns only attempt-root allocation; this API atomically and exclusively creates its `workspace/` and `session/` children plus immutable `prepared.json`, so the two layers never create the same path.
|
||||
|
||||
Copy only manifest-declared asset `source` files to their canonical `workspace_path`; never copy the repository or prompt unless declared as an asset. Recompute the predecessor's canonical fixture checksum before copying and a workspace checksum after copying. Require empty session state and a unique session identity derived from the full attempt identity plus fresh randomness; expose that locator to later adapters without reading host-global conversation/resume directories. Record `setup_cache_policy="isolated"`; no writable shared cache is supported.
|
||||
|
||||
Resolve the exact `../iop-s2` path only for preflight provenance. Require a Git checkout and clean `git status --porcelain=v1 --untracked-files=all`, record branch/HEAD/status digest before and after preparation, and never traverse/copy/write its source tree.
|
||||
|
||||
Before (`SDD.md:102`):
|
||||
|
||||
```text
|
||||
same fixture -> clean workspace + fresh caller session; iop-s2 source is not changed
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
prepared = prepare_workspace(manifest, attempt_root, identity)
|
||||
assert prepared.workspace_checksum == manifest.fixture.checksum
|
||||
assert prepared.session_is_fresh
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_benchmark/workspace.py` with shared identity/path validation, exclusive child allocation under a caller-owned attempt root, asset mapping, checksum verification, empty session state, testbed provenance, and frozen metadata.
|
||||
- [ ] Update `scripts/agent_benchmark/__init__.py` to export only stable workspace APIs.
|
||||
|
||||
**Test Strategy:** Use temporary repository/testbed/output fixtures. Cover the exact run/cell/repetition/attempt grammar, missing/foreign/non-empty/symlink attempt roots, exclusive child collision, exact asset mapping, prompt exclusion unless declared, checksum mismatch, absolute/escaping/symlink destination, special-file source, unsupported cache policy, dirty testbed, and immutable metadata.
|
||||
|
||||
**Verification:** `python3 -m unittest scripts.agent_benchmark.workspace_test` exits 0.
|
||||
|
||||
### [API-2] Prove cross-attempt isolation and source integrity
|
||||
|
||||
**Problem:** A successful copy does not prove fresh conversation state, matrix isolation, or testbed non-mutation.
|
||||
|
||||
**Solution:** Exclusively allocate four empty attempt roots through a test seam, then prepare two cells and two repetitions from the same fixture. Assert four distinct workspace/session identities and identical initial workspace digests, mutate one workspace/session, and prove all peers remain byte-identical and empty. Capture testbed branch/HEAD/status digest before and after and assert equality. Assert no path under the real/temporary testbed appears beneath any workspace.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
No executable S03 evidence exists.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
self.assertEqual({p.workspace_checksum for p in prepared}, {manifest.fixture.checksum})
|
||||
self.assertEqual(len({p.session_id for p in prepared}), 4)
|
||||
self.assertEqual(testbed_before, testbed_after)
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_benchmark/workspace_test.py` with normal/boundary/isolation/session/non-mutation cases.
|
||||
- [ ] Record exact output in `CODE_REVIEW-cloud-G05.md`.
|
||||
|
||||
**Test Strategy:** Tests write only under `TemporaryDirectory`; one read-only preflight may inspect the configured real testbed, but unit tests never modify it.
|
||||
|
||||
**Verification:** Focused and aggregate tests pass and leave no temporary workspace/session state.
|
||||
|
||||
## Dependencies and Execution Order
|
||||
|
||||
1. Required predecessor index `01_benchmark_manifest` is encoded by `02+01_...`.
|
||||
2. At planning time no valid active/archive `01_*/complete.log` or `01+*/complete.log` exists. Runtime must wait.
|
||||
3. Before implementation, require exactly one matching completion across allowed active/monthly-archive locations; missing or multiple matches fail closed.
|
||||
4. Resolve final manifest/asset/digest symbols, implement API-1, then API-2.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Items |
|
||||
|------|-------|
|
||||
| `scripts/agent_benchmark/workspace.py` | API-1 |
|
||||
| `scripts/agent_benchmark/workspace_test.py` | API-2 |
|
||||
| `scripts/agent_benchmark/__init__.py` | API-1 |
|
||||
| `agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/CODE_REVIEW-cloud-G05.md` | API-2 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
1. `python3 -c 'from pathlib import Path; g="m-agent-comparison-benchmark-pipeline"; ids=("01",); a=Path("agent-task")/g; r=Path("agent-task/archive"); found={i:sorted([*a.glob(f"{i}_*/complete.log"),*a.glob(f"{i}+*/complete.log"),*r.glob(f"*/*/{g}/{i}_*/complete.log"),*r.glob(f"*/*/{g}/{i}+*/complete.log")],key=str) for i in ids}; bad={i:[str(p) for p in ps] for i,ps in found.items() if len(ps)!=1}; assert not bad,bad; print("\n".join(str(found[i][0]) for i in ids))'`
|
||||
- Expected: exactly one completion path for predecessor `01` before implementation.
|
||||
2. `python3 -m unittest scripts.agent_benchmark.workspace_test`
|
||||
- Expected: fixture, containment, session freshness, isolation, and testbed non-mutation cases pass.
|
||||
3. `make test-agent-comparison-benchmark`
|
||||
- Expected: all credential-free benchmark tests pass.
|
||||
4. `git diff --check`
|
||||
- Expected: exit 0 with no whitespace errors.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,123 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace plan=0 tag=API milestone-task=isolated-workspace -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-09
|
||||
task=m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace, plan=0, tag=API
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files and verify that output in `Verification Results` matches code.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals.
|
||||
2. Archive `CODE_REVIEW-cloud-G05.md` → `code_review_cloud_G05_0.log` and `PLAN-cloud-G05.md` → `plan_cloud_G05_0.log`.
|
||||
3. If PASS, write `complete.log` and move active task directory to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/`. If WARN/FAIL, fully write the next filesystem state required by the code-review skill.
|
||||
4. If PASS, preserve `milestone-task=isolated-workspace` in `complete.log` and report it for runtime aggregation; roadmap evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check applicable `Review-Only Checklist` items at the final `.log` location before reporting.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|--------|
|
||||
| API-1 Prepare a contained workspace | [ ] |
|
||||
| API-2 Prove isolation and source integrity | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement contained, exclusive, integrity-checked workspace preparation from a validated manifest.
|
||||
- [ ] Add deterministic normal, boundary, symlink, collision, and source-nonmutation tests.
|
||||
- [ ] Run predecessor, unit, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Archive active `CODE_REVIEW-*-G??.md` to `code_review_cloud_G05_0.log`.
|
||||
- [ ] Archive active `PLAN-*-G??.md` to `plan_cloud_G05_0.log`.
|
||||
- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write `complete.log` based on `agent-ops/skills/common/code-review/templates/complete-log-template.md` and leave no active `.md` files.
|
||||
- [ ] If PASS, move active task directory `agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/` to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/` and update this checklist at the final archive path.
|
||||
- [ ] If PASS, preserve and report `milestone-task=isolated-workspace` for runtime aggregation, without modifying roadmap or directly calling `update-roadmap`.
|
||||
- [ ] If PASS for split work, remove empty active parent only when no sibling remains.
|
||||
- [ ] If WARN/FAIL, write the next filesystem state matching code-review verdict and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record deviations and rationale._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record implementation decisions._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Workspace destinations are canonical, contained, and exclusively created.
|
||||
- Unsafe entries, dirty sources, collisions, and integrity drift fail closed.
|
||||
- Repetitions/cells are isolated and the source Git state and content remain unchanged.
|
||||
- No process, provider, retry, scoring, or reporting behavior entered this packet.
|
||||
|
||||
## Verification Results
|
||||
|
||||
### `test -f agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/complete.log`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 -m unittest scripts.agent_benchmark.workspace_test`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `make test-agent-comparison-benchmark`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `git diff --check`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results | Fixed headings/commands | Implementing agent fills actual output only |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,131 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace plan=1 tag=API milestone-task=isolated-workspace -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-09
|
||||
task=m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace, plan=1, tag=API
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/plan_cloud_G05_0.log` and `agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/code_review_cloud_G05_0.log`.
|
||||
- Review state: the prior pair was unimplemented and had no official verdict; it was archived by the requested Epic self-review replan.
|
||||
- Self-review defects: its predecessor check named only the active task path even though PASS moves it to `agent-task/archive/YYYY/MM/`, and its sibling-checkout preflight asserted a stale `dev` branch.
|
||||
- Scope carried forward: `isolated-workspace`, SDD S03, implementation files, and credential-free verification remain unchanged; dependency resolution, source provenance, and paired evidence are corrected.
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files and verify that output in `Verification Results` matches code.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals.
|
||||
2. Archive `CODE_REVIEW-cloud-G05.md` → `code_review_cloud_G05_1.log` and `PLAN-cloud-G05.md` → `plan_cloud_G05_1.log`.
|
||||
3. If PASS, write `complete.log` and move active task directory to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/`. If WARN/FAIL, fully write the next filesystem state required by the code-review skill.
|
||||
4. If PASS, preserve `milestone-task=isolated-workspace` in `complete.log` and report it for runtime aggregation; roadmap evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check applicable `Review-Only Checklist` items at the final `.log` location before reporting.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|--------|
|
||||
| API-1 Prepare a contained workspace | [ ] |
|
||||
| API-2 Prove isolation and source integrity | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement contained, exclusive, integrity-checked workspace preparation from a validated manifest.
|
||||
- [ ] Add deterministic normal, boundary, symlink, collision, and source-nonmutation tests.
|
||||
- [ ] Resolve the indexed predecessor from the allowed active/archive candidates, then run unit, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Archive active `CODE_REVIEW-*-G??.md` to `code_review_cloud_G05_1.log`.
|
||||
- [ ] Archive active `PLAN-*-G??.md` to `plan_cloud_G05_1.log`.
|
||||
- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write `complete.log` based on `agent-ops/skills/common/code-review/templates/complete-log-template.md` and leave no active `.md` files.
|
||||
- [ ] If PASS, move active task directory `agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/` to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/` and update this checklist at the final archive path.
|
||||
- [ ] If PASS, preserve and report `milestone-task=isolated-workspace` for runtime aggregation, without modifying roadmap or directly calling `update-roadmap`.
|
||||
- [ ] If PASS for split work, remove empty active parent only when no sibling remains.
|
||||
- [ ] If WARN/FAIL, write the next filesystem state matching code-review verdict and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record deviations and rationale._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record implementation decisions._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Workspace destinations are canonical, contained, and exclusively created.
|
||||
- Unsafe entries, dirty sources, collisions, and integrity drift fail closed.
|
||||
- Repetitions/cells are isolated, the source Git state and content remain unchanged, and the actual branch/HEAD is recorded without a hard-coded branch gate.
|
||||
- Predecessor index `01` resolves to exactly one active or monthly-archive `complete.log`; an absent or ambiguous match fails before implementation.
|
||||
- No process, provider, retry, scoring, or reporting behavior entered this packet.
|
||||
|
||||
## Verification Results
|
||||
|
||||
### `python3 -c 'from pathlib import Path; g="m-agent-comparison-benchmark-pipeline"; ids=("01",); a=Path("agent-task")/g; r=Path("agent-task/archive"); found={i:sorted([*a.glob(f"{i}_*/complete.log"),*a.glob(f"{i}+*/complete.log"),*r.glob(f"*/*/{g}/{i}_*/complete.log"),*r.glob(f"*/*/{g}/{i}+*/complete.log")],key=str) for i in ids}; bad={i:[str(p) for p in ps] for i,ps in found.items() if len(ps)!=1}; assert not bad,bad; print("\n".join(str(found[i][0]) for i in ids))'`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 -m unittest scripts.agent_benchmark.workspace_test`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `make test-agent-comparison-benchmark`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `git diff --check`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results | Fixed headings/commands | Implementing agent fills actual output only |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,132 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace plan=2 tag=API milestone-task=isolated-workspace -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-09
|
||||
task=m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace, plan=2, tag=API
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/plan_cloud_G05_1.log` and `agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/code_review_cloud_G05_1.log` (the prior snapshot points to generation 0).
|
||||
- Review state: unimplemented, no official verdict, replaced through explicit plan `write` mode.
|
||||
- Self-review defects: the prior plan treated `../iop-s2` as the workspace copy source and omitted the fresh caller-session/history boundary required by S03.
|
||||
- Scope carried forward: contained exclusive allocation, integrity evidence, source non-mutation, predecessor resolution, and credential-free tests.
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files and verify that output in `Verification Results` matches code.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals.
|
||||
2. Archive `CODE_REVIEW-cloud-G05.md` → `code_review_cloud_G05_2.log` and `PLAN-cloud-G05.md` → `plan_cloud_G05_2.log`.
|
||||
3. If PASS, write `complete.log` and move the active task directory to its monthly group archive. If WARN/FAIL, write the next state required by the code-review skill.
|
||||
4. If PASS, preserve `milestone-task=isolated-workspace` and report it for runtime aggregation; roadmap evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check applicable `Review-Only Checklist` items at the final `.log` location.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|---------|
|
||||
| API-1 Materialize one clean workspace and session | [ ] |
|
||||
| API-2 Prove cross-attempt isolation and source integrity | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement exclusive fixture-seeded workspace and empty caller-session materialization without copying or writing the `../iop-s2` runtime testbed.
|
||||
- [ ] Add deterministic containment, checksum, collision, session-freshness, matrix-isolation, and testbed-nonmutation tests.
|
||||
- [ ] Resolve predecessor index `01`, then run focused, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Archive active `CODE_REVIEW-*-G??.md` to `code_review_cloud_G05_2.log`.
|
||||
- [ ] Archive active `PLAN-*-G??.md` to `plan_cloud_G05_2.log`.
|
||||
- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write `complete.log` from the canonical template and leave no active `.md` files.
|
||||
- [ ] If PASS, move the active task directory to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/` and update this checklist there.
|
||||
- [ ] If PASS, preserve/report `milestone-task=isolated-workspace` without modifying roadmap or calling `update-roadmap` directly.
|
||||
- [ ] If PASS for split work, remove empty active parent or verify it remains for sibling files.
|
||||
- [ ] If WARN/FAIL, write the next filesystem state matching the verdict and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record any deviations from the plan and the rationale here._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record key design decisions here._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Exactly one predecessor completion is resolved before implementation and final manifest APIs are reused.
|
||||
- Only declared source/destination assets seed exclusive workspaces; prompt/testbed content is not copied implicitly.
|
||||
- Every cell/repetition/attempt gets a distinct empty session locator with no host history/resume reuse.
|
||||
- Testbed Git provenance is checked before/after without traversal or mutation; unsupported dirty/cache states fail closed.
|
||||
- Containment, symlink, collision, checksum, isolation, and atomic prepared-evidence boundaries are exercised.
|
||||
|
||||
## Verification Results
|
||||
|
||||
### Predecessor completion check from `PLAN-cloud-G05.md`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 -m unittest scripts.agent_benchmark.workspace_test`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `make test-agent-comparison-benchmark`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `git diff --check`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these |
|
||||
| Archive Evidence Snapshot | Fixed at stub creation from plan when present | Read only cited archive evidence when more detail is required |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results (section headings + commands) | Fixed at stub creation | Fill output only; changed commands require a deviation entry |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,159 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace plan=0 tag=API milestone-task=isolated-workspace -->
|
||||
|
||||
# Isolated Benchmark Workspace
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections of `CODE_REVIEW-cloud-G05.md` is the mandatory final implementation step. Run every verification command and paste actual output. Do not finalize, archive, write `complete.log`, ask the user, or change predecessor ownership.
|
||||
|
||||
## Background
|
||||
|
||||
Each matrix cell and repetition needs a clean, exclusive copy of the fixed `../iop-s2` testbed. This packet adds only contained workspace preparation and integrity evidence; process execution and retry state remain downstream.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `AGENTS.md`
|
||||
- `agent-ops/rules/project/rules.md`
|
||||
- `agent-ops/rules/project/domain/testing/rules.md`
|
||||
- `agent-test/local/rules.md`
|
||||
- `agent-test/local/testing-smoke.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/agent-comparison-benchmark-pipeline.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md`
|
||||
- `agent-spec/input/openai-compatible-surface.md`
|
||||
- `agent-spec/runtime/provider-pool-config-refresh.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
- `scripts/e2e-single-request-claude.sh`
|
||||
- `Makefile`
|
||||
- `.gitignore`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- Approved/released SDD scenario S03 requires a fresh isolated workspace for every cell/repetition and proof that the testbed remains unchanged.
|
||||
- Evidence Map row S03 directly becomes API-1's containment/exclusive-copy contract, API-2's matrix/source-nonmutation tests, and the focused plus aggregate Final Verification commands.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- Local Python 3.12.3/Linux arm64; standard library only; no network or provider call.
|
||||
- Handoff: no separate verification handoff was supplied; the SDD, repository-native harness, local test rules, and observed sibling checkout are the verification sources.
|
||||
- The fixed source checkout is `/config/workspace/iop-s2`, branch `dev`, observed clean at planning time. Tests must use temporary fixture repositories and must not mutate that checkout.
|
||||
- The existing single-request self-test has a baseline cleanup-bound failure outside this packet; it is not a pass criterion.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
No current test allocates one clean benchmark workspace per attempt or rejects symlink/special-file/path-escape inputs while preserving source integrity.
|
||||
|
||||
### Symbol References
|
||||
|
||||
The predecessor will introduce `scripts.agent_benchmark.manifest.load_manifest` and the public CLI dispatcher. Resolve exact symbols against the completed predecessor before editing; do not duplicate its parser.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
Workspace creation has filesystem side effects and a containment/integrity responsibility boundary, so it is not direct-small. Its stable output is prepared-workspace metadata consumable by lifecycle code. Predecessor `01_benchmark_manifest` is missing its active or archived `complete.log` at planning time.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
Excluded: caller execution, attempt retries, scoring, provider adapters, web validation, and report rendering.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- evaluation_mode: `first-pass`
|
||||
- finalizer: `finalize-task-policy.sh`, mode `pair`
|
||||
- build closures all `true`; scores `1/1/1/1/1`; route `risk-boundary`; grade `G05`; catalog `worker/cloud/G05`; filename `PLAN-cloud-G05.md`.
|
||||
- review closures all `true`; scores `1/1/1/1/1`; route `official-review`; grade `G05`; catalog `review/cloud/G05`; filename `CODE_REVIEW-cloud-G05.md`.
|
||||
- risks: `temporal_state`, `boundary_contract`, `structured_interpretation`, `variant_product`; large_indivisible_context `false`; rework `0`; evidence failure `false`; capability gap none.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement contained, exclusive, integrity-checked workspace preparation from a validated manifest.
|
||||
- [ ] Add deterministic normal, boundary, symlink, collision, and source-nonmutation tests.
|
||||
- [ ] Run predecessor, unit, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Prepare a contained workspace
|
||||
|
||||
**Problem:** The SDD requires per-attempt isolation, but no implementation currently proves that the copied workspace is fresh, contained, or source-preserving.
|
||||
|
||||
**Solution:** Add a standard-library workspace module that accepts only a validated manifest plus canonical cell/repetition/attempt identity. Resolve the testbed without following unsafe entries; reject path escape, symlink, FIFO/device/socket, dirty source, destination collision, and ambiguous setup state. Copy into an exclusively created `run/cells/<cell>/<repetition>/<attempt>/workspace` path, record source HEAD/tree and setup/cache policy, and verify source and copied regular-file digests before returning frozen metadata. Add a `prepare` CLI that uses the same implementation and emits only sanitized metadata.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
No benchmark workspace API or prepare command exists.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
prepared = prepare_workspace(manifest, cell_id, repetition, attempt_id)
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_benchmark/workspace.py` with validation, exclusive allocation, safe copy, and integrity metadata.
|
||||
- [ ] Update `scripts/agent_benchmark/__init__.py` to export the stable workspace API.
|
||||
- [ ] Update `scripts/agent_comparison_benchmark.py` with the `prepare` subcommand.
|
||||
|
||||
**Test Strategy:** Use temporary Git repositories and output roots. Cover a clean source, source dirtiness, symlink/special entries, escaping output paths, pre-existing destinations, checksum drift, and immutable metadata.
|
||||
|
||||
**Verification:** `python3 -m unittest scripts.agent_benchmark.workspace_test` exits 0.
|
||||
|
||||
### [API-2] Prove isolation and source integrity
|
||||
|
||||
**Problem:** A successful copy alone does not prove that repetitions are exclusive or that preparation left the fixed testbed untouched.
|
||||
|
||||
**Solution:** Add two-cell/two-repetition tests that allocate distinct destinations, mutate one copy, and prove all other copies and the source digest/HEAD/status remain unchanged. Assert the setup/cache policy is recorded rather than inferred.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
Copy isolation and source nonmutation have no executable evidence.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
self.assertNotEqual(first.workspace, second.workspace)
|
||||
self.assertEqual(source_before, source_after)
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_benchmark/workspace_test.py` with boundary, matrix-isolation, and source-nonmutation cases.
|
||||
- [ ] Record exact command output in `CODE_REVIEW-cloud-G05.md`.
|
||||
|
||||
**Test Strategy:** Tests run wholly under `TemporaryDirectory`; no tracked or sibling checkout is written.
|
||||
|
||||
**Verification:** The unit and aggregate targets pass, and `git diff --check` reports no errors.
|
||||
|
||||
## Dependencies and Execution Order
|
||||
|
||||
1. Required predecessor: `agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/complete.log`.
|
||||
2. At planning time that completion evidence is absent; do not implement this packet until it exists.
|
||||
3. After the predecessor completes, resolve its final manifest/CLI symbols, then implement API-1 followed by API-2.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Items |
|
||||
|------|-------|
|
||||
| `scripts/agent_benchmark/workspace.py` | API-1 |
|
||||
| `scripts/agent_benchmark/workspace_test.py` | API-2 |
|
||||
| `scripts/agent_benchmark/__init__.py` | API-1 |
|
||||
| `scripts/agent_comparison_benchmark.py` | API-1 |
|
||||
| `agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/CODE_REVIEW-cloud-G05.md` | API-2 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
1. `test -f agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/complete.log`
|
||||
- Expected: exit 0 before implementation begins.
|
||||
2. `python3 -m unittest scripts.agent_benchmark.workspace_test`
|
||||
- Expected: all containment, collision, isolation, and nonmutation cases pass.
|
||||
3. `make test-agent-comparison-benchmark`
|
||||
- Expected: the complete credential-free benchmark suite passes.
|
||||
4. `git diff --check`
|
||||
- Expected: exit 0 with no whitespace errors.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,168 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace plan=1 tag=API milestone-task=isolated-workspace -->
|
||||
|
||||
# Isolated Benchmark Workspace
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections of `CODE_REVIEW-cloud-G05.md` is the mandatory final implementation step. Run every verification command and paste actual output. Do not finalize, archive, write `complete.log`, ask the user, or change predecessor ownership.
|
||||
|
||||
## Background
|
||||
|
||||
Each matrix cell and repetition needs a clean, exclusive copy of the fixed `../iop-s2` testbed. This packet adds only contained workspace preparation and integrity evidence; process execution and retry state remain downstream.
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/plan_cloud_G05_0.log` and `agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/code_review_cloud_G05_0.log`.
|
||||
- Review state: the prior pair was unimplemented and had no official verdict; it was archived by the requested Epic self-review replan.
|
||||
- Self-review defects: its predecessor check named only the active task path even though PASS moves it to `agent-task/archive/YYYY/MM/`, and its sibling-checkout preflight asserted a stale `dev` branch.
|
||||
- Scope carried forward: `isolated-workspace`, SDD S03, implementation files, and credential-free verification remain unchanged; dependency resolution, source provenance, and paired evidence are corrected.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `AGENTS.md`
|
||||
- `agent-ops/rules/project/rules.md`
|
||||
- `agent-ops/rules/project/domain/testing/rules.md`
|
||||
- `agent-test/local/rules.md`
|
||||
- `agent-test/local/testing-smoke.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/agent-comparison-benchmark-pipeline.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md`
|
||||
- `agent-spec/input/openai-compatible-surface.md`
|
||||
- `agent-spec/runtime/provider-pool-config-refresh.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
- `agent-ops/skills/common/plan/SKILL.md`
|
||||
- `scripts/e2e-single-request-claude.sh`
|
||||
- `Makefile`
|
||||
- `.gitignore`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- Approved/released SDD scenario S03 requires a fresh isolated workspace for every cell/repetition and proof that the testbed remains unchanged.
|
||||
- Evidence Map row S03 directly becomes API-1's containment/exclusive-copy contract, API-2's matrix/source-nonmutation tests, and the focused plus aggregate Final Verification commands.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- Local Python 3.12.3/Linux arm64; standard library only; no network or provider call.
|
||||
- Handoff: no separate verification handoff was supplied; the SDD, repository-native harness, local test rules, and observed sibling checkout are the verification sources.
|
||||
- The fixed source checkout is `/config/workspace/iop-s2`; self-review observed branch `feature/single-request-plan-review-templates` with a clean worktree. The SDD fixes the `../iop-s2` testbed identity, not a branch name, so preparation must record the actual branch, HEAD, tree digest, and status and must not hard-code `dev`. Tests use temporary fixture repositories and must not mutate that checkout.
|
||||
- The existing single-request self-test has a baseline cleanup-bound failure outside this packet; it is not a pass criterion.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
No current test allocates one clean benchmark workspace per attempt or rejects symlink/special-file/path-escape inputs while preserving source integrity.
|
||||
|
||||
### Symbol References
|
||||
|
||||
The predecessor will introduce `scripts.agent_benchmark.manifest.load_manifest` and the public CLI dispatcher. Resolve exact symbols against the completed predecessor before editing; do not duplicate its parser.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
Workspace creation has filesystem side effects and a containment/integrity responsibility boundary, so it is not direct-small. Its stable output is prepared-workspace metadata consumable by lifecycle code. Predecessor `01_benchmark_manifest` is missing its active or archived `complete.log` at planning time.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
Excluded: caller execution, attempt retries, scoring, provider adapters, web validation, and report rendering.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- evaluation_mode: `isolated-reassessment`
|
||||
- finalizer: `finalize-task-policy.sh`, mode `pair`
|
||||
- build closures all `true`; scores `1/1/1/1/1`; route `risk-boundary`; grade `G05`; catalog `worker/cloud/G05`; filename `PLAN-cloud-G05.md`.
|
||||
- review closures all `true`; scores `1/1/1/1/1`; route `official-review`; grade `G05`; catalog `review/cloud/G05`; filename `CODE_REVIEW-cloud-G05.md`.
|
||||
- risks: `temporal_state`, `boundary_contract`, `structured_interpretation`, `variant_product`; large_indivisible_context `false`; rework `0`; evidence failure `false`; capability gap none.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement contained, exclusive, integrity-checked workspace preparation from a validated manifest.
|
||||
- [ ] Add deterministic normal, boundary, symlink, collision, and source-nonmutation tests.
|
||||
- [ ] Resolve the indexed predecessor from the allowed active/archive candidates, then run unit, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Prepare a contained workspace
|
||||
|
||||
**Problem:** The SDD requires per-attempt isolation, but no implementation currently proves that the copied workspace is fresh, contained, or source-preserving.
|
||||
|
||||
**Solution:** Add a standard-library workspace module that accepts only a validated manifest plus canonical cell/repetition/attempt identity. Resolve the testbed without following unsafe entries; reject path escape, symlink, FIFO/device/socket, dirty source, destination collision, and ambiguous setup state. Copy into an exclusively created `run/cells/<cell>/<repetition>/<attempt>/workspace` path, record source branch/HEAD/tree and setup/cache policy, and verify source and copied regular-file digests before returning frozen metadata. Add a `prepare` CLI that uses the same implementation and emits only sanitized metadata. The manifest's `environment=dev` and `testbed=../iop-s2` identify the environment and source path; they do not require a Git branch named `dev`, so any clean branch/HEAD is accepted and preserved as provenance.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
No benchmark workspace API or prepare command exists.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
prepared = prepare_workspace(manifest, cell_id, repetition, attempt_id)
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_benchmark/workspace.py` with validation, exclusive allocation, safe copy, and integrity metadata.
|
||||
- [ ] Update `scripts/agent_benchmark/__init__.py` to export the stable workspace API.
|
||||
- [ ] Update `scripts/agent_comparison_benchmark.py` with the `prepare` subcommand.
|
||||
|
||||
**Test Strategy:** Use temporary Git repositories and output roots. Cover a clean source, source dirtiness, symlink/special entries, escaping output paths, pre-existing destinations, checksum drift, and immutable metadata.
|
||||
|
||||
**Verification:** `python3 -m unittest scripts.agent_benchmark.workspace_test` exits 0.
|
||||
|
||||
### [API-2] Prove isolation and source integrity
|
||||
|
||||
**Problem:** A successful copy alone does not prove that repetitions are exclusive or that preparation left the fixed testbed untouched.
|
||||
|
||||
**Solution:** Add two-cell/two-repetition tests that allocate distinct destinations, mutate one copy, and prove all other copies and the source digest/HEAD/status remain unchanged. Assert the setup/cache policy is recorded rather than inferred.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
Copy isolation and source nonmutation have no executable evidence.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
self.assertNotEqual(first.workspace, second.workspace)
|
||||
self.assertEqual(source_before, source_after)
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_benchmark/workspace_test.py` with boundary, matrix-isolation, and source-nonmutation cases.
|
||||
- [ ] Record exact command output in `CODE_REVIEW-cloud-G05.md`.
|
||||
|
||||
**Test Strategy:** Tests run wholly under `TemporaryDirectory`; no tracked or sibling checkout is written.
|
||||
|
||||
**Verification:** The unit and aggregate targets pass, and `git diff --check` reports no errors.
|
||||
|
||||
## Dependencies and Execution Order
|
||||
|
||||
1. Required predecessor index: `01_benchmark_manifest`, encoded by the `02+01_...` task directory.
|
||||
2. At planning time no allowed active/archive `01_*/complete.log` or `01+*/complete.log` candidate exists. Runtime scheduling must wait; a completed predecessor may be under the active task group or `agent-task/archive/YYYY/MM/`.
|
||||
3. Before implementation, require exactly one matching predecessor completion record across those allowed locations. Multiple matches are ambiguous and must not be guessed.
|
||||
4. After the predecessor completes, resolve its final manifest/CLI symbols, then implement API-1 followed by API-2.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Items |
|
||||
|------|-------|
|
||||
| `scripts/agent_benchmark/workspace.py` | API-1 |
|
||||
| `scripts/agent_benchmark/workspace_test.py` | API-2 |
|
||||
| `scripts/agent_benchmark/__init__.py` | API-1 |
|
||||
| `scripts/agent_comparison_benchmark.py` | API-1 |
|
||||
| `agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/CODE_REVIEW-cloud-G05.md` | API-2 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
1. `python3 -c 'from pathlib import Path; g="m-agent-comparison-benchmark-pipeline"; ids=("01",); a=Path("agent-task")/g; r=Path("agent-task/archive"); found={i:sorted([*a.glob(f"{i}_*/complete.log"),*a.glob(f"{i}+*/complete.log"),*r.glob(f"*/*/{g}/{i}_*/complete.log"),*r.glob(f"*/*/{g}/{i}+*/complete.log")],key=str) for i in ids}; bad={i:[str(p) for p in ps] for i,ps in found.items() if len(ps)!=1}; assert not bad,bad; print("\n".join(str(found[i][0]) for i in ids))'`
|
||||
- Expected: exit 0 and print exactly one active or archived completion path for predecessor index `01` before implementation begins.
|
||||
2. `python3 -m unittest scripts.agent_benchmark.workspace_test`
|
||||
- Expected: all containment, collision, isolation, and nonmutation cases pass.
|
||||
3. `make test-agent-comparison-benchmark`
|
||||
- Expected: the complete credential-free benchmark suite passes.
|
||||
4. `git diff --check`
|
||||
- Expected: exit 0 with no whitespace errors.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,179 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace plan=2 tag=API milestone-task=isolated-workspace -->
|
||||
|
||||
# Isolated Benchmark Workspace
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections of `CODE_REVIEW-cloud-G05.md` is mandatory. Resolve the encoded predecessor before editing, run every verification command, paste actual output, and leave both active files in place for official review. Do not finalize, archive, write `complete.log`, ask the user, or change predecessor ownership.
|
||||
|
||||
## Background
|
||||
|
||||
Every cell/repetition needs a clean workspace and fresh caller-session locator derived from the same manifest fixture. `../iop-s2` is the dev IOP runtime testbed whose provenance and non-mutation are recorded; it is not copied into the scored workspace. Process execution and attempt retry remain downstream.
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/plan_cloud_G05_1.log` and `agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/code_review_cloud_G05_1.log` (the prior snapshot points to generation 0).
|
||||
- Review state: unimplemented, no official verdict, replaced through explicit plan `write` mode.
|
||||
- Self-review defects: the prior plan treated `../iop-s2` as the workspace copy source and omitted the fresh caller-session/history boundary required by S03.
|
||||
- Scope carried forward: contained exclusive allocation, integrity evidence, source non-mutation, predecessor resolution, and credential-free tests.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `AGENTS.md`
|
||||
- `agent-ops/rules/project/rules.md`
|
||||
- `agent-ops/rules/project/domain/testing/rules.md`
|
||||
- `agent-ops/rules/common/rules-roadmap.md`
|
||||
- `agent-ops/skills/common/plan/SKILL.md`
|
||||
- `agent-ops/skills/common/finalize-task-routing/SKILL.md`
|
||||
- `agent-test/local/rules.md`
|
||||
- `agent-test/local/testing-smoke.md`
|
||||
- `agent-roadmap/current.md`
|
||||
- `agent-roadmap/priority-queue.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/PHASE.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/agent-comparison-benchmark-pipeline.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md`
|
||||
- `agent-spec/index.md`
|
||||
- `agent-spec/input/openai-compatible-surface.md`
|
||||
- `agent-spec/runtime/provider-pool-config-refresh.md`
|
||||
- `agent-contract/index.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
- `Makefile`
|
||||
- `.gitignore`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- S03 requires identical fixture checksum, clean isolated workspace, fresh caller session, no history/resume reuse, and no `../iop-s2` mutation.
|
||||
- Evidence Map S03 requires checksum/containment/non-mutation tests. D05 and Milestone lines 39/51 explicitly separate runtime testbed source from result workspace.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- Local Python 3.12.3/Linux arm64; standard library and temporary Git fixtures only.
|
||||
- Current fixed testbed is `/config/workspace/iop-s2`, branch `feature/single-request-plan-review-templates`, HEAD `6b5ee6c366a26a30cec9180fe30ce442d4df27fd`, clean at review time. Branch name is provenance, not an acceptance constant.
|
||||
- No provider/network call and no write to the real sibling checkout. `bench-01` remains the unblocked `bench` lane head.
|
||||
- Predecessor index `01` has no active/archive `complete.log`; implementation must wait.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
No current test materializes declared fixture assets into exclusive attempt roots, creates fresh caller-session state, rejects destination escape/collision, or proves the runtime testbed is untouched.
|
||||
|
||||
### Symbol References
|
||||
|
||||
Resolve the completed predecessor's immutable manifest, asset mapping, digest, and CLI dispatcher symbols before editing. Do not duplicate manifest parsing.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
Workspace/session creation is one filesystem containment and provenance boundary. It independently returns immutable prepared-workspace metadata consumed by lifecycle/attempt code. Predecessor `01` is currently unsatisfied.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
Excluded: copying the IOP repository, caller execution, attempts/retry, scoring, provider adapters, browser validation, and reports.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- evaluation_mode: `isolated-reassessment`
|
||||
- finalizer: `finalize-task-policy.sh`, mode `pair`
|
||||
- build closures all `true`; scores `1/1/1/1/1`; base `local-fit`, route `risk-boundary`; grade `G05`; catalog `worker/cloud/G05`; filename `PLAN-cloud-G05.md`.
|
||||
- review closures all `true`; scores `1/1/1/1/1`; route `official-review`; grade `G05`; catalog `review/cloud/G05`; filename `CODE_REVIEW-cloud-G05.md`.
|
||||
- large_indivisible_context `false`; risks `temporal_state`, `boundary_contract`, `structured_interpretation`, `variant_product`; rework `0`; evidence integrity failure `false`; capability gap none.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement exclusive fixture-seeded workspace and empty caller-session materialization without copying or writing the `../iop-s2` runtime testbed.
|
||||
- [ ] Add deterministic containment, checksum, collision, session-freshness, matrix-isolation, and testbed-nonmutation tests.
|
||||
- [ ] Resolve predecessor index `01`, then run focused, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Materialize one clean workspace and session
|
||||
|
||||
**Problem:** `SDD.md:69` and `SDD.md:102` require fixture identity plus a fresh caller session, while Milestone lines 39/51 state that `../iop-s2` is separate from result workspaces. The prior plan instead proposed copying the testbed and did not allocate isolated session state.
|
||||
|
||||
**Solution:** Add a standard-library `prepare_workspace` API accepting only a validated manifest plus canonical run/cell/repetition/attempt identity. Exclusively create `<output_root>/<run_id>/cells/<cell_id>/<repetition>/<attempt_id>/`, with sibling `workspace/`, `session/`, and immutable `prepared.json`. Copy only manifest-declared asset `source` files to their canonical `workspace_path`; never copy the repository or prompt into the workspace unless declared as an asset. Recompute the predecessor's canonical fixture checksum before copying and a workspace checksum after copying. Require empty session state and a unique session identity; expose that locator to later adapters without reading host-global conversation/resume directories. Record the declared setup/cache policy; the initial implementation supports no writable shared cache and fails closed on unsupported policy.
|
||||
|
||||
Resolve the exact `../iop-s2` path only for preflight provenance. Require a Git checkout and clean `git status --porcelain=v1 --untracked-files=all`, record branch/HEAD/status digest before and after preparation, and never traverse/copy/write its source tree.
|
||||
|
||||
Before (`SDD.md:102`):
|
||||
|
||||
```text
|
||||
same fixture -> clean workspace + fresh caller session; iop-s2 source is not changed
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
prepared = prepare_workspace(manifest, run_id, cell_id, repetition, attempt_id)
|
||||
assert prepared.workspace_checksum == manifest.fixture.checksum
|
||||
assert prepared.session_is_fresh
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_benchmark/workspace.py` with exclusive allocation, asset mapping, checksum verification, empty session state, testbed provenance, and frozen metadata.
|
||||
- [ ] Update `scripts/agent_benchmark/__init__.py` to export only stable workspace APIs.
|
||||
- [ ] Update `scripts/agent_comparison_benchmark.py` with a sanitized `prepare` subcommand.
|
||||
|
||||
**Test Strategy:** Use temporary repository/testbed/output fixtures. Cover exact asset mapping, prompt exclusion unless declared, checksum mismatch, absolute/escaping/symlink destination, special-file source, collision, unsupported cache policy, dirty testbed, and immutable metadata.
|
||||
|
||||
**Verification:** `python3 -m unittest scripts.agent_benchmark.workspace_test` exits 0.
|
||||
|
||||
### [API-2] Prove cross-attempt isolation and source integrity
|
||||
|
||||
**Problem:** A successful copy does not prove fresh conversation state, matrix isolation, or testbed non-mutation.
|
||||
|
||||
**Solution:** Prepare two cells and two repetitions from the same fixture, assert four distinct workspace/session identities and identical initial workspace digests, mutate one workspace/session, and prove all peers remain byte-identical and empty. Capture testbed branch/HEAD/status digest before and after and assert equality. Assert no path under the real/temporary testbed appears beneath any workspace.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
No executable S03 evidence exists.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
self.assertEqual({p.workspace_checksum for p in prepared}, {manifest.fixture.checksum})
|
||||
self.assertEqual(len({p.session_id for p in prepared}), 4)
|
||||
self.assertEqual(testbed_before, testbed_after)
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_benchmark/workspace_test.py` with normal/boundary/isolation/session/non-mutation cases.
|
||||
- [ ] Record exact output in `CODE_REVIEW-cloud-G05.md`.
|
||||
|
||||
**Test Strategy:** Tests write only under `TemporaryDirectory`; one read-only preflight may inspect the configured real testbed, but unit tests never modify it.
|
||||
|
||||
**Verification:** Focused and aggregate tests pass and leave no temporary workspace/session state.
|
||||
|
||||
## Dependencies and Execution Order
|
||||
|
||||
1. Required predecessor index `01_benchmark_manifest` is encoded by `02+01_...`.
|
||||
2. At planning time no valid active/archive `01_*/complete.log` or `01+*/complete.log` exists. Runtime must wait.
|
||||
3. Before implementation, require exactly one matching completion across allowed active/monthly-archive locations; missing or multiple matches fail closed.
|
||||
4. Resolve final manifest/asset/digest symbols, implement API-1, then API-2.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Items |
|
||||
|------|-------|
|
||||
| `scripts/agent_benchmark/workspace.py` | API-1 |
|
||||
| `scripts/agent_benchmark/workspace_test.py` | API-2 |
|
||||
| `scripts/agent_benchmark/__init__.py` | API-1 |
|
||||
| `scripts/agent_comparison_benchmark.py` | API-1 |
|
||||
| `agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/CODE_REVIEW-cloud-G05.md` | API-2 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
1. `python3 -c 'from pathlib import Path; g="m-agent-comparison-benchmark-pipeline"; ids=("01",); a=Path("agent-task")/g; r=Path("agent-task/archive"); found={i:sorted([*a.glob(f"{i}_*/complete.log"),*a.glob(f"{i}+*/complete.log"),*r.glob(f"*/*/{g}/{i}_*/complete.log"),*r.glob(f"*/*/{g}/{i}+*/complete.log")],key=str) for i in ids}; bad={i:[str(p) for p in ps] for i,ps in found.items() if len(ps)!=1}; assert not bad,bad; print("\n".join(str(found[i][0]) for i in ids))'`
|
||||
- Expected: exactly one completion path for predecessor `01` before implementation.
|
||||
2. `python3 -m unittest scripts.agent_benchmark.workspace_test`
|
||||
- Expected: fixture, containment, session freshness, isolation, and testbed non-mutation cases pass.
|
||||
3. `make test-agent-comparison-benchmark`
|
||||
- Expected: all credential-free benchmark tests pass.
|
||||
4. `git diff --check`
|
||||
- Expected: exit 0 with no whitespace errors.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,134 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle plan=3 tag=API milestone-task=run-lifecycle -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-09
|
||||
task=m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle, plan=3, tag=API
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/plan_cloud_G08_2.log` and `agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/code_review_cloud_G09_2.log` (generation 2 retains earlier history).
|
||||
- Review state: unimplemented, no official verdict, replaced through explicit plan `write` mode.
|
||||
- Self-review defects: the prior parent-owned `Popen` then `on_started` callback left a controller-crash window in which a caller could exist without a durable locator. It also promised arbitrary descendant verification although a portable POSIX harness can safely own only the caller's dedicated process group.
|
||||
- Scope carried forward: generic lifecycle/events, bounded capture, timeout/cancel races, predecessor resolution, and credential-free subprocess tests.
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files and verify that output in `Verification Results` matches code.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and verified routing signals.
|
||||
2. Archive this review as `code_review_cloud_G09_3.log` and the plan as `plan_cloud_G08_3.log`.
|
||||
3. If PASS, create `complete.log` and move the task directory to its monthly group archive; otherwise write the required next state.
|
||||
4. If PASS, preserve/report `milestone-task=run-lifecycle`; roadmap evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check the review-only list at the final log location.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|---------|
|
||||
| API-1 Execute one bounded invocation | [ ] |
|
||||
| API-2 Clean every owned process group | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement a registered supervisor that durably proves ownership before launching exactly one caller and performing exactly one harness-owned task submission, plus normalized event journal, bounded/redacted capture, and strict finish→idle→quiet completion policies.
|
||||
- [ ] Implement single-owner terminal arbitration, controller-loss handling, authenticated recovery, and bounded owned-process-group cleanup on success, failure, timeout, cancel, malformed events, and reader errors.
|
||||
- [ ] Resolve predecessor `01`, add deterministic registration/crash/race/orphan/redaction tests, and run focused, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Archive active review to `code_review_cloud_G09_3.log`.
|
||||
- [ ] Archive active plan to `plan_cloud_G08_3.log`.
|
||||
- [ ] Verify `.gitignore` managed rules unignore task markdown/logs and ignore `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write canonical `complete.log` and leave no active `.md` files.
|
||||
- [ ] If PASS, move to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/` and update this checklist there.
|
||||
- [ ] If PASS, preserve/report `milestone-task=run-lifecycle` without directly changing roadmap.
|
||||
- [ ] If PASS for split work, remove empty parent or verify remaining sibling ownership.
|
||||
- [ ] If WARN/FAIL, write the next state and do not create `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record any deviations from the plan and the rationale here._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record key design decisions here._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- The controller launches only the internal supervisor first; the supervisor durably registers a private authenticated control endpoint, and the caller cannot start until `on_started` commits the locator and sends `START`.
|
||||
- Controller loss before `START` launches no caller; controller loss after `START` routes through the supervisor's bounded cleanup without persisting raw argv/environment.
|
||||
- Closed `argv_task`/`stdin_once` modes produce one harness-owned `submitted` event; caller output cannot synthesize it and no second harness input is accepted.
|
||||
- Normalized submission→finish→idle ordering plus quiet window is enforced for both closed completion modes.
|
||||
- Output bounds and exact/fallback redaction apply before atomic journal/result publication.
|
||||
- One terminal arbiter covers success, nonzero, malformed event, reader failure, timeout, cancel, and races.
|
||||
- Every return path gracefully stops/escalates/reaps/drains/verifies the dedicated caller group before publication; authenticated recovery refuses stale/forged/mismatched ownership and never blindly signals a recorded pid.
|
||||
|
||||
## Verification Results
|
||||
|
||||
### Predecessor completion check from `PLAN-cloud-G08.md`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 -m unittest scripts.agent_benchmark.lifecycle_test`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `make test-agent-comparison-benchmark`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `git diff --check`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementer must not modify or execute these |
|
||||
| Archive Evidence Snapshot | Fixed at stub creation from plan when present | Read only cited archive evidence when required |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementer checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementer checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementer must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholders with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results (section headings + commands) | Fixed at stub creation | Fill output only; changes require a deviation entry |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,180 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle plan=3 tag=API milestone-task=run-lifecycle -->
|
||||
|
||||
# Benchmark Run Lifecycle
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections of `CODE_REVIEW-cloud-G09.md` is mandatory. Resolve predecessor `01`, run every verification command, paste actual output, and leave the active pair for official review. Do not finalize, archive, write `complete.log`, ask the user, or change predecessor ownership.
|
||||
|
||||
## Background
|
||||
|
||||
Caller adapters need one generic, bounded subprocess lifecycle. This packet owns exactly-one invocation, normalized finish/idle evidence, output quiescence, redaction, completion policy, timeout/cancel, and descendant cleanup. It does not encode Claude, agy, or Codex command lines.
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/plan_cloud_G08_2.log` and `agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/code_review_cloud_G09_2.log` (generation 2 retains earlier history).
|
||||
- Review state: unimplemented, no official verdict, replaced through explicit plan `write` mode.
|
||||
- Self-review defects: the prior parent-owned `Popen` then `on_started` callback left a controller-crash window in which a caller could exist without a durable locator. It also promised arbitrary descendant verification although a portable POSIX harness can safely own only the caller's dedicated process group.
|
||||
- Scope carried forward: generic lifecycle/events, bounded capture, timeout/cancel races, predecessor resolution, and credential-free subprocess tests.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `AGENTS.md`
|
||||
- `agent-ops/rules/project/rules.md`
|
||||
- `agent-ops/rules/project/domain/testing/rules.md`
|
||||
- `agent-ops/rules/common/rules-roadmap.md`
|
||||
- `agent-ops/skills/common/plan/SKILL.md`
|
||||
- `agent-ops/skills/common/code-review/SKILL.md`
|
||||
- `agent-ops/skills/common/finalize-task-routing/SKILL.md`
|
||||
- `agent-ops/skills/common/plan/templates/review-stub-template.md`
|
||||
- `agent-test/local/rules.md`
|
||||
- `agent-test/local/testing-smoke.md`
|
||||
- `agent-roadmap/current.md`
|
||||
- `agent-roadmap/priority-queue.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/PHASE.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/agent-comparison-benchmark-pipeline.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md`
|
||||
- `agent-spec/index.md`
|
||||
- `agent-spec/input/openai-compatible-surface.md`
|
||||
- `agent-contract/index.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `Makefile`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- S04 requires one user submission, caller-specific finish/complete normalization, idle, bounded waiting, and a common terminal timeline.
|
||||
- State invariant line 72 requires idle plus process/output quiescence; Evidence Map S04 requires fixture streams and a real subprocess lifecycle probe.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- Local Python 3.12.3/Linux arm64. Tests use real temporary subprocesses/process groups and deterministic event fixtures; no external provider/network/credential.
|
||||
- The `bench-01` lane and SDD gates are open. Predecessor `01` is not complete, so runtime scheduling must wait.
|
||||
- The lifecycle must be POSIX-safe for the current Linux/macOS benchmark target. Unsupported platforms fail before launch rather than weakening cleanup.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
No current shared module proves durable registration before caller launch, exactly one harness-owned user-task submission, controller-loss cleanup, ordered finish→idle→quiet, both exit and harness-stop completion modes, bounded/redacted capture, or no surviving owned process-group members on success, error, malformed event, timeout, and cancel.
|
||||
|
||||
### Symbol References
|
||||
|
||||
Resolve the completed manifest's timeout types before implementation. The lifecycle accepts an immutable invocation specification plus parser/redactor hooks and must not import future caller adapters.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
Process state, event ordering, output readers, terminal arbitration, signals, and cleanup share one race-sensitive invariant. Splitting would allow success publication before cleanup, so the packet stays atomic.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
Excluded: caller command construction/protocol parsing, workspace creation, attempt persistence/recovery, provider usage normalization, scoring, browser checks, and reports.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- evaluation_mode: `isolated-reassessment`
|
||||
- finalizer: `finalize-task-policy.sh`, mode `pair`
|
||||
- build closures all `true`; scores `2/2/1/1/2`; base `local-fit`, route `risk-boundary`; grade `G08`; catalog `worker/cloud/G08`; filename `PLAN-cloud-G08.md`.
|
||||
- review closures all `true`; scores `2/2/1/2/2`; route `official-review`; grade `G09`; catalog `review/cloud/G09`; filename `CODE_REVIEW-cloud-G09.md`.
|
||||
- large_indivisible_context `false`; risks `temporal_state`, `concurrent_consistency`, `boundary_contract`, `structured_interpretation`, `variant_product`; rework `0`; evidence integrity failure `false`; capability gap none.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement a registered supervisor that durably proves ownership before launching exactly one caller and performing exactly one harness-owned task submission, plus normalized event journal, bounded/redacted capture, and strict finish→idle→quiet completion policies.
|
||||
- [ ] Implement single-owner terminal arbitration, controller-loss handling, authenticated recovery, and bounded owned-process-group cleanup on success, failure, timeout, cancel, malformed events, and reader errors.
|
||||
- [ ] Resolve predecessor `01`, add deterministic registration/crash/race/orphan/redaction tests, and run focused, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Execute one bounded invocation
|
||||
|
||||
**Problem:** `SDD.md:72`, `SDD.md:85`, and S04 require finish/complete, idle, and quiescence, but no generic runner exists. Natural process exit cannot be the only success policy because some callers remain alive while idle.
|
||||
|
||||
**Solution:** Add frozen invocation/result/event/supervisor-locator types and an internal supervisor mode in `lifecycle.py`. The invocation has one closed `submission_mode`: `argv_task` means the adapter already placed the one user task in the in-memory argv, while `stdin_once` supplies one bounded task payload that the supervisor writes exactly once and then closes the submission channel. No mode permits a second harness input. The controller starts only the supervisor in a new POSIX session and passes the caller specification over private inherited pipes; it never launches the caller directly. The supervisor creates a mode-`0600` Unix-domain control endpoint, atomically writes a locator containing supervisor pid, process start identity, socket path, and an unguessable challenge marker, then sends `REGISTERED`. The controller must durably commit that exact locator through required `on_started(locator)` and send `START` only after the callback succeeds. If the controller disappears or callback fails before `START`, the supervisor exits without launching a caller.
|
||||
|
||||
After `START`, the supervisor launches exactly one argv in a separate owned process group with explicit cwd and a minimal allowlisted environment. It emits the sole harness-owned `submitted` event only after successful child exec for `argv_task`, or after the sole full stdin write for `stdin_once`; caller output can never synthesize or duplicate submission. It then proxies bounded output/control frames and cleans the caller group on controller-channel EOF. Raw argv, environment, prompt, and secret values are never written to locator/evidence files.
|
||||
|
||||
Recovery authenticates the live supervisor by connecting to the recorded socket and completing a marker challenge before requesting status/stop; pid/start checks are corroboration and never authorize an OS signal alone. A missing/stale socket, failed challenge, or mismatched start identity fails closed. Inject a caller parser in the controller that converts proxied native output only to closed `finish` and `idle` terminal evidence; optional later metric events remain data and cannot synthesize submission or completion. Capture stdout/stderr with hard byte/line bounds, exact secret-value redaction supplied by the adapter, secret-shaped fallback redaction, truncation markers, and monotonic observation times. Publish JSONL journal and terminal result through temp-file + fsync + replace only after readers stop and authenticated supervisor cleanup completes.
|
||||
|
||||
Support two closed completion modes in the invocation spec: `exit_after_idle` requires exit 0 after ordered finish→idle→quiet; `stop_after_idle` treats ordered finish→idle→quiet as task completion, then performs a separate expected graceful harness stop. In both modes, duplicate/out-of-order/malformed events, nonzero unexpected exit, or missing idle fail closed.
|
||||
|
||||
Before (`SDD.md:103`):
|
||||
|
||||
```text
|
||||
one submission -> finish/complete -> idle -> common terminal outcome
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
result = run_invocation(spec, parse_event=parse_event, redact=redactor, cancellation=token,
|
||||
on_started=persist_locator)
|
||||
assert (not result.success) or result.finish_then_idle_then_quiet
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_benchmark/lifecycle.py` with immutable contracts, registered internal supervisor/control protocol, durable callback-before-START gate, readers, event validation, redaction, completion modes, authenticated recovery, and atomic evidence publication.
|
||||
- [ ] Update `scripts/agent_benchmark/__init__.py` to export only stable lifecycle APIs.
|
||||
|
||||
**Test Strategy:** Use real temporary executable scripts. Cover both completion modes, exact invocation/submission count for `argv_task` and `stdin_once`, second-input rejection, caller-spoofed `submitted` rejection, locator fsync/replace before `START`, callback failure and controller loss before/after `START`, exit/nonzero, missing/duplicate/out-of-order finish/idle, malformed parser output, reader error, capture truncation, exact/fallback redaction, quiet-window boundaries, and journal/result immutability.
|
||||
|
||||
**Verification:** `python3 -m unittest scripts.agent_benchmark.lifecycle_test` exits 0.
|
||||
|
||||
### [API-2] Clean every owned process group
|
||||
|
||||
**Problem:** `SDD.md:65` requires cleanup evidence. An owned caller process-group member surviving any return path can contaminate later attempts; timeout/cancel-only cleanup is insufficient. Claiming cleanup for descendants that deliberately detach into another session would exceed the portable ownership boundary.
|
||||
|
||||
**Solution:** Route success, nonzero exit, event failure, output-reader failure, timeout, explicit cancel, and controller EOF through the supervisor's one terminal arbiter and cleanup path. If the caller group is live, send the mode-appropriate graceful signal, wait the manifest-bounded cleanup grace, escalate to group kill, reap the leader, drain/join proxy readers, and verify `killpg(..., 0)` reports no owned member before acknowledging cleanup. The supervisor itself remains outside the caller group and exits only after writing an atomic cleanup receipt. `recover_invocation(locator)` uses the authenticated control endpoint to request this same cleanup and verifies the receipt; it never signals a pid/group from an unauthenticated or stale locator. Adapters that deliberately daemonize/detach are unsupported and must fail preflight rather than be claimed as contained. Near-deadline finish/timeout/cancel races select exactly one terminal reason; cleanup failure is terminal and cannot be relabeled success.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
No shared all-terminal cleanup owner exists.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
terminal = supervisor.finish(reason)
|
||||
assert terminal.cleanup_complete
|
||||
assert not terminal.process_group_alive
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Complete all-path supervisor/group cleanup, challenge-authenticated `recover_invocation`, and cleanup receipts in `scripts/agent_benchmark/lifecycle.py`.
|
||||
- [ ] Add `scripts/agent_benchmark/lifecycle_test.py` with owned-group cases for natural success, nonzero exit, malformed event, ignored graceful stop, controller loss, near-deadline timeout, explicit cancel, reader failure, and stale/forged locators.
|
||||
- [ ] Record exact output in `CODE_REVIEW-cloud-G09.md`.
|
||||
|
||||
**Test Strategy:** Poll process identities with bounded deadlines; do not use unbounded or arbitrary long sleeps. Assert no owned fixture group member survives each case, authenticated recovery succeeds, forged socket/marker and pid/start mismatch refuse cleanup, controller death leaves either no caller or a supervisor that cleans it, and cleanup latency stays below a documented generous bound.
|
||||
|
||||
**Verification:** Focused and aggregate tests pass with no surviving fixture process.
|
||||
|
||||
## Dependencies and Execution Order
|
||||
|
||||
1. Required predecessor `01_benchmark_manifest` is encoded by `03+01_...`.
|
||||
2. No valid active/archive predecessor completion exists at planning time; runtime waits.
|
||||
3. Before implementation, require exactly one matching allowed completion record; missing/multiple matches fail closed.
|
||||
4. Resolve timeout/value symbols, implement API-1 and API-2 as one terminal invariant, then verify from a clean process state.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Items |
|
||||
|------|-------|
|
||||
| `scripts/agent_benchmark/lifecycle.py` | API-1, API-2 |
|
||||
| `scripts/agent_benchmark/lifecycle_test.py` | API-2 |
|
||||
| `scripts/agent_benchmark/__init__.py` | API-1 |
|
||||
| `agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/CODE_REVIEW-cloud-G09.md` | API-2 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
1. `python3 -c 'from pathlib import Path; g="m-agent-comparison-benchmark-pipeline"; ids=("01",); a=Path("agent-task")/g; r=Path("agent-task/archive"); found={i:sorted([*a.glob(f"{i}_*/complete.log"),*a.glob(f"{i}+*/complete.log"),*r.glob(f"*/*/{g}/{i}_*/complete.log"),*r.glob(f"*/*/{g}/{i}+*/complete.log")],key=str) for i in ids}; bad={i:[str(p) for p in ps] for i,ps in found.items() if len(ps)!=1}; assert not bad,bad; print("\n".join(str(found[i][0]) for i in ids))'`
|
||||
- Expected: exactly one completion path for predecessor `01` before implementation.
|
||||
2. `python3 -m unittest scripts.agent_benchmark.lifecycle_test`
|
||||
- Expected: event, completion-mode, redaction, race, timeout/cancel, and all-terminal cleanup cases pass.
|
||||
3. `make test-agent-comparison-benchmark`
|
||||
- Expected: all credential-free benchmark tests pass with no fixture process left alive.
|
||||
4. `git diff --check`
|
||||
- Expected: exit 0 with no whitespace errors.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,123 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle plan=0 tag=API milestone-task=run-lifecycle -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-09
|
||||
task=m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle, plan=0, tag=API
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files and verify that output in `Verification Results` matches code.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals.
|
||||
2. Archive `CODE_REVIEW-cloud-G09.md` → `code_review_cloud_G09_0.log` and `PLAN-cloud-G08.md` → `plan_cloud_G08_0.log`.
|
||||
3. If PASS, write `complete.log` and move active task directory to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/`. If WARN/FAIL, fully write the next filesystem state required by the code-review skill.
|
||||
4. If PASS, preserve `milestone-task=run-lifecycle` in `complete.log` and report it for runtime aggregation; roadmap evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check applicable `Review-Only Checklist` items at the final `.log` location before reporting.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|--------|
|
||||
| API-1 Execute one bounded invocation | [ ] |
|
||||
| API-2 Bound timeout and cancellation cleanup | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement the generic one-process lifecycle, bounded capture, normalized event journal, and strict completion gate.
|
||||
- [ ] Implement cancellation/timeout escalation and descendant cleanup with deterministic race tests.
|
||||
- [ ] Run predecessor, lifecycle, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Archive active `CODE_REVIEW-*-G??.md` to `code_review_cloud_G09_0.log`.
|
||||
- [ ] Archive active `PLAN-*-G??.md` to `plan_cloud_G08_0.log`.
|
||||
- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write `complete.log` based on `agent-ops/skills/common/code-review/templates/complete-log-template.md` and leave no active `.md` files.
|
||||
- [ ] If PASS, move active task directory `agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/` to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/` and update this checklist at the final archive path.
|
||||
- [ ] If PASS, preserve and report `milestone-task=run-lifecycle` for runtime aggregation, without modifying roadmap or directly calling `update-roadmap`.
|
||||
- [ ] If PASS for split work, remove empty active parent only when no sibling remains.
|
||||
- [ ] If WARN/FAIL, write the next filesystem state matching code-review verdict and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record deviations and rationale._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record implementation decisions._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Exactly one child invocation occurs and success requires ordered finish, idle, quiet, and exit 0.
|
||||
- Capture/event artifacts are bounded, sanitized, immutable, and atomically published.
|
||||
- Timeout/cancel races have one terminal outcome; descendants receive escalation and are reaped.
|
||||
- No caller-specific adapter, provider protocol, retry, scoring, or reporting behavior entered this packet.
|
||||
|
||||
## Verification Results
|
||||
|
||||
### `test -f agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/complete.log`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 -m unittest scripts.agent_benchmark.lifecycle_test`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `make test-agent-comparison-benchmark`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `git diff --check`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results | Fixed headings/commands | Implementing agent fills actual output only |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,131 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle plan=1 tag=API milestone-task=run-lifecycle -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-09
|
||||
task=m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle, plan=1, tag=API
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/plan_cloud_G08_0.log` and `agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/code_review_cloud_G09_0.log`.
|
||||
- Review state: the prior pair was unimplemented and had no official verdict; it was archived by the requested Epic self-review replan.
|
||||
- Self-review defect: its predecessor check named only the active task path, but a PASS predecessor is normally moved under `agent-task/archive/YYYY/MM/`.
|
||||
- Scope carried forward: `run-lifecycle`, SDD S04, implementation files, and credential-free verification remain unchanged; only dependency resolution and paired evidence are corrected.
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files and verify that output in `Verification Results` matches code.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals.
|
||||
2. Archive `CODE_REVIEW-cloud-G09.md` → `code_review_cloud_G09_1.log` and `PLAN-cloud-G08.md` → `plan_cloud_G08_1.log`.
|
||||
3. If PASS, write `complete.log` and move active task directory to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/`. If WARN/FAIL, fully write the next filesystem state required by the code-review skill.
|
||||
4. If PASS, preserve `milestone-task=run-lifecycle` in `complete.log` and report it for runtime aggregation; roadmap evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check applicable `Review-Only Checklist` items at the final `.log` location before reporting.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|--------|
|
||||
| API-1 Execute one bounded invocation | [ ] |
|
||||
| API-2 Bound timeout and cancellation cleanup | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement the generic one-process lifecycle, bounded capture, normalized event journal, and strict completion gate.
|
||||
- [ ] Implement cancellation/timeout escalation and descendant cleanup with deterministic race tests.
|
||||
- [ ] Resolve the indexed predecessor from the allowed active/archive candidates, then run lifecycle, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Archive active `CODE_REVIEW-*-G??.md` to `code_review_cloud_G09_1.log`.
|
||||
- [ ] Archive active `PLAN-*-G??.md` to `plan_cloud_G08_1.log`.
|
||||
- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write `complete.log` based on `agent-ops/skills/common/code-review/templates/complete-log-template.md` and leave no active `.md` files.
|
||||
- [ ] If PASS, move active task directory `agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/` to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/` and update this checklist at the final archive path.
|
||||
- [ ] If PASS, preserve and report `milestone-task=run-lifecycle` for runtime aggregation, without modifying roadmap or directly calling `update-roadmap`.
|
||||
- [ ] If PASS for split work, remove empty active parent only when no sibling remains.
|
||||
- [ ] If WARN/FAIL, write the next filesystem state matching code-review verdict and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record deviations and rationale._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record implementation decisions._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Exactly one child invocation occurs and success requires ordered finish, idle, quiet, and exit 0.
|
||||
- Capture/event artifacts are bounded, sanitized, immutable, and atomically published.
|
||||
- Timeout/cancel races have one terminal outcome; descendants receive escalation and are reaped.
|
||||
- Predecessor index `01` resolves to exactly one active or monthly-archive `complete.log`; an absent or ambiguous match fails before implementation.
|
||||
- No caller-specific adapter, provider protocol, retry, scoring, or reporting behavior entered this packet.
|
||||
|
||||
## Verification Results
|
||||
|
||||
### `python3 -c 'from pathlib import Path; g="m-agent-comparison-benchmark-pipeline"; ids=("01",); a=Path("agent-task")/g; r=Path("agent-task/archive"); found={i:sorted([*a.glob(f"{i}_*/complete.log"),*a.glob(f"{i}+*/complete.log"),*r.glob(f"*/*/{g}/{i}_*/complete.log"),*r.glob(f"*/*/{g}/{i}+*/complete.log")],key=str) for i in ids}; bad={i:[str(p) for p in ps] for i,ps in found.items() if len(ps)!=1}; assert not bad,bad; print("\n".join(str(found[i][0]) for i in ids))'`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 -m unittest scripts.agent_benchmark.lifecycle_test`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `make test-agent-comparison-benchmark`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `git diff --check`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results | Fixed headings/commands | Implementing agent fills actual output only |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,131 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle plan=2 tag=API milestone-task=run-lifecycle -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-09
|
||||
task=m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle, plan=2, tag=API
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/plan_cloud_G08_1.log` and `agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/code_review_cloud_G09_1.log`.
|
||||
- Review state: unimplemented, no official verdict, replaced through explicit plan `write` mode.
|
||||
- Corrected scope: closed idle completion modes and all-terminal descendant cleanup before evidence publication.
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files and verify that output in `Verification Results` matches code.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and verified routing signals.
|
||||
2. Archive this review as `code_review_cloud_G09_2.log` and the plan as `plan_cloud_G08_2.log`.
|
||||
3. If PASS, create `complete.log` and move the task directory to its monthly group archive; otherwise write the required next state.
|
||||
4. If PASS, preserve/report `milestone-task=run-lifecycle`; roadmap evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check the review-only list at the final log location.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|---------|
|
||||
| API-1 Execute one bounded invocation | [ ] |
|
||||
| API-2 Clean every terminal process tree | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement exactly-one invocation, durable process-start locator handoff, normalized event journal, bounded/redacted capture, and strict finish→idle→quiet completion policies.
|
||||
- [ ] Implement single-owner terminal arbitration plus identity-checked recovery and bounded descendant cleanup on success, failure, timeout, cancel, malformed events, and reader errors.
|
||||
- [ ] Resolve predecessor `01`, add deterministic race/orphan/redaction tests, and run focused, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Archive active review to `code_review_cloud_G09_2.log`.
|
||||
- [ ] Archive active plan to `plan_cloud_G08_2.log`.
|
||||
- [ ] Verify `.gitignore` managed rules unignore task markdown/logs and ignore `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write canonical `complete.log` and leave no active `.md` files.
|
||||
- [ ] If PASS, move to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/` and update this checklist there.
|
||||
- [ ] If PASS, preserve/report `milestone-task=run-lifecycle` without directly changing roadmap.
|
||||
- [ ] If PASS for split work, remove empty parent or verify remaining sibling ownership.
|
||||
- [ ] If WARN/FAIL, write the next state and do not create `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record any deviations from the plan and the rationale here._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record key design decisions here._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Exactly one process is launched in a new group with minimal environment; a durable identity/ownership locator is persisted before submission and raw argv/environment is never persisted.
|
||||
- Normalized submission→finish→idle ordering plus quiet window is enforced for both closed completion modes.
|
||||
- Output bounds and exact/fallback redaction apply before atomic journal/result publication.
|
||||
- One terminal arbiter covers success, nonzero, malformed event, reader failure, timeout, cancel, and races.
|
||||
- Every return path gracefully stops/escalates/reaps/drains/verifies descendants before publication; recovery refuses mismatched/reused/unowned process identity before signaling.
|
||||
|
||||
## Verification Results
|
||||
|
||||
### Predecessor completion check from `PLAN-cloud-G08.md`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 -m unittest scripts.agent_benchmark.lifecycle_test`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `make test-agent-comparison-benchmark`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `git diff --check`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementer must not modify or execute these |
|
||||
| Archive Evidence Snapshot | Fixed at stub creation from plan when present | Read only cited archive evidence when required |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementer checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementer checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementer must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholders with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results (section headings + commands) | Fixed at stub creation | Fill output only; changes require a deviation entry |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,159 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle plan=0 tag=API milestone-task=run-lifecycle -->
|
||||
|
||||
# Benchmark Run Lifecycle
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections of `CODE_REVIEW-cloud-G09.md` is the mandatory final implementation step. Run every verification command and paste actual output. Do not finalize, archive, write `complete.log`, ask the user, or change predecessor ownership.
|
||||
|
||||
## Background
|
||||
|
||||
The benchmark needs one bounded lifecycle contract before caller-specific adapters are added. This packet owns generic child invocation, normalized events, completion gates, cancellation, and descendant cleanup; it does not know Claude/Agy/Codex command lines.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `AGENTS.md`
|
||||
- `agent-ops/rules/project/rules.md`
|
||||
- `agent-ops/rules/project/domain/testing/rules.md`
|
||||
- `agent-test/local/rules.md`
|
||||
- `agent-test/local/testing-smoke.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/agent-comparison-benchmark-pipeline.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md`
|
||||
- `agent-spec/input/openai-compatible-surface.md`
|
||||
- `agent-spec/runtime/provider-pool-config-refresh.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
- `scripts/e2e-single-request-claude.sh`
|
||||
- `Makefile`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- Approved/released SDD scenario S04 requires one invocation per attempt, normalized submitted/finish/idle events, success only after finish+idle+quiet confirmation, and bounded timeout/cancel cleanup.
|
||||
- Evidence Map row S04 directly becomes API-1's event/completion assertions, API-2's cancellation/tree-cleanup cases, and the focused plus aggregate Final Verification commands.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- Local Python 3.12.3/Linux arm64; tests use temporary executable fixtures only, no provider credentials or network.
|
||||
- Handoff: no separate verification handoff was supplied; the SDD, local test rules, and repository-native single-request harness are the verification sources.
|
||||
- Existing harness sections for process groups, bounded capture, redaction, and atomic publication are repository-native references, not code to copy blindly.
|
||||
- The existing harness self-test has a baseline cleanup-bound failure outside this packet; the new focused lifecycle suite must establish its own deterministic bounds.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
No shared library currently proves exactly-one invocation, event-order completion, bounded capture, cancellation races, or descendant process cleanup.
|
||||
|
||||
### Symbol References
|
||||
|
||||
Resolve the completed predecessor's manifest timeout/value symbols before editing. The lifecycle public boundary should accept a generic immutable invocation specification and event parser callback rather than import future caller adapters.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
Process state, concurrency, signals, event interpretation, and cleanup form one indivisible lifecycle responsibility boundary, so this slice is large. Predecessor `01_benchmark_manifest` is missing its active or archived `complete.log` at planning time.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
Excluded: caller-specific command construction, provider protocol parsing, workspace copying, retry allocation, scoring, browser checks, and reports.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- evaluation_mode: `first-pass`
|
||||
- finalizer: `finalize-task-policy.sh`, mode `pair`
|
||||
- build closures all `true`; scores `2/2/1/1/2`; route `risk-boundary`; grade `G08`; catalog `worker/cloud/G08`; filename `PLAN-cloud-G08.md`.
|
||||
- review closures all `true`; scores `2/2/1/2/2`; route `official-review`; grade `G09`; catalog `review/cloud/G09`; filename `CODE_REVIEW-cloud-G09.md`.
|
||||
- risks: `temporal_state`, `concurrent_consistency`, `boundary_contract`, `structured_interpretation`, `variant_product`; large_indivisible_context `false`; rework `0`; evidence failure `false`; capability gap none.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement the generic one-process lifecycle, bounded capture, normalized event journal, and strict completion gate.
|
||||
- [ ] Implement cancellation/timeout escalation and descendant cleanup with deterministic race tests.
|
||||
- [ ] Run predecessor, lifecycle, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Execute one bounded invocation
|
||||
|
||||
**Problem:** There is no reusable lifecycle that can distinguish process exit from benchmark completion or retain bounded, sanitized evidence.
|
||||
|
||||
**Solution:** Add immutable invocation/result/event types and a runner using an injected argv/env/cwd plus event parser. Start exactly one child in a new process session, stream stdout/stderr into bounded redacted captures, append normalized JSONL events atomically, and recognize `submitted`, `finish`, and `idle` without depending on provider-specific optional fields. Success requires exit 0, one finish, one subsequent idle, and a bounded quiet window; duplicates/out-of-order/malformed events fail closed.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
No generic benchmark lifecycle or completion gate exists.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
result = run_invocation(spec, parse_event=parse_event, cancellation=token)
|
||||
assert (not result.success) or result.finish_then_idle_then_quiet
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_benchmark/lifecycle.py` with immutable contracts, process supervision, bounded/redacted capture, event validation, and atomic evidence publication.
|
||||
- [ ] Update `scripts/agent_benchmark/__init__.py` to export only the stable generic lifecycle API.
|
||||
|
||||
**Test Strategy:** Use generated temporary executables to cover success, nonzero exit, missing/duplicate/out-of-order events, malformed parser output, capture truncation, redaction, and exact invocation count.
|
||||
|
||||
**Verification:** `python3 -m unittest scripts.agent_benchmark.lifecycle_test` exits 0.
|
||||
|
||||
### [API-2] Bound timeout and cancellation cleanup
|
||||
|
||||
**Problem:** A child may hang or spawn descendants; returning while any process survives contaminates later benchmark cells.
|
||||
|
||||
**Solution:** On deadline or cancellation, mark the terminal reason once, signal the process group, wait a bounded grace interval, escalate to kill, reap the leader, and verify the group no longer exists before returning. Resolve finish/timeout/cancel races under one lock/state transition and never relabel a terminal result as success.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
There is no shared timeout/cancel state machine for benchmark children.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
terminal = supervisor.cancel_or_timeout(reason)
|
||||
assert terminal.reason in {"timeout", "cancelled"}
|
||||
assert not terminal.process_group_alive
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Complete timeout/cancel state transitions and cleanup in `scripts/agent_benchmark/lifecycle.py`.
|
||||
- [ ] Add `scripts/agent_benchmark/lifecycle_test.py` with child-tree, ignored-signal, near-deadline, explicit-cancel, truncation, and no-survivor cases.
|
||||
- [ ] Record exact verification output in `CODE_REVIEW-cloud-G09.md`.
|
||||
|
||||
**Test Strategy:** Poll bounded deadlines and process existence; avoid arbitrary long sleeps and assert cleanup latency against a documented generous bound.
|
||||
|
||||
**Verification:** Focused and aggregate tests pass with no surviving fixture process.
|
||||
|
||||
## Dependencies and Execution Order
|
||||
|
||||
1. Required predecessor: `agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/complete.log`.
|
||||
2. At planning time that completion evidence is absent; do not implement until it exists.
|
||||
3. Resolve final predecessor symbols, implement API-1, then API-2, then run all verification once from a clean process state.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Items |
|
||||
|------|-------|
|
||||
| `scripts/agent_benchmark/lifecycle.py` | API-1, API-2 |
|
||||
| `scripts/agent_benchmark/lifecycle_test.py` | API-2 |
|
||||
| `scripts/agent_benchmark/__init__.py` | API-1 |
|
||||
| `agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/CODE_REVIEW-cloud-G09.md` | API-2 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
1. `test -f agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/complete.log`
|
||||
- Expected: exit 0 before implementation begins.
|
||||
2. `python3 -m unittest scripts.agent_benchmark.lifecycle_test`
|
||||
- Expected: success/failure/event/cancel/timeout/tree-cleanup cases all pass within their bounds.
|
||||
3. `make test-agent-comparison-benchmark`
|
||||
- Expected: all credential-free benchmark tests pass with no fixture process left running.
|
||||
4. `git diff --check`
|
||||
- Expected: exit 0 with no whitespace errors.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,168 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle plan=1 tag=API milestone-task=run-lifecycle -->
|
||||
|
||||
# Benchmark Run Lifecycle
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections of `CODE_REVIEW-cloud-G09.md` is the mandatory final implementation step. Run every verification command and paste actual output. Do not finalize, archive, write `complete.log`, ask the user, or change predecessor ownership.
|
||||
|
||||
## Background
|
||||
|
||||
The benchmark needs one bounded lifecycle contract before caller-specific adapters are added. This packet owns generic child invocation, normalized events, completion gates, cancellation, and descendant cleanup; it does not know Claude/Agy/Codex command lines.
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/plan_cloud_G08_0.log` and `agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/code_review_cloud_G09_0.log`.
|
||||
- Review state: the prior pair was unimplemented and had no official verdict; it was archived by the requested Epic self-review replan.
|
||||
- Self-review defect: its predecessor check named only the active task path, but a PASS predecessor is normally moved under `agent-task/archive/YYYY/MM/`.
|
||||
- Scope carried forward: `run-lifecycle`, SDD S04, implementation files, and credential-free verification remain unchanged; only dependency resolution and paired evidence are corrected.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `AGENTS.md`
|
||||
- `agent-ops/rules/project/rules.md`
|
||||
- `agent-ops/rules/project/domain/testing/rules.md`
|
||||
- `agent-test/local/rules.md`
|
||||
- `agent-test/local/testing-smoke.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/agent-comparison-benchmark-pipeline.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md`
|
||||
- `agent-spec/input/openai-compatible-surface.md`
|
||||
- `agent-spec/runtime/provider-pool-config-refresh.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
- `agent-ops/skills/common/plan/SKILL.md`
|
||||
- `scripts/e2e-single-request-claude.sh`
|
||||
- `Makefile`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- Approved/released SDD scenario S04 requires one invocation per attempt, normalized submitted/finish/idle events, success only after finish+idle+quiet confirmation, and bounded timeout/cancel cleanup.
|
||||
- Evidence Map row S04 directly becomes API-1's event/completion assertions, API-2's cancellation/tree-cleanup cases, and the focused plus aggregate Final Verification commands.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- Local Python 3.12.3/Linux arm64; tests use temporary executable fixtures only, no provider credentials or network.
|
||||
- Handoff: no separate verification handoff was supplied; the SDD, local test rules, and repository-native single-request harness are the verification sources.
|
||||
- Existing harness sections for process groups, bounded capture, redaction, and atomic publication are repository-native references, not code to copy blindly.
|
||||
- The existing harness self-test has a baseline cleanup-bound failure outside this packet; the new focused lifecycle suite must establish its own deterministic bounds.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
No shared library currently proves exactly-one invocation, event-order completion, bounded capture, cancellation races, or descendant process cleanup.
|
||||
|
||||
### Symbol References
|
||||
|
||||
Resolve the completed predecessor's manifest timeout/value symbols before editing. The lifecycle public boundary should accept a generic immutable invocation specification and event parser callback rather than import future caller adapters.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
Process state, concurrency, signals, event interpretation, and cleanup form one indivisible lifecycle responsibility boundary, so this slice is large. Predecessor `01_benchmark_manifest` is missing its active or archived `complete.log` at planning time.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
Excluded: caller-specific command construction, provider protocol parsing, workspace copying, retry allocation, scoring, browser checks, and reports.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- evaluation_mode: `isolated-reassessment`
|
||||
- finalizer: `finalize-task-policy.sh`, mode `pair`
|
||||
- build closures all `true`; scores `2/2/1/1/2`; route `risk-boundary`; grade `G08`; catalog `worker/cloud/G08`; filename `PLAN-cloud-G08.md`.
|
||||
- review closures all `true`; scores `2/2/1/2/2`; route `official-review`; grade `G09`; catalog `review/cloud/G09`; filename `CODE_REVIEW-cloud-G09.md`.
|
||||
- risks: `temporal_state`, `concurrent_consistency`, `boundary_contract`, `structured_interpretation`, `variant_product`; large_indivisible_context `false`; rework `0`; evidence failure `false`; capability gap none.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement the generic one-process lifecycle, bounded capture, normalized event journal, and strict completion gate.
|
||||
- [ ] Implement cancellation/timeout escalation and descendant cleanup with deterministic race tests.
|
||||
- [ ] Resolve the indexed predecessor from the allowed active/archive candidates, then run lifecycle, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Execute one bounded invocation
|
||||
|
||||
**Problem:** There is no reusable lifecycle that can distinguish process exit from benchmark completion or retain bounded, sanitized evidence.
|
||||
|
||||
**Solution:** Add immutable invocation/result/event types and a runner using an injected argv/env/cwd plus event parser. Start exactly one child in a new process session, stream stdout/stderr into bounded redacted captures, append normalized JSONL events atomically, and recognize `submitted`, `finish`, and `idle` without depending on provider-specific optional fields. Success requires exit 0, one finish, one subsequent idle, and a bounded quiet window; duplicates/out-of-order/malformed events fail closed.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
No generic benchmark lifecycle or completion gate exists.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
result = run_invocation(spec, parse_event=parse_event, cancellation=token)
|
||||
assert (not result.success) or result.finish_then_idle_then_quiet
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_benchmark/lifecycle.py` with immutable contracts, process supervision, bounded/redacted capture, event validation, and atomic evidence publication.
|
||||
- [ ] Update `scripts/agent_benchmark/__init__.py` to export only the stable generic lifecycle API.
|
||||
|
||||
**Test Strategy:** Use generated temporary executables to cover success, nonzero exit, missing/duplicate/out-of-order events, malformed parser output, capture truncation, redaction, and exact invocation count.
|
||||
|
||||
**Verification:** `python3 -m unittest scripts.agent_benchmark.lifecycle_test` exits 0.
|
||||
|
||||
### [API-2] Bound timeout and cancellation cleanup
|
||||
|
||||
**Problem:** A child may hang or spawn descendants; returning while any process survives contaminates later benchmark cells.
|
||||
|
||||
**Solution:** On deadline or cancellation, mark the terminal reason once, signal the process group, wait a bounded grace interval, escalate to kill, reap the leader, and verify the group no longer exists before returning. Resolve finish/timeout/cancel races under one lock/state transition and never relabel a terminal result as success.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
There is no shared timeout/cancel state machine for benchmark children.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
terminal = supervisor.cancel_or_timeout(reason)
|
||||
assert terminal.reason in {"timeout", "cancelled"}
|
||||
assert not terminal.process_group_alive
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Complete timeout/cancel state transitions and cleanup in `scripts/agent_benchmark/lifecycle.py`.
|
||||
- [ ] Add `scripts/agent_benchmark/lifecycle_test.py` with child-tree, ignored-signal, near-deadline, explicit-cancel, truncation, and no-survivor cases.
|
||||
- [ ] Record exact verification output in `CODE_REVIEW-cloud-G09.md`.
|
||||
|
||||
**Test Strategy:** Poll bounded deadlines and process existence; avoid arbitrary long sleeps and assert cleanup latency against a documented generous bound.
|
||||
|
||||
**Verification:** Focused and aggregate tests pass with no surviving fixture process.
|
||||
|
||||
## Dependencies and Execution Order
|
||||
|
||||
1. Required predecessor index: `01_benchmark_manifest`, encoded by the `03+01_...` task directory.
|
||||
2. At planning time no allowed active/archive `01_*/complete.log` or `01+*/complete.log` candidate exists. Runtime scheduling must wait; a completed predecessor may be under the active task group or `agent-task/archive/YYYY/MM/`.
|
||||
3. Before implementation, require exactly one matching predecessor completion record across those allowed locations. Multiple matches are ambiguous and must not be guessed.
|
||||
4. Resolve final predecessor symbols, implement API-1, then API-2, then run all verification once from a clean process state.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Items |
|
||||
|------|-------|
|
||||
| `scripts/agent_benchmark/lifecycle.py` | API-1, API-2 |
|
||||
| `scripts/agent_benchmark/lifecycle_test.py` | API-2 |
|
||||
| `scripts/agent_benchmark/__init__.py` | API-1 |
|
||||
| `agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/CODE_REVIEW-cloud-G09.md` | API-2 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
1. `python3 -c 'from pathlib import Path; g="m-agent-comparison-benchmark-pipeline"; ids=("01",); a=Path("agent-task")/g; r=Path("agent-task/archive"); found={i:sorted([*a.glob(f"{i}_*/complete.log"),*a.glob(f"{i}+*/complete.log"),*r.glob(f"*/*/{g}/{i}_*/complete.log"),*r.glob(f"*/*/{g}/{i}+*/complete.log")],key=str) for i in ids}; bad={i:[str(p) for p in ps] for i,ps in found.items() if len(ps)!=1}; assert not bad,bad; print("\n".join(str(found[i][0]) for i in ids))'`
|
||||
- Expected: exit 0 and print exactly one active or archived completion path for predecessor index `01` before implementation begins.
|
||||
2. `python3 -m unittest scripts.agent_benchmark.lifecycle_test`
|
||||
- Expected: success/failure/event/cancel/timeout/tree-cleanup cases all pass within their bounds.
|
||||
3. `make test-agent-comparison-benchmark`
|
||||
- Expected: all credential-free benchmark tests pass with no fixture process left running.
|
||||
4. `git diff --check`
|
||||
- Expected: exit 0 with no whitespace errors.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,173 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle plan=2 tag=API milestone-task=run-lifecycle -->
|
||||
|
||||
# Benchmark Run Lifecycle
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections of `CODE_REVIEW-cloud-G09.md` is mandatory. Resolve predecessor `01`, run every verification command, paste actual output, and leave the active pair for official review. Do not finalize, archive, write `complete.log`, ask the user, or change predecessor ownership.
|
||||
|
||||
## Background
|
||||
|
||||
Caller adapters need one generic, bounded subprocess lifecycle. This packet owns exactly-one invocation, normalized finish/idle evidence, output quiescence, redaction, completion policy, timeout/cancel, and descendant cleanup. It does not encode Claude, agy, or Codex command lines.
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/plan_cloud_G08_1.log` and `agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/code_review_cloud_G09_1.log` (generation 1 retains generation 0 history).
|
||||
- Review state: unimplemented, no official verdict, replaced through explicit plan `write` mode.
|
||||
- Self-review defects: the prior success rule required natural exit after idle, which cannot represent a caller that remains idle until the harness stops it, and descendant cleanup was explicit only for timeout/cancel rather than every terminal path.
|
||||
- Scope carried forward: generic lifecycle/events, bounded capture, timeout/cancel races, predecessor resolution, and credential-free subprocess tests.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `AGENTS.md`
|
||||
- `agent-ops/rules/project/rules.md`
|
||||
- `agent-ops/rules/project/domain/testing/rules.md`
|
||||
- `agent-ops/rules/common/rules-roadmap.md`
|
||||
- `agent-ops/skills/common/plan/SKILL.md`
|
||||
- `agent-ops/skills/common/finalize-task-routing/SKILL.md`
|
||||
- `agent-test/local/rules.md`
|
||||
- `agent-test/local/testing-smoke.md`
|
||||
- `agent-roadmap/current.md`
|
||||
- `agent-roadmap/priority-queue.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/PHASE.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/agent-comparison-benchmark-pipeline.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md`
|
||||
- `agent-spec/index.md`
|
||||
- `agent-spec/input/openai-compatible-surface.md`
|
||||
- `agent-contract/index.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `Makefile`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- S04 requires one user submission, caller-specific finish/complete normalization, idle, bounded waiting, and a common terminal timeline.
|
||||
- State invariant line 72 requires idle plus process/output quiescence; Evidence Map S04 requires fixture streams and a real subprocess lifecycle probe.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- Local Python 3.12.3/Linux arm64. Tests use real temporary subprocesses/process groups and deterministic event fixtures; no external provider/network/credential.
|
||||
- The `bench-01` lane and SDD gates are open. Predecessor `01` is not complete, so runtime scheduling must wait.
|
||||
- The lifecycle must be POSIX-safe for the current Linux/macOS benchmark target. Unsupported platforms fail before launch rather than weakening cleanup.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
No current shared module proves ordered finish→idle→quiet, both exit and harness-stop completion modes, bounded/redacted capture, or no surviving descendants on success, error, malformed event, timeout, and cancel.
|
||||
|
||||
### Symbol References
|
||||
|
||||
Resolve the completed manifest's timeout types before implementation. The lifecycle accepts an immutable invocation specification plus parser/redactor hooks and must not import future caller adapters.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
Process state, event ordering, output readers, terminal arbitration, signals, and cleanup share one race-sensitive invariant. Splitting would allow success publication before cleanup, so the packet stays atomic.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
Excluded: caller command construction/protocol parsing, workspace creation, attempt persistence/recovery, provider usage normalization, scoring, browser checks, and reports.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- evaluation_mode: `isolated-reassessment`
|
||||
- finalizer: `finalize-task-policy.sh`, mode `pair`
|
||||
- build closures all `true`; scores `2/2/1/1/2`; base `local-fit`, route `risk-boundary`; grade `G08`; catalog `worker/cloud/G08`; filename `PLAN-cloud-G08.md`.
|
||||
- review closures all `true`; scores `2/2/1/2/2`; route `official-review`; grade `G09`; catalog `review/cloud/G09`; filename `CODE_REVIEW-cloud-G09.md`.
|
||||
- large_indivisible_context `false`; risks `temporal_state`, `concurrent_consistency`, `boundary_contract`, `structured_interpretation`, `variant_product`; rework `0`; evidence integrity failure `false`; capability gap none.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement exactly-one invocation, durable process-start locator handoff, normalized event journal, bounded/redacted capture, and strict finish→idle→quiet completion policies.
|
||||
- [ ] Implement single-owner terminal arbitration plus identity-checked recovery and bounded descendant cleanup on success, failure, timeout, cancel, malformed events, and reader errors.
|
||||
- [ ] Resolve predecessor `01`, add deterministic race/orphan/redaction tests, and run focused, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Execute one bounded invocation
|
||||
|
||||
**Problem:** `SDD.md:72`, `SDD.md:85`, and S04 require finish/complete, idle, and quiescence, but no generic runner exists. Natural process exit cannot be the only success policy because some callers remain alive while idle.
|
||||
|
||||
**Solution:** Add frozen invocation/result/event/process-locator types. Start exactly one argv in a new POSIX session with explicit cwd and a minimal allowlisted environment. Never persist raw argv/environment. After spawn and before task submission or output observation, invoke a required durable `on_started(locator)` callback carrying pid, process-group id, platform start identity, and an unguessable ownership marker; callback failure enters cleanup and no task is submitted. Inject a caller parser that converts native output to the closed normalized sequence `submitted`, `finish`, `idle`; optional later metric events remain data and cannot synthesize completion. Capture stdout/stderr with hard byte/line bounds, exact secret-value redaction supplied by the adapter, secret-shaped fallback redaction, truncation markers, and monotonic observation times. Publish a JSONL journal and terminal result through temp-file + fsync + replace only after readers stop and cleanup completes.
|
||||
|
||||
Support two closed completion modes in the invocation spec: `exit_after_idle` requires exit 0 after ordered finish→idle→quiet; `stop_after_idle` treats ordered finish→idle→quiet as task completion, then performs a separate expected graceful harness stop. In both modes, duplicate/out-of-order/malformed events, nonzero unexpected exit, or missing idle fail closed.
|
||||
|
||||
Before (`SDD.md:103`):
|
||||
|
||||
```text
|
||||
one submission -> finish/complete -> idle -> common terminal outcome
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
result = run_invocation(spec, parse_event=parse_event, redact=redactor, cancellation=token)
|
||||
assert (not result.success) or result.finish_then_idle_then_quiet
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_benchmark/lifecycle.py` with immutable contracts, durable start callback, readers, event validation, redaction, completion modes, terminal ownership, recovery cleanup, and atomic evidence publication.
|
||||
- [ ] Update `scripts/agent_benchmark/__init__.py` to export only stable lifecycle APIs.
|
||||
|
||||
**Test Strategy:** Use real temporary executable scripts. Cover both completion modes, exact invocation count, durable start callback ordering/failure, exit/nonzero, missing/duplicate/out-of-order events, malformed parser output, reader error, capture truncation, exact/fallback redaction, quiet-window boundaries, and journal/result immutability.
|
||||
|
||||
**Verification:** `python3 -m unittest scripts.agent_benchmark.lifecycle_test` exits 0.
|
||||
|
||||
### [API-2] Clean every terminal process tree
|
||||
|
||||
**Problem:** `SDD.md:65` requires cleanup evidence. A leader or descendant surviving any return path can contaminate later attempts; timeout/cancel-only cleanup is insufficient.
|
||||
|
||||
**Solution:** Route success, nonzero exit, event failure, output-reader failure, timeout, and explicit cancel through one terminal arbiter and one cleanup `finally` path. If the group is still live, send the mode-appropriate graceful signal, wait a bounded grace interval, escalate to group kill, reap the leader, drain/join readers, and verify no member survives before result publication. Expose the same bounded cleanup as `recover_invocation(locator)`: it first proves pid/group/start identity and ownership marker still match, refuses an ambiguous/reused/unowned process, then terminates and verifies the recorded group. Near-deadline finish/timeout/cancel races select exactly one terminal reason; cleanup failure is a terminal failure and cannot be relabeled success.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
No shared all-terminal cleanup owner exists.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
terminal = supervisor.finish(reason)
|
||||
assert terminal.cleanup_complete
|
||||
assert not terminal.process_group_alive
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Complete all-path terminal/cleanup state and identity-checked `recover_invocation` in `scripts/agent_benchmark/lifecycle.py`.
|
||||
- [ ] Add `scripts/agent_benchmark/lifecycle_test.py` with child-tree cases for natural success, nonzero exit, malformed event, ignored graceful stop, near-deadline timeout, explicit cancel, and reader failure.
|
||||
- [ ] Record exact output in `CODE_REVIEW-cloud-G09.md`.
|
||||
|
||||
**Test Strategy:** Poll process identities with bounded deadlines; do not use unbounded or arbitrary long sleeps. Assert no fixture leader/descendant survives each case, matching recovery succeeds, pid/start/marker mismatch refuses to signal, and cleanup latency stays below a documented generous bound.
|
||||
|
||||
**Verification:** Focused and aggregate tests pass with no surviving fixture process.
|
||||
|
||||
## Dependencies and Execution Order
|
||||
|
||||
1. Required predecessor `01_benchmark_manifest` is encoded by `03+01_...`.
|
||||
2. No valid active/archive predecessor completion exists at planning time; runtime waits.
|
||||
3. Before implementation, require exactly one matching allowed completion record; missing/multiple matches fail closed.
|
||||
4. Resolve timeout/value symbols, implement API-1 and API-2 as one terminal invariant, then verify from a clean process state.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Items |
|
||||
|------|-------|
|
||||
| `scripts/agent_benchmark/lifecycle.py` | API-1, API-2 |
|
||||
| `scripts/agent_benchmark/lifecycle_test.py` | API-2 |
|
||||
| `scripts/agent_benchmark/__init__.py` | API-1 |
|
||||
| `agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/CODE_REVIEW-cloud-G09.md` | API-2 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
1. `python3 -c 'from pathlib import Path; g="m-agent-comparison-benchmark-pipeline"; ids=("01",); a=Path("agent-task")/g; r=Path("agent-task/archive"); found={i:sorted([*a.glob(f"{i}_*/complete.log"),*a.glob(f"{i}+*/complete.log"),*r.glob(f"*/*/{g}/{i}_*/complete.log"),*r.glob(f"*/*/{g}/{i}+*/complete.log")],key=str) for i in ids}; bad={i:[str(p) for p in ps] for i,ps in found.items() if len(ps)!=1}; assert not bad,bad; print("\n".join(str(found[i][0]) for i in ids))'`
|
||||
- Expected: exactly one completion path for predecessor `01` before implementation.
|
||||
2. `python3 -m unittest scripts.agent_benchmark.lifecycle_test`
|
||||
- Expected: event, completion-mode, redaction, race, timeout/cancel, and all-terminal cleanup cases pass.
|
||||
3. `make test-agent-comparison-benchmark`
|
||||
- Expected: all credential-free benchmark tests pass with no fixture process left alive.
|
||||
4. `git diff --check`
|
||||
- Expected: exit 0 with no whitespace errors.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,132 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt plan=3 tag=API milestone-task=repeat-attempt -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record exact evidence and the resume condition only in implementation-owned fields.
|
||||
> Do not ask the user, call user-input tools, create stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only.
|
||||
> Follow the ownership table at the bottom of this file.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-09
|
||||
task=m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt, plan=3, tag=API
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/plan_cloud_G07_2.log` and `agent-task/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/code_review_cloud_G08_2.log` (generation 2 retains earlier history).
|
||||
- Review state: unimplemented, no official verdict, replaced through explicit plan `write` mode.
|
||||
- Self-review defects: the prior pair used exclusive attempt allocation without a run-level single-writer lease, so concurrent `run`/`resume` processes could start two attempts for one slot. It also marked every recovered nonterminal attempt `interrupted`, losing a valid lifecycle terminal that was published just before controller failure.
|
||||
- Scope carried forward: `repeat-attempt`, S05, append-only attempts, run/resume/status/retry, predecessor resolution, and credential-free tests.
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** Implementing agents must not execute this section.
|
||||
|
||||
Compare every item to source and verify pasted command output.
|
||||
|
||||
1. Append verdict and verified routing signals.
|
||||
2. Archive this review to `code_review_cloud_G08_3.log` and the plan to `plan_cloud_G07_3.log`.
|
||||
3. On PASS, create `complete.log` and move the task directory to its monthly group archive; otherwise write the required next state.
|
||||
4. On PASS, preserve/report `milestone-task=repeat-attempt`; roadmap evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check review-only items at the final log location.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|---------|
|
||||
| API-1 Bind runs and attempts to immutable identity | [ ] |
|
||||
| API-2 Run, recover, resume, and retry safely | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Persist an immutable canonical manifest snapshot/digest, deterministic repetition slots, and append-only attempt records under one fail-fast run-level writer lease.
|
||||
- [ ] Integrate prepared workspaces and lifecycle outcomes into run/resume/status plus explicit failed-attempt retry without evidence overwrite.
|
||||
- [ ] Recover nonterminal attempts by reconciling a valid terminal first or authenticating and cleaning the registered supervisor before sealing interruption and allocating a successor; add crash/concurrency/retry tests.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one `PASS`, `WARN`, or `FAIL` verdict plus verified `review_rework_count` and `evidence_integrity_failure`.
|
||||
- [ ] Verify verdict, dimensions, and Required/Suggested/Nit classifications agree.
|
||||
- [ ] Archive active review to `code_review_cloud_G08_3.log`.
|
||||
- [ ] Archive active plan to `plan_cloud_G07_3.log`.
|
||||
- [ ] Verify `.gitignore` managed task/roadmap rules.
|
||||
- [ ] If PASS, write canonical `complete.log` and leave no active `.md` files.
|
||||
- [ ] If PASS, move to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/` and update this checklist there.
|
||||
- [ ] If PASS, preserve/report `milestone-task=repeat-attempt` without directly changing roadmap.
|
||||
- [ ] If PASS for split work, remove empty parent or justify remaining siblings.
|
||||
- [ ] If WARN/FAIL, materialize the required next state and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record any deviations from the plan and the rationale here._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record key design decisions here._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Run creation persists and verifies byte-immutable canonical manifest identity before any slot work.
|
||||
- Run/attempt ids are harness-generated closed basenames; resume/status reject traversal, symlink, and foreign-root state.
|
||||
- `run`/`resume` hold one fail-fast kernel-released writer lease; concurrent writers cause zero mutation, every slot has at most one nonterminal attempt, and terminal records remain byte-identical.
|
||||
- Each allocated attempt calls predecessor workspace and lifecycle APIs exactly once in the documented order.
|
||||
- Resume first commits a matching atomic lifecycle terminal and cleanup receipt; only a genuinely unterminated attempt becomes `interrupted`.
|
||||
- A registered survivor is challenge-authenticated and cleaned before interruption/successor allocation; unverifiable absence or ownership fails closed without mutation.
|
||||
- Missing adapters cause zero run/attempt/workspace/supervisor side effects; status is read-only and errors remain deterministic/redacted.
|
||||
|
||||
## Verification Results
|
||||
|
||||
### Predecessor completion check from `PLAN-cloud-G07.md`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 -m unittest scripts.agent_benchmark.attempts_test`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `make test-agent-comparison-benchmark`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `git diff --check`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementer must not modify or execute these |
|
||||
| Archive Evidence Snapshot | Fixed at stub creation from plan when present | Read only cited evidence when required |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementer checks status only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementer checks status only |
|
||||
| Review-Only Checklist | Review agent only | Implementer must not modify |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholders with evidence |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results (section headings + commands) | Fixed at stub creation | Fill output only; changes require deviation |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,182 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt plan=3 tag=API milestone-task=repeat-attempt -->
|
||||
|
||||
# Repeat and Attempt State
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections of `CODE_REVIEW-cloud-G08.md` is mandatory. Resolve predecessors `01`, `02`, and `03` before editing, run every verification command, paste actual output, and leave the active pair for official review. Do not finalize, archive, write `complete.log`, ask the user, or change predecessor ownership.
|
||||
|
||||
## Background
|
||||
|
||||
The validated manifest, isolated workspace, and bounded lifecycle need one durable orchestration boundary. This packet owns immutable run identity, append-only attempt allocation, crash-safe resume, explicit retry, and sanitized status. Real caller adapters remain a later Epic.
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/plan_cloud_G07_2.log` and `agent-task/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/code_review_cloud_G08_2.log` (generation 2 retains earlier history).
|
||||
- Review state: unimplemented, no official verdict, replaced through explicit plan `write` mode.
|
||||
- Self-review defects: the prior pair used exclusive attempt allocation without a run-level single-writer lease, so concurrent `run`/`resume` processes could start two attempts for one slot. It also marked every recovered nonterminal attempt `interrupted`, losing a valid lifecycle terminal that was published just before controller failure.
|
||||
- Scope carried forward: `repeat-attempt`, S05, append-only attempts, run/resume/status/retry, predecessor resolution, and credential-free tests.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `AGENTS.md`
|
||||
- `agent-ops/rules/project/rules.md`
|
||||
- `agent-ops/rules/project/domain/testing/rules.md`
|
||||
- `agent-ops/rules/common/rules-roadmap.md`
|
||||
- `agent-ops/skills/common/plan/SKILL.md`
|
||||
- `agent-ops/skills/common/code-review/SKILL.md`
|
||||
- `agent-ops/skills/common/finalize-task-routing/SKILL.md`
|
||||
- `agent-ops/skills/common/plan/templates/review-stub-template.md`
|
||||
- `agent-test/local/rules.md`
|
||||
- `agent-test/local/testing-smoke.md`
|
||||
- `agent-roadmap/current.md`
|
||||
- `agent-roadmap/priority-queue.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/PHASE.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/agent-comparison-benchmark-pipeline.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md`
|
||||
- `agent-spec/index.md`
|
||||
- `agent-spec/input/openai-compatible-surface.md`
|
||||
- `agent-spec/runtime/provider-pool-config-refresh.md`
|
||||
- `agent-contract/index.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
- `Makefile`
|
||||
- `.gitignore`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- S05 requires default repetition `1`, independent append-only attempts, preserved failures, and resume without evidence overwrite.
|
||||
- S01 supplies canonical matrix identity; S03 supplies clean workspace/session identity; S04 supplies the bounded lifecycle outcome consumed by each attempt.
|
||||
- Evidence Map S05 requires deterministic repetition, retry, concurrency, and resume evidence.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- Local Python 3.12.3/Linux arm64; standard-library tests use temporary directories and injected adapters only.
|
||||
- `bench-01` is unblocked, but predecessors `01`, `02`, and `03` have no completion records, so implementation must wait.
|
||||
- No provider, network, browser, or sibling-checkout mutation is permitted. Missing adapters fail before workspace or process creation.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
No current code persists an immutable run snapshot, serializes all run mutation under one crash-released writer lease, owns a repetition slot, atomically allocates attempts, integrates workspace/lifecycle, reconciles already-published lifecycle terminals, cleans a surviving supervisor/caller group, or preserves old evidence through resume/retry.
|
||||
|
||||
### Symbol References
|
||||
|
||||
Resolve completed manifest canonicalization, workspace preparation, and lifecycle invocation/result/process-locator types before editing. Do not duplicate those contracts.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
Manifest identity, slot ownership, process recovery, transition atomicity, and retry policy form one crash-consistency boundary. Splitting them would permit a new attempt while an old process or mutable identity remains active.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
Excluded: real caller command adapters, provider translation, scoring, browser checks, aggregate reporting, and adapter-specific retry policy.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- evaluation_mode: `isolated-reassessment`
|
||||
- finalizer: `finalize-task-policy.sh`, mode `pair`
|
||||
- build closures all `true`; scores `2/2/1/1/1`; base `local-fit`, route `risk-boundary`; grade `G07`; catalog `worker/cloud/G07`; filename `PLAN-cloud-G07.md`.
|
||||
- review closures all `true`; scores `2/2/1/2/1`; route `official-review`; grade `G08`; catalog `review/cloud/G08`; filename `CODE_REVIEW-cloud-G08.md`.
|
||||
- large_indivisible_context `false`; risks `temporal_state`, `concurrent_consistency`, `boundary_contract`, `variant_product`; rework `0`; evidence integrity failure `false`; capability gap none.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Persist an immutable canonical manifest snapshot/digest, deterministic repetition slots, and append-only attempt records under one fail-fast run-level writer lease.
|
||||
- [ ] Integrate prepared workspaces and lifecycle outcomes into run/resume/status plus explicit failed-attempt retry without evidence overwrite.
|
||||
- [ ] Recover nonterminal attempts by reconciling a valid terminal first or authenticating and cleaning the registered supervisor before sealing interruption and allocating a successor; add crash/concurrency/retry tests.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Bind runs and attempts to immutable identity
|
||||
|
||||
**Problem:** A run can be resumed against changed manifest bytes or allocate overlapping work for the same cell/repetition, making evidence incomparable or overwritable.
|
||||
|
||||
**Solution:** On new run creation, consume the predecessor loader's canonical manifest form/digest, generate a harness-owned `run-YYYYMMDDTHHMMSSZ-<12 lowercase hex>` id through an injectable clock/random source, exclusively create its non-symlink directory under the validated `output_root`, and atomically persist the byte-immutable snapshot plus digest. `resume`/`status` accept only that exact basename grammar, resolve it under the same output root, and reject absolute, traversal, symlink, missing, or foreign roots. Attempt directories use monotonic `attempt-000001` basenames and exclusive non-symlink creation; user text never becomes a path component. Reject an existing generated run id and reject resume when supplied manifest bytes/digest or canonical expansion differ. Expand cells in canonical id order and repetitions `1..N`; define each `(manifest_digest, run_id, cell_id, repetition)` as one slot. Before any mutation, `run` and `resume` acquire a non-blocking POSIX `flock` on a stable run lock file and hold it through allocation, lifecycle reconciliation, and terminal publication; a concurrent writer returns sanitized `run-busy` with zero mutation. Kernel lock release on process death is the stale-lock policy. `status` opens immutable/atomically replaced records read-only and never claims the writer lease.
|
||||
|
||||
Under the lease, allow at most one nonterminal attempt per slot, allocate monotonically increasing attempt ids with exclusive directory creation, and atomically publish transitions. Store each attempt's workspace metadata, lifecycle journal/result, supervisor locator, and cleanup receipt only beneath that attempt directory. Terminal attempt records never change.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
No durable run snapshot, slot identity, or attempt allocation exists.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
run = store.create(manifest_snapshot, manifest_digest)
|
||||
attempt = store.allocate(run.slot(cell_id, repetition))
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_benchmark/attempts.py` with frozen run/slot/attempt types, snapshot verification, canonical expansion, exclusive allocation, and atomic transitions.
|
||||
- [ ] Update `scripts/agent_benchmark/__init__.py` to export only stable attempt APIs.
|
||||
|
||||
**Test Strategy:** Cover deterministic injected run-id generation/collision, run/resume/status basename containment and symlink rejection, repetition default/override, canonical order, resume digest/byte mismatch, corrupt/foreign state, concurrent run/resume writer contention with exactly one mutation owner, one-nonterminal-per-slot, monotonic attempt ids, crash-released lock recovery, read-only status during a writer, and terminal byte immutability.
|
||||
|
||||
**Verification:** `python3 -m unittest scripts.agent_benchmark.attempts_test` exits 0.
|
||||
|
||||
### [API-2] Run, recover, resume, and retry safely
|
||||
|
||||
**Problem:** A crashed controller can leave both a nonterminal record and a live process tree. Sealing it interrupted and starting again before cleanup creates concurrent scored work and contaminates evidence.
|
||||
|
||||
**Solution:** Add `run`, `resume`, and `status`; permit `--retry-failed` only on resume. Before creating a run id/directory, resolve every manifest caller against the injected adapter registry; while this Epic has no real adapters, return `capability-unavailable: caller-adapter` with no run/attempt/workspace/supervisor side effect. For each pending slot under the writer lease after adapters exist, allocate an attempt, call the completed workspace preparation API, and invoke the completed lifecycle API with an `on_started` callback that atomically persists its registered supervisor locator before the lifecycle controller can send `START`. Then atomically publish the lifecycle terminal outcome and its cleanup receipt.
|
||||
|
||||
On resume, inspect each nonterminal attempt while holding the writer lease. First validate any atomic lifecycle terminal result, cleanup receipt, attempt identity, and digests; if complete and consistent, publish that exact success/failure/timeout/cancel outcome instead of inventing `interrupted` or allocating a successor. Otherwise, a missing locator is safe to seal `interrupted` because the predecessor supervisor contract forbids caller launch before locator commit. For a present locator, authenticate the supervisor control challenge, request cleanup, and verify the receipt/group absence before sealing `interrupted`. A stale/forged/ambiguous locator or unreachable supervisor whose absence cannot be proven blocks resume without state mutation. Only after reconciliation may resume allocate a successor. Skip successes; preserve failed/timeout/cancelled records unless explicit retry is requested. Status is deterministic and redacted.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
No supported crash-safe continuation or explicit retry exists.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```text
|
||||
run|resume|status preserve the manifest snapshot and old attempt bytes.
|
||||
resume cleans an identity-matched survivor before sealing and reallocating.
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Complete writer-leased workspace/lifecycle orchestration, terminal reconciliation, and authenticated recovery policy in `scripts/agent_benchmark/attempts.py`.
|
||||
- [ ] Update `scripts/agent_comparison_benchmark.py` with sanitized `run`, `resume`, and `status` contracts.
|
||||
- [ ] Add `scripts/agent_benchmark/attempts_test.py` with integration, crash-survivor, ambiguous-identity, concurrency, skip, retry, redaction, and unavailable-adapter cases.
|
||||
- [ ] Record exact output in `CODE_REVIEW-cloud-G08.md`.
|
||||
|
||||
**Test Strategy:** Inject deterministic adapters and real short-lived supervisor/process fixtures. Assert exactly one writer and one active attempt per slot, workspace/lifecycle call counts, terminal-result-before-interruption reconciliation, cleanup-before-interruption ordering, no surviving owned process group, old-record byte identity, and no run directory or downstream side effect on unavailable capability.
|
||||
|
||||
**Verification:** Focused and aggregate tests pass with no fixture process or temporary state left alive.
|
||||
|
||||
## Dependencies and Execution Order
|
||||
|
||||
1. Required predecessors `01`, `02`, and `03` are encoded by `04+01,02,03_...`.
|
||||
2. At planning time none has exactly one valid active/archive `complete.log`; runtime must wait.
|
||||
3. Before implementation, require exactly one allowed completion for every predecessor. Missing or multiple matches fail closed.
|
||||
4. Resolve final predecessor symbols, implement immutable identity, integrate workspace/lifecycle, add recovery/retry, then verify.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Items |
|
||||
|------|-------|
|
||||
| `scripts/agent_benchmark/attempts.py` | API-1, API-2 |
|
||||
| `scripts/agent_benchmark/attempts_test.py` | API-1, API-2 |
|
||||
| `scripts/agent_benchmark/__init__.py` | API-1 |
|
||||
| `scripts/agent_comparison_benchmark.py` | API-2 |
|
||||
| `agent-task/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/CODE_REVIEW-cloud-G08.md` | API-2 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
1. `python3 -c 'from pathlib import Path; g="m-agent-comparison-benchmark-pipeline"; ids=("01","02","03"); a=Path("agent-task")/g; r=Path("agent-task/archive"); found={i:sorted([*a.glob(f"{i}_*/complete.log"),*a.glob(f"{i}+*/complete.log"),*r.glob(f"*/*/{g}/{i}_*/complete.log"),*r.glob(f"*/*/{g}/{i}+*/complete.log")],key=str) for i in ids}; bad={i:[str(p) for p in ps] for i,ps in found.items() if len(ps)!=1}; assert not bad,bad; print("\n".join(str(found[i][0]) for i in ids))'`
|
||||
- Expected: exactly one allowed completion path for each predecessor before implementation.
|
||||
2. `python3 -m unittest scripts.agent_benchmark.attempts_test`
|
||||
- Expected: immutable identity, integration, crash recovery, concurrency, retry, redaction, and unavailable-adapter cases pass.
|
||||
3. `make test-agent-comparison-benchmark`
|
||||
- Expected: all credential-free benchmark tests pass with no process left alive.
|
||||
4. `git diff --check`
|
||||
- Expected: exit 0 with no whitespace errors.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,123 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt plan=0 tag=API milestone-task=repeat-attempt -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-09
|
||||
task=m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt, plan=0, tag=API
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files and verify that output in `Verification Results` matches code.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals.
|
||||
2. Archive `CODE_REVIEW-cloud-G08.md` → `code_review_cloud_G08_0.log` and `PLAN-cloud-G07.md` → `plan_cloud_G07_0.log`.
|
||||
3. If PASS, write `complete.log` and move active task directory to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/`. If WARN/FAIL, fully write the next filesystem state required by the code-review skill.
|
||||
4. If PASS, preserve `milestone-task=repeat-attempt` in `complete.log` and report it for runtime aggregation; roadmap evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check applicable `Review-Only Checklist` items at the final `.log` location before reporting.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|--------|
|
||||
| API-1 Persist canonical attempts without overwrite | [ ] |
|
||||
| API-2 Expose run, resume, status, and explicit retry | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement canonical run expansion and append-only, atomically allocated attempt records.
|
||||
- [ ] Implement run/resume/status and explicit failed-attempt retry without evidence overwrite.
|
||||
- [ ] Add crash, concurrency, retry, ordering, and unavailable-adapter tests and run all verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Archive active `CODE_REVIEW-*-G??.md` to `code_review_cloud_G08_0.log`.
|
||||
- [ ] Archive active `PLAN-*-G??.md` to `plan_cloud_G07_0.log`.
|
||||
- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write `complete.log` based on `agent-ops/skills/common/code-review/templates/complete-log-template.md` and leave no active `.md` files.
|
||||
- [ ] If PASS, move active task directory `agent-task/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/` to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/` and update this checklist at the final archive path.
|
||||
- [ ] If PASS, preserve and report `milestone-task=repeat-attempt` for runtime aggregation, without modifying roadmap or directly calling `update-roadmap`.
|
||||
- [ ] If PASS for split work, remove empty active parent only when no sibling remains.
|
||||
- [ ] If WARN/FAIL, write the next filesystem state matching code-review verdict and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record deviations and rationale._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record implementation decisions._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Canonical cell/repetition identity and attempt allocation are deterministic and collision-safe.
|
||||
- Terminal evidence is byte-preserved; recovery and explicit retry always allocate a new attempt.
|
||||
- Resume skip/retry rules and status output match the documented CLI contract.
|
||||
- Missing caller adapters fail before workspace/process work; no fake success or downstream report work exists.
|
||||
|
||||
## Verification Results
|
||||
|
||||
### `test -f agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/complete.log && test -f agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/complete.log && test -f agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/complete.log`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 -m unittest scripts.agent_benchmark.attempts_test`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `make test-agent-comparison-benchmark`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `git diff --check`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results | Fixed headings/commands | Implementing agent fills actual output only |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,131 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt plan=1 tag=API milestone-task=repeat-attempt -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-09
|
||||
task=m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt, plan=1, tag=API
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/plan_cloud_G07_0.log` and `agent-task/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/code_review_cloud_G08_0.log`.
|
||||
- Review state: the prior pair was unimplemented and had no official verdict; it was archived by the requested Epic self-review replan.
|
||||
- Self-review defect: its predecessor checks named only active task paths, but PASS predecessors are normally moved under `agent-task/archive/YYYY/MM/`.
|
||||
- Scope carried forward: `repeat-attempt`, SDD S05, implementation files, and credential-free verification remain unchanged; only dependency resolution and paired evidence are corrected.
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files and verify that output in `Verification Results` matches code.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals.
|
||||
2. Archive `CODE_REVIEW-cloud-G08.md` → `code_review_cloud_G08_1.log` and `PLAN-cloud-G07.md` → `plan_cloud_G07_1.log`.
|
||||
3. If PASS, write `complete.log` and move active task directory to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/`. If WARN/FAIL, fully write the next filesystem state required by the code-review skill.
|
||||
4. If PASS, preserve `milestone-task=repeat-attempt` in `complete.log` and report it for runtime aggregation; roadmap evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check applicable `Review-Only Checklist` items at the final `.log` location before reporting.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|--------|
|
||||
| API-1 Persist canonical attempts without overwrite | [ ] |
|
||||
| API-2 Expose run, resume, status, and explicit retry | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement canonical run expansion and append-only, atomically allocated attempt records.
|
||||
- [ ] Implement run/resume/status and explicit failed-attempt retry without evidence overwrite.
|
||||
- [ ] Resolve all indexed predecessors from the allowed active/archive candidates, add crash, concurrency, retry, ordering, and unavailable-adapter tests, and run all verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Archive active `CODE_REVIEW-*-G??.md` to `code_review_cloud_G08_1.log`.
|
||||
- [ ] Archive active `PLAN-*-G??.md` to `plan_cloud_G07_1.log`.
|
||||
- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write `complete.log` based on `agent-ops/skills/common/code-review/templates/complete-log-template.md` and leave no active `.md` files.
|
||||
- [ ] If PASS, move active task directory `agent-task/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/` to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/` and update this checklist at the final archive path.
|
||||
- [ ] If PASS, preserve and report `milestone-task=repeat-attempt` for runtime aggregation, without modifying roadmap or directly calling `update-roadmap`.
|
||||
- [ ] If PASS for split work, remove empty active parent only when no sibling remains.
|
||||
- [ ] If WARN/FAIL, write the next filesystem state matching code-review verdict and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record deviations and rationale._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record implementation decisions._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Canonical cell/repetition identity and attempt allocation are deterministic and collision-safe.
|
||||
- Terminal evidence is byte-preserved; recovery and explicit retry always allocate a new attempt.
|
||||
- Resume skip/retry rules and status output match the documented CLI contract.
|
||||
- Predecessor indices `01`, `02`, and `03` each resolve to exactly one active or monthly-archive `complete.log`; missing or ambiguous matches fail before implementation.
|
||||
- Missing caller adapters fail before workspace/process work; no fake success or downstream report work exists.
|
||||
|
||||
## Verification Results
|
||||
|
||||
### `python3 -c 'from pathlib import Path; g="m-agent-comparison-benchmark-pipeline"; ids=("01","02","03"); a=Path("agent-task")/g; r=Path("agent-task/archive"); found={i:sorted([*a.glob(f"{i}_*/complete.log"),*a.glob(f"{i}+*/complete.log"),*r.glob(f"*/*/{g}/{i}_*/complete.log"),*r.glob(f"*/*/{g}/{i}+*/complete.log")],key=str) for i in ids}; bad={i:[str(p) for p in ps] for i,ps in found.items() if len(ps)!=1}; assert not bad,bad; print("\n".join(str(found[i][0]) for i in ids))'`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 -m unittest scripts.agent_benchmark.attempts_test`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `make test-agent-comparison-benchmark`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `git diff --check`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results | Fixed headings/commands | Implementing agent fills actual output only |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,129 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt plan=2 tag=API milestone-task=repeat-attempt -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record exact evidence and the resume condition only in implementation-owned fields.
|
||||
> Do not ask the user, call user-input tools, create stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only.
|
||||
> Follow the ownership table at the bottom of this file.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-09
|
||||
task=m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt, plan=2, tag=API
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/plan_cloud_G07_1.log` and `agent-task/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/code_review_cloud_G08_1.log`.
|
||||
- Review state: unimplemented, no official verdict, replaced through explicit plan `write` mode.
|
||||
- Corrected scope: immutable manifest/run identity, explicit workspace/lifecycle integration, and identity-checked cleanup before crash recovery.
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** Implementing agents must not execute this section.
|
||||
|
||||
Compare every item to source and verify pasted command output.
|
||||
|
||||
1. Append verdict and verified routing signals.
|
||||
2. Archive this review to `code_review_cloud_G08_2.log` and the plan to `plan_cloud_G07_2.log`.
|
||||
3. On PASS, create `complete.log` and move the task directory to its monthly group archive; otherwise write the required next state.
|
||||
4. On PASS, preserve/report `milestone-task=repeat-attempt`; roadmap evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check review-only items at the final log location.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|---------|
|
||||
| API-1 Bind runs and attempts to immutable identity | [ ] |
|
||||
| API-2 Run, recover, resume, and retry safely | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Persist an immutable canonical manifest snapshot/digest, deterministic repetition slots, and append-only atomically allocated attempt records.
|
||||
- [ ] Integrate prepared workspaces and lifecycle outcomes into run/resume/status plus explicit failed-attempt retry without evidence overwrite.
|
||||
- [ ] Recover nonterminal attempts by identity-checking and cleaning any surviving process before sealing interruption and allocating a successor; add crash/concurrency/retry tests.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one `PASS`, `WARN`, or `FAIL` verdict plus verified `review_rework_count` and `evidence_integrity_failure`.
|
||||
- [ ] Verify verdict, dimensions, and Required/Suggested/Nit classifications agree.
|
||||
- [ ] Archive active review to `code_review_cloud_G08_2.log`.
|
||||
- [ ] Archive active plan to `plan_cloud_G07_2.log`.
|
||||
- [ ] Verify `.gitignore` managed task/roadmap rules.
|
||||
- [ ] If PASS, write canonical `complete.log` and leave no active `.md` files.
|
||||
- [ ] If PASS, move to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/` and update this checklist there.
|
||||
- [ ] If PASS, preserve/report `milestone-task=repeat-attempt` without directly changing roadmap.
|
||||
- [ ] If PASS for split work, remove empty parent or justify remaining siblings.
|
||||
- [ ] If WARN/FAIL, materialize the required next state and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record any deviations from the plan and the rationale here._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record key design decisions here._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Run creation persists and verifies byte-immutable canonical manifest identity before any slot work.
|
||||
- Slot/attempt allocation is exclusive, ordered, append-only, and terminal records remain byte-identical.
|
||||
- Each allocated attempt calls predecessor workspace and lifecycle APIs exactly once in the documented order.
|
||||
- Resume identity-checks and cleans a recorded survivor before sealing interruption or allocating a successor; ambiguity fails closed.
|
||||
- Missing adapters cause zero allocation/workspace/process side effects; status and errors remain deterministic and redacted.
|
||||
|
||||
## Verification Results
|
||||
|
||||
### Predecessor completion check from `PLAN-cloud-G07.md`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 -m unittest scripts.agent_benchmark.attempts_test`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `make test-agent-comparison-benchmark`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `git diff --check`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementer must not modify or execute these |
|
||||
| Archive Evidence Snapshot | Fixed at stub creation from plan when present | Read only cited evidence when required |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementer checks status only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementer checks status only |
|
||||
| Review-Only Checklist | Review agent only | Implementer must not modify |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholders with evidence |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results (section headings + commands) | Fixed at stub creation | Fill output only; changes require deviation |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,164 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt plan=0 tag=API milestone-task=repeat-attempt -->
|
||||
|
||||
# Repeat and Attempt State
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections of `CODE_REVIEW-cloud-G08.md` is the mandatory final implementation step. Run every verification command and paste actual output. Do not finalize, archive, write `complete.log`, ask the user, or change predecessor ownership.
|
||||
|
||||
## Background
|
||||
|
||||
The manifest matrix, isolated workspace, and lifecycle need durable orchestration semantics. This packet creates append-only attempt identity plus run/resume/status behavior while deliberately leaving real caller adapters to the next Epic.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `AGENTS.md`
|
||||
- `agent-ops/rules/project/rules.md`
|
||||
- `agent-ops/rules/project/domain/testing/rules.md`
|
||||
- `agent-test/local/rules.md`
|
||||
- `agent-test/local/testing-smoke.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/agent-comparison-benchmark-pipeline.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md`
|
||||
- `agent-spec/input/openai-compatible-surface.md`
|
||||
- `agent-spec/runtime/provider-pool-config-refresh.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
- `scripts/e2e-single-request-claude.sh`
|
||||
- `Makefile`
|
||||
- `.gitignore`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- Approved/released SDD scenario S05 requires default repetition `1`, independent append-only attempts, preserved failures, and resume without overwriting evidence.
|
||||
- Evidence Map row S05 directly becomes API-1's append-only/concurrency evidence, API-2's resume/retry cases, and the focused plus aggregate Final Verification commands; S01 supplies canonical matrix expansion and S03/S04 supply predecessor contracts.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- Local Python 3.12.3/Linux arm64; standard-library tests use temporary directories and injected fake adapters only.
|
||||
- Handoff: no separate verification handoff was supplied; the approved SDD, local test rules, and predecessor contracts are the verification sources.
|
||||
- Real Claude/Agy/Codex adapters are not yet present. The command must fail closed with a stable capability error rather than simulate a scored success.
|
||||
- No provider, network, browser, or sibling-checkout mutation is allowed.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
No current code allocates attempts atomically, preserves terminal records, resumes interrupted matrices, or distinguishes explicit retry from overwrite.
|
||||
|
||||
### Symbol References
|
||||
|
||||
Resolve the final predecessor manifest/workspace/lifecycle symbols after their completion logs exist. Do not duplicate their validation, copying, supervision, or event logic.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
Attempt identity, atomic persistence, matrix progress, resume, and retry are one state-consistency boundary with filesystem side effects; this is a large slice. Predecessors `01_benchmark_manifest`, `02+01_isolated_workspace`, and `03+01_run_lifecycle` are each missing their active or archived `complete.log` at planning time.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
Excluded: actual caller command adapters, provider translation, usage normalization beyond predecessor events, scoring, web checks, aggregate report generation, and adapter-specific retry policy.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- evaluation_mode: `first-pass`
|
||||
- finalizer: `finalize-task-policy.sh`, mode `pair`
|
||||
- build closures all `true`; scores `2/2/1/1/1`; route `risk-boundary`; grade `G07`; catalog `worker/cloud/G07`; filename `PLAN-cloud-G07.md`.
|
||||
- review closures all `true`; scores `2/2/1/2/1`; route `official-review`; grade `G08`; catalog `review/cloud/G08`; filename `CODE_REVIEW-cloud-G08.md`.
|
||||
- risks: `temporal_state`, `concurrent_consistency`, `boundary_contract`, `variant_product`; large_indivisible_context `false`; rework `0`; evidence failure `false`; capability gap none.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement canonical run expansion and append-only, atomically allocated attempt records.
|
||||
- [ ] Implement run/resume/status and explicit failed-attempt retry without evidence overwrite.
|
||||
- [ ] Add crash, concurrency, retry, ordering, and unavailable-adapter tests and run all verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Persist canonical attempts without overwrite
|
||||
|
||||
**Problem:** Repetitions and failures have no durable identity; rerunning could silently replace evidence or lose which matrix cell was attempted.
|
||||
|
||||
**Solution:** Expand validated cells in canonical order and repetitions from `1..N`. Allocate each attempt using exclusive creation under the prepared run root, with immutable identity tying manifest digest, run id, cell id, repetition, and monotonically increasing attempt number. Publish state transitions atomically. Terminal `success`, `failure`, `timeout`, `cancelled`, and `interrupted` records are never mutated; a recovered nonterminal record is sealed interrupted before a new attempt is allocated.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
No durable run, repetition, or attempt state exists.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
attempt = store.allocate(run_id, cell_id, repetition)
|
||||
store.transition(attempt, terminal_state)
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_benchmark/attempts.py` with frozen state types, canonical expansion, exclusive allocation, atomic transitions, and read-only status projection.
|
||||
- [ ] Update `scripts/agent_benchmark/__init__.py` to export the stable attempt API.
|
||||
|
||||
**Test Strategy:** Cover repetition default/override, matrix ordering, identity stability, terminal immutability, crash recovery, corrupt/foreign manifest state, and concurrent allocation.
|
||||
|
||||
**Verification:** `python3 -m unittest scripts.agent_benchmark.attempts_test` exits 0.
|
||||
|
||||
### [API-2] Expose run, resume, status, and explicit retry
|
||||
|
||||
**Problem:** There is no supported command path to continue an interrupted matrix or retry a failed cell while retaining original evidence.
|
||||
|
||||
**Solution:** Extend the CLI with `run`, `resume`, and `status`; add `--retry-failed` only to resume. `run` refuses an existing run id. `resume` skips successful terminal cells, seals interrupted records, and allocates new attempt ids for unfinished or explicitly retried failed cells. Failed/timeout/cancelled records remain terminal unless explicit retry is requested. Resolve adapters through an injected registry; absent real adapters return a stable unavailable-capability result before workspace/process work, never a fabricated success.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
The benchmark CLI cannot run, resume, inspect, or explicitly retry a matrix.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```text
|
||||
agent_comparison_benchmark.py run|resume|status ...
|
||||
resume --retry-failed allocates new attempts and preserves prior evidence.
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Complete orchestration and resume/retry policy in `scripts/agent_benchmark/attempts.py`.
|
||||
- [ ] Update `scripts/agent_comparison_benchmark.py` with `run`, `resume`, and `status` argument/exit contracts.
|
||||
- [ ] Add `scripts/agent_benchmark/attempts_test.py` with CLI, concurrency, crash, skip, retry, and adapter-unavailable cases.
|
||||
- [ ] Record exact output in `CODE_REVIEW-cloud-G08.md`.
|
||||
|
||||
**Test Strategy:** Inject deterministic fake adapters for lifecycle outcomes. Assert invocation counts, state order, old-record byte identity, stable sanitized status, and no filesystem work when adapter capability is absent.
|
||||
|
||||
**Verification:** Focused and aggregate tests pass; no external caller runs.
|
||||
|
||||
## Dependencies and Execution Order
|
||||
|
||||
1. Required predecessors:
|
||||
- `agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/complete.log`
|
||||
- `agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/complete.log`
|
||||
- `agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/complete.log`
|
||||
2. At planning time all three completion records are absent; do not implement until every path exists.
|
||||
3. Resolve completed APIs, implement API-1, then API-2, then run verification from fresh temporary state.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Items |
|
||||
|------|-------|
|
||||
| `scripts/agent_benchmark/attempts.py` | API-1, API-2 |
|
||||
| `scripts/agent_benchmark/attempts_test.py` | API-2 |
|
||||
| `scripts/agent_benchmark/__init__.py` | API-1 |
|
||||
| `scripts/agent_comparison_benchmark.py` | API-2 |
|
||||
| `agent-task/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/CODE_REVIEW-cloud-G08.md` | API-2 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
1. `test -f agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/complete.log && test -f agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/complete.log && test -f agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/complete.log`
|
||||
- Expected: exit 0 before implementation begins.
|
||||
2. `python3 -m unittest scripts.agent_benchmark.attempts_test`
|
||||
- Expected: canonical, concurrency, resume, retry, immutability, and unavailable-adapter cases pass.
|
||||
3. `make test-agent-comparison-benchmark`
|
||||
- Expected: all credential-free benchmark tests pass without a provider call.
|
||||
4. `git diff --check`
|
||||
- Expected: exit 0 with no whitespace errors.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,170 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt plan=1 tag=API milestone-task=repeat-attempt -->
|
||||
|
||||
# Repeat and Attempt State
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections of `CODE_REVIEW-cloud-G08.md` is the mandatory final implementation step. Run every verification command and paste actual output. Do not finalize, archive, write `complete.log`, ask the user, or change predecessor ownership.
|
||||
|
||||
## Background
|
||||
|
||||
The manifest matrix, isolated workspace, and lifecycle need durable orchestration semantics. This packet creates append-only attempt identity plus run/resume/status behavior while deliberately leaving real caller adapters to the next Epic.
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/plan_cloud_G07_0.log` and `agent-task/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/code_review_cloud_G08_0.log`.
|
||||
- Review state: the prior pair was unimplemented and had no official verdict; it was archived by the requested Epic self-review replan.
|
||||
- Self-review defect: its predecessor checks named only active task paths, but PASS predecessors are normally moved under `agent-task/archive/YYYY/MM/`.
|
||||
- Scope carried forward: `repeat-attempt`, SDD S05, implementation files, and credential-free verification remain unchanged; only dependency resolution and paired evidence are corrected.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `AGENTS.md`
|
||||
- `agent-ops/rules/project/rules.md`
|
||||
- `agent-ops/rules/project/domain/testing/rules.md`
|
||||
- `agent-test/local/rules.md`
|
||||
- `agent-test/local/testing-smoke.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/agent-comparison-benchmark-pipeline.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md`
|
||||
- `agent-spec/input/openai-compatible-surface.md`
|
||||
- `agent-spec/runtime/provider-pool-config-refresh.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
- `agent-ops/skills/common/plan/SKILL.md`
|
||||
- `scripts/e2e-single-request-claude.sh`
|
||||
- `Makefile`
|
||||
- `.gitignore`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- Approved/released SDD scenario S05 requires default repetition `1`, independent append-only attempts, preserved failures, and resume without overwriting evidence.
|
||||
- Evidence Map row S05 directly becomes API-1's append-only/concurrency evidence, API-2's resume/retry cases, and the focused plus aggregate Final Verification commands; S01 supplies canonical matrix expansion and S03/S04 supply predecessor contracts.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- Local Python 3.12.3/Linux arm64; standard-library tests use temporary directories and injected fake adapters only.
|
||||
- Handoff: no separate verification handoff was supplied; the approved SDD, local test rules, and predecessor contracts are the verification sources.
|
||||
- Real Claude/Agy/Codex adapters are not yet present. The command must fail closed with a stable capability error rather than simulate a scored success.
|
||||
- No provider, network, browser, or sibling-checkout mutation is allowed.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
No current code allocates attempts atomically, preserves terminal records, resumes interrupted matrices, or distinguishes explicit retry from overwrite.
|
||||
|
||||
### Symbol References
|
||||
|
||||
Resolve the final predecessor manifest/workspace/lifecycle symbols after their completion logs exist. Do not duplicate their validation, copying, supervision, or event logic.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
Attempt identity, atomic persistence, matrix progress, resume, and retry are one state-consistency boundary with filesystem side effects; this is a large slice. Predecessors `01_benchmark_manifest`, `02+01_isolated_workspace`, and `03+01_run_lifecycle` are each missing their active or archived `complete.log` at planning time.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
Excluded: actual caller command adapters, provider translation, usage normalization beyond predecessor events, scoring, web checks, aggregate report generation, and adapter-specific retry policy.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- evaluation_mode: `isolated-reassessment`
|
||||
- finalizer: `finalize-task-policy.sh`, mode `pair`
|
||||
- build closures all `true`; scores `2/2/1/1/1`; route `risk-boundary`; grade `G07`; catalog `worker/cloud/G07`; filename `PLAN-cloud-G07.md`.
|
||||
- review closures all `true`; scores `2/2/1/2/1`; route `official-review`; grade `G08`; catalog `review/cloud/G08`; filename `CODE_REVIEW-cloud-G08.md`.
|
||||
- risks: `temporal_state`, `concurrent_consistency`, `boundary_contract`, `variant_product`; large_indivisible_context `false`; rework `0`; evidence failure `false`; capability gap none.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Implement canonical run expansion and append-only, atomically allocated attempt records.
|
||||
- [ ] Implement run/resume/status and explicit failed-attempt retry without evidence overwrite.
|
||||
- [ ] Resolve all indexed predecessors from the allowed active/archive candidates, add crash, concurrency, retry, ordering, and unavailable-adapter tests, and run all verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Persist canonical attempts without overwrite
|
||||
|
||||
**Problem:** Repetitions and failures have no durable identity; rerunning could silently replace evidence or lose which matrix cell was attempted.
|
||||
|
||||
**Solution:** Expand validated cells in canonical order and repetitions from `1..N`. Allocate each attempt using exclusive creation under the prepared run root, with immutable identity tying manifest digest, run id, cell id, repetition, and monotonically increasing attempt number. Publish state transitions atomically. Terminal `success`, `failure`, `timeout`, `cancelled`, and `interrupted` records are never mutated; a recovered nonterminal record is sealed interrupted before a new attempt is allocated.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
No durable run, repetition, or attempt state exists.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
attempt = store.allocate(run_id, cell_id, repetition)
|
||||
store.transition(attempt, terminal_state)
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_benchmark/attempts.py` with frozen state types, canonical expansion, exclusive allocation, atomic transitions, and read-only status projection.
|
||||
- [ ] Update `scripts/agent_benchmark/__init__.py` to export the stable attempt API.
|
||||
|
||||
**Test Strategy:** Cover repetition default/override, matrix ordering, identity stability, terminal immutability, crash recovery, corrupt/foreign manifest state, and concurrent allocation.
|
||||
|
||||
**Verification:** `python3 -m unittest scripts.agent_benchmark.attempts_test` exits 0.
|
||||
|
||||
### [API-2] Expose run, resume, status, and explicit retry
|
||||
|
||||
**Problem:** There is no supported command path to continue an interrupted matrix or retry a failed cell while retaining original evidence.
|
||||
|
||||
**Solution:** Extend the CLI with `run`, `resume`, and `status`; add `--retry-failed` only to resume. `run` refuses an existing run id. `resume` skips successful terminal cells, seals interrupted records, and allocates new attempt ids for unfinished or explicitly retried failed cells. Failed/timeout/cancelled records remain terminal unless explicit retry is requested. Resolve adapters through an injected registry; absent real adapters return a stable unavailable-capability result before workspace/process work, never a fabricated success.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
The benchmark CLI cannot run, resume, inspect, or explicitly retry a matrix.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```text
|
||||
agent_comparison_benchmark.py run|resume|status ...
|
||||
resume --retry-failed allocates new attempts and preserves prior evidence.
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Complete orchestration and resume/retry policy in `scripts/agent_benchmark/attempts.py`.
|
||||
- [ ] Update `scripts/agent_comparison_benchmark.py` with `run`, `resume`, and `status` argument/exit contracts.
|
||||
- [ ] Add `scripts/agent_benchmark/attempts_test.py` with CLI, concurrency, crash, skip, retry, and adapter-unavailable cases.
|
||||
- [ ] Record exact output in `CODE_REVIEW-cloud-G08.md`.
|
||||
|
||||
**Test Strategy:** Inject deterministic fake adapters for lifecycle outcomes. Assert invocation counts, state order, old-record byte identity, stable sanitized status, and no filesystem work when adapter capability is absent.
|
||||
|
||||
**Verification:** Focused and aggregate tests pass; no external caller runs.
|
||||
|
||||
## Dependencies and Execution Order
|
||||
|
||||
1. Required predecessor indices are `01_benchmark_manifest`, `02+01_isolated_workspace`, and `03+01_run_lifecycle`, encoded by the `04+01,02,03_...` task directory.
|
||||
2. At planning time none has an allowed active/archive `complete.log` candidate. Runtime scheduling must wait; completed predecessors may be under the active task group or `agent-task/archive/YYYY/MM/`.
|
||||
3. Before implementation, require exactly one matching completion record for each index across the allowed `NN_*/complete.log` and `NN+*/complete.log` locations. Missing or multiple matches fail closed and must not be guessed.
|
||||
4. Resolve completed APIs, implement API-1, then API-2, then run verification from fresh temporary state.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Items |
|
||||
|------|-------|
|
||||
| `scripts/agent_benchmark/attempts.py` | API-1, API-2 |
|
||||
| `scripts/agent_benchmark/attempts_test.py` | API-2 |
|
||||
| `scripts/agent_benchmark/__init__.py` | API-1 |
|
||||
| `scripts/agent_comparison_benchmark.py` | API-2 |
|
||||
| `agent-task/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/CODE_REVIEW-cloud-G08.md` | API-2 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
1. `python3 -c 'from pathlib import Path; g="m-agent-comparison-benchmark-pipeline"; ids=("01","02","03"); a=Path("agent-task")/g; r=Path("agent-task/archive"); found={i:sorted([*a.glob(f"{i}_*/complete.log"),*a.glob(f"{i}+*/complete.log"),*r.glob(f"*/*/{g}/{i}_*/complete.log"),*r.glob(f"*/*/{g}/{i}+*/complete.log")],key=str) for i in ids}; bad={i:[str(p) for p in ps] for i,ps in found.items() if len(ps)!=1}; assert not bad,bad; print("\n".join(str(found[i][0]) for i in ids))'`
|
||||
- Expected: exit 0 and print exactly one active or archived completion path for each predecessor index before implementation begins.
|
||||
2. `python3 -m unittest scripts.agent_benchmark.attempts_test`
|
||||
- Expected: canonical, concurrency, resume, retry, immutability, and unavailable-adapter cases pass.
|
||||
3. `make test-agent-comparison-benchmark`
|
||||
- Expected: all credential-free benchmark tests pass without a provider call.
|
||||
4. `git diff --check`
|
||||
- Expected: exit 0 with no whitespace errors.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,178 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt plan=2 tag=API milestone-task=repeat-attempt -->
|
||||
|
||||
# Repeat and Attempt State
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections of `CODE_REVIEW-cloud-G08.md` is mandatory. Resolve predecessors `01`, `02`, and `03` before editing, run every verification command, paste actual output, and leave the active pair for official review. Do not finalize, archive, write `complete.log`, ask the user, or change predecessor ownership.
|
||||
|
||||
## Background
|
||||
|
||||
The validated manifest, isolated workspace, and bounded lifecycle need one durable orchestration boundary. This packet owns immutable run identity, append-only attempt allocation, crash-safe resume, explicit retry, and sanitized status. Real caller adapters remain a later Epic.
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/plan_cloud_G07_1.log` and `agent-task/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/code_review_cloud_G08_1.log` (generation 1 retains generation 0 history).
|
||||
- Review state: unimplemented, no official verdict, replaced through explicit plan `write` mode.
|
||||
- Self-review defects: the prior plan did not bind resume to an immutable canonical manifest snapshot, did not make workspace/lifecycle calls explicit, and would seal an interrupted attempt before identity-checking and cleaning a possibly surviving process.
|
||||
- Scope carried forward: `repeat-attempt`, S05, append-only attempts, run/resume/status/retry, predecessor resolution, and credential-free tests.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `AGENTS.md`
|
||||
- `agent-ops/rules/project/rules.md`
|
||||
- `agent-ops/rules/project/domain/testing/rules.md`
|
||||
- `agent-ops/rules/common/rules-roadmap.md`
|
||||
- `agent-ops/skills/common/plan/SKILL.md`
|
||||
- `agent-ops/skills/common/finalize-task-routing/SKILL.md`
|
||||
- `agent-test/local/rules.md`
|
||||
- `agent-test/local/testing-smoke.md`
|
||||
- `agent-roadmap/current.md`
|
||||
- `agent-roadmap/priority-queue.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/PHASE.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/agent-comparison-benchmark-pipeline.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md`
|
||||
- `agent-spec/index.md`
|
||||
- `agent-spec/input/openai-compatible-surface.md`
|
||||
- `agent-spec/runtime/provider-pool-config-refresh.md`
|
||||
- `agent-contract/index.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
- `Makefile`
|
||||
- `.gitignore`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- S05 requires default repetition `1`, independent append-only attempts, preserved failures, and resume without evidence overwrite.
|
||||
- S01 supplies canonical matrix identity; S03 supplies clean workspace/session identity; S04 supplies the bounded lifecycle outcome consumed by each attempt.
|
||||
- Evidence Map S05 requires deterministic repetition, retry, concurrency, and resume evidence.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- Local Python 3.12.3/Linux arm64; standard-library tests use temporary directories and injected adapters only.
|
||||
- `bench-01` is unblocked, but predecessors `01`, `02`, and `03` have no completion records, so implementation must wait.
|
||||
- No provider, network, browser, or sibling-checkout mutation is permitted. Missing adapters fail before workspace or process creation.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
No current code persists an immutable run snapshot, owns a repetition slot, atomically allocates attempts, integrates workspace/lifecycle, cleans a surviving crashed process, or preserves old evidence through resume/retry.
|
||||
|
||||
### Symbol References
|
||||
|
||||
Resolve completed manifest canonicalization, workspace preparation, and lifecycle invocation/result/process-locator types before editing. Do not duplicate those contracts.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
Manifest identity, slot ownership, process recovery, transition atomicity, and retry policy form one crash-consistency boundary. Splitting them would permit a new attempt while an old process or mutable identity remains active.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
Excluded: real caller command adapters, provider translation, scoring, browser checks, aggregate reporting, and adapter-specific retry policy.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- evaluation_mode: `isolated-reassessment`
|
||||
- finalizer: `finalize-task-policy.sh`, mode `pair`
|
||||
- build closures all `true`; scores `2/2/1/1/1`; base `local-fit`, route `risk-boundary`; grade `G07`; catalog `worker/cloud/G07`; filename `PLAN-cloud-G07.md`.
|
||||
- review closures all `true`; scores `2/2/1/2/1`; route `official-review`; grade `G08`; catalog `review/cloud/G08`; filename `CODE_REVIEW-cloud-G08.md`.
|
||||
- large_indivisible_context `false`; risks `temporal_state`, `concurrent_consistency`, `boundary_contract`, `variant_product`; rework `0`; evidence integrity failure `false`; capability gap none.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Persist an immutable canonical manifest snapshot/digest, deterministic repetition slots, and append-only atomically allocated attempt records.
|
||||
- [ ] Integrate prepared workspaces and lifecycle outcomes into run/resume/status plus explicit failed-attempt retry without evidence overwrite.
|
||||
- [ ] Recover nonterminal attempts by identity-checking and cleaning any surviving process before sealing interruption and allocating a successor; add crash/concurrency/retry tests.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Bind runs and attempts to immutable identity
|
||||
|
||||
**Problem:** A run can be resumed against changed manifest bytes or allocate overlapping work for the same cell/repetition, making evidence incomparable or overwritable.
|
||||
|
||||
**Solution:** On new run creation, consume the predecessor loader's canonical manifest form/digest and atomically persist its byte-immutable snapshot plus digest under `<output_root>/<run_id>/`. Reject an existing run id and reject resume when supplied manifest bytes/digest or canonical expansion differ. Expand cells in canonical id order and repetitions `1..N`; define each `(manifest_digest, run_id, cell_id, repetition)` as one slot. Allocate monotonically increasing attempt ids with exclusive filesystem creation. Atomically publish transitions. Terminal records never change.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
No durable run snapshot, slot identity, or attempt allocation exists.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
run = store.create(manifest_snapshot, manifest_digest)
|
||||
attempt = store.allocate(run.slot(cell_id, repetition))
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_benchmark/attempts.py` with frozen run/slot/attempt types, snapshot verification, canonical expansion, exclusive allocation, and atomic transitions.
|
||||
- [ ] Update `scripts/agent_benchmark/__init__.py` to export only stable attempt APIs.
|
||||
|
||||
**Test Strategy:** Cover repetition default/override, canonical order, resume digest/byte mismatch, corrupt/foreign state, concurrent slot allocation, monotonic ids, and terminal byte immutability.
|
||||
|
||||
**Verification:** `python3 -m unittest scripts.agent_benchmark.attempts_test` exits 0.
|
||||
|
||||
### [API-2] Run, recover, resume, and retry safely
|
||||
|
||||
**Problem:** A crashed controller can leave both a nonterminal record and a live process tree. Sealing it interrupted and starting again before cleanup creates concurrent scored work and contaminates evidence.
|
||||
|
||||
**Solution:** Add `run`, `resume`, and `status`; permit `--retry-failed` only on resume. For each pending slot, allocate an attempt, call the completed workspace preparation API, and invoke the completed lifecycle API with an `on_started` callback that atomically persists its identity/ownership process locator before submission or observation. Then atomically publish the lifecycle terminal outcome. If an adapter is unavailable, return `capability-unavailable: caller-adapter` before allocation, workspace creation, or process launch.
|
||||
|
||||
On resume, inspect each nonterminal attempt. If it has a process locator, prove identity matches the recorded process; use the lifecycle cleanup primitive and verify the group is gone before atomically sealing `interrupted`. Refuse resume when the locator is ambiguous, reused, or denotes a live process the harness cannot own. Only then allocate a new attempt. Skip successes; preserve failed/timeout/cancelled records unless explicit retry is requested. Status is read-only, deterministic, and redacted.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
No supported crash-safe continuation or explicit retry exists.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```text
|
||||
run|resume|status preserve the manifest snapshot and old attempt bytes.
|
||||
resume cleans an identity-matched survivor before sealing and reallocating.
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Complete workspace/lifecycle orchestration and recovery policy in `scripts/agent_benchmark/attempts.py`.
|
||||
- [ ] Update `scripts/agent_comparison_benchmark.py` with sanitized `run`, `resume`, and `status` contracts.
|
||||
- [ ] Add `scripts/agent_benchmark/attempts_test.py` with integration, crash-survivor, ambiguous-identity, concurrency, skip, retry, redaction, and unavailable-adapter cases.
|
||||
- [ ] Record exact output in `CODE_REVIEW-cloud-G08.md`.
|
||||
|
||||
**Test Strategy:** Inject deterministic adapters and real short-lived process fixtures. Assert workspace/lifecycle call counts, cleanup-before-interruption ordering, no surviving process, old-record byte identity, and no side effects on unavailable capability.
|
||||
|
||||
**Verification:** Focused and aggregate tests pass with no fixture process or temporary state left alive.
|
||||
|
||||
## Dependencies and Execution Order
|
||||
|
||||
1. Required predecessors `01`, `02`, and `03` are encoded by `04+01,02,03_...`.
|
||||
2. At planning time none has exactly one valid active/archive `complete.log`; runtime must wait.
|
||||
3. Before implementation, require exactly one allowed completion for every predecessor. Missing or multiple matches fail closed.
|
||||
4. Resolve final predecessor symbols, implement immutable identity, integrate workspace/lifecycle, add recovery/retry, then verify.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Items |
|
||||
|------|-------|
|
||||
| `scripts/agent_benchmark/attempts.py` | API-1, API-2 |
|
||||
| `scripts/agent_benchmark/attempts_test.py` | API-1, API-2 |
|
||||
| `scripts/agent_benchmark/__init__.py` | API-1 |
|
||||
| `scripts/agent_comparison_benchmark.py` | API-2 |
|
||||
| `agent-task/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/CODE_REVIEW-cloud-G08.md` | API-2 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
1. `python3 -c 'from pathlib import Path; g="m-agent-comparison-benchmark-pipeline"; ids=("01","02","03"); a=Path("agent-task")/g; r=Path("agent-task/archive"); found={i:sorted([*a.glob(f"{i}_*/complete.log"),*a.glob(f"{i}+*/complete.log"),*r.glob(f"*/*/{g}/{i}_*/complete.log"),*r.glob(f"*/*/{g}/{i}+*/complete.log")],key=str) for i in ids}; bad={i:[str(p) for p in ps] for i,ps in found.items() if len(ps)!=1}; assert not bad,bad; print("\n".join(str(found[i][0]) for i in ids))'`
|
||||
- Expected: exactly one allowed completion path for each predecessor before implementation.
|
||||
2. `python3 -m unittest scripts.agent_benchmark.attempts_test`
|
||||
- Expected: immutable identity, integration, crash recovery, concurrency, retry, redaction, and unavailable-adapter cases pass.
|
||||
3. `make test-agent-comparison-benchmark`
|
||||
- Expected: all credential-free benchmark tests pass with no process left alive.
|
||||
4. `git diff --check`
|
||||
- Expected: exit 0 with no whitespace errors.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,136 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill plan=3 tag=API milestone-task=benchmark-skill -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record exact evidence and the resume condition only in implementation-owned fields.
|
||||
> Do not ask the user, call user-input tools, create stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only.
|
||||
> Follow the ownership table at the bottom of this file.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-09
|
||||
task=m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill, plan=3, tag=API
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/plan_local_G02_2.log` and `agent-task/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/code_review_cloud_G03_2.log` (generation 2 retains earlier history).
|
||||
- Review state: unimplemented, no official verdict, replaced through explicit plan `write` mode.
|
||||
- Self-review defects: the prior pair advertised a public `prepare` operation even though the corrected workspace contract is internal and the attempt runner owns allocation. Leaving it in the skill would create a second stateful surface with no safe attempt-root owner.
|
||||
- Scope carried forward: `benchmark-skill`, S02, project routing, CLI parity tests, predecessor resolution, and credential-free verification.
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** Implementing agents must not execute this section.
|
||||
|
||||
Compare every item to source and verify pasted command output.
|
||||
|
||||
1. Append verdict and verified routing signals.
|
||||
2. Archive this review to `code_review_cloud_G03_3.log` and the plan to `plan_local_G02_3.log`.
|
||||
3. On PASS, create `complete.log` and move the task directory to its monthly group archive; otherwise write the required next state.
|
||||
4. On PASS, preserve/report `milestone-task=benchmark-skill`; roadmap evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check review-only items at the final log location.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|---------|
|
||||
| API-1 Create the project benchmark operator skill | [ ] |
|
||||
| API-2 Lock the skill to the executable surface | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Run create-skill preflight/template validation and create the project benchmark skill with exact validate/run/resume/status/report-readiness triggers, CLI delegation, safety rules, and capability gates.
|
||||
- [ ] Route the benchmark request family in project rules and add deterministic skill/frontmatter/routing/CLI-help contract tests.
|
||||
- [ ] Resolve predecessors `01` through `04`, then run skill contract, CLI help, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one `PASS`, `WARN`, or `FAIL` verdict plus verified `review_rework_count` and `evidence_integrity_failure`.
|
||||
- [ ] Verify verdict, dimensions, and Required/Suggested/Nit classifications agree.
|
||||
- [ ] Archive active review to `code_review_cloud_G03_3.log`.
|
||||
- [ ] Archive active plan to `plan_local_G02_3.log`.
|
||||
- [ ] Verify `.gitignore` managed task/roadmap rules.
|
||||
- [ ] If PASS, write canonical `complete.log` and leave no active `.md` files.
|
||||
- [ ] If PASS, move to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/` and update this checklist there.
|
||||
- [ ] If PASS, preserve/report `milestone-task=benchmark-skill` without directly changing roadmap.
|
||||
- [ ] If PASS for split work, remove empty parent or justify remaining siblings.
|
||||
- [ ] If WARN/FAIL, materialize the required next state and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record any deviations from the plan and the rationale here._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record key design decisions here._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Router/create-skill preflight, destination ownership, duplicate check, template, and frontmatter validation were followed.
|
||||
- Project routing narrowly recognizes validate/run/resume/status/report-readiness benchmark intent; the internal workspace API is not a user command.
|
||||
- Supported stateful work delegates only to the deterministic CLI; the skill contains no second implementation or dispatcher.
|
||||
- Missing caller and report capabilities return exact `capability-unavailable: caller-adapter` and `capability-unavailable: report-output` results without fallback.
|
||||
- Contract tests bind frontmatter/routing/documented commands to real CLI help and forbid secret/provider/dispatcher behavior.
|
||||
|
||||
## Verification Results
|
||||
|
||||
### Predecessor completion check from `PLAN-local-G02.md`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 -m unittest scripts.agent_benchmark.skill_contract_test`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 scripts/agent_comparison_benchmark.py --help`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `make test-agent-comparison-benchmark`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `git diff --check`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementer must not modify or execute these |
|
||||
| Archive Evidence Snapshot | Fixed at stub creation from plan when present | Read only cited evidence when required |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementer checks status only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementer checks status only |
|
||||
| Review-Only Checklist | Review agent only | Implementer must not modify |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholders with evidence |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results (section headings + commands) | Fixed at stub creation | Fill output only; changes require deviation |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,182 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill plan=3 tag=API milestone-task=benchmark-skill -->
|
||||
|
||||
# Benchmark Operator Skill
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections of `CODE_REVIEW-cloud-G03.md` is mandatory. Resolve predecessors `01` through `04`, reread the router and create-skill instructions, run their preflight/template flow, execute every verification command, paste actual output, and leave the active pair for official review. Do not finalize, archive, write `complete.log`, ask the user, or change predecessor ownership.
|
||||
|
||||
## Background
|
||||
|
||||
The deterministic pipeline needs a project-owned operator skill that recognizes supported benchmark requests and delegates to the CLI. The skill is routing/documentation, not a second implementation or orchestration dispatcher. Caller adapters and report rendering remain later Epics and must fail closed.
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/plan_local_G02_2.log` and `agent-task/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/code_review_cloud_G03_2.log` (generation 2 retains earlier history).
|
||||
- Review state: unimplemented, no official verdict, replaced through explicit plan `write` mode.
|
||||
- Self-review defects: the prior pair advertised a public `prepare` operation even though the corrected workspace contract is internal and the attempt runner owns allocation. Leaving it in the skill would create a second stateful surface with no safe attempt-root owner.
|
||||
- Scope carried forward: `benchmark-skill`, S02, project routing, CLI parity tests, predecessor resolution, and credential-free verification.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `AGENTS.md`
|
||||
- `agent-ops/rules/project/rules.md`
|
||||
- `agent-ops/rules/project/domain/testing/rules.md`
|
||||
- `agent-ops/rules/common/rules-roadmap.md`
|
||||
- `agent-ops/skills/common/router.md`
|
||||
- `agent-ops/skills/common/create-skill/SKILL.md`
|
||||
- `agent-ops/skills/common/create-skill/templates/SKILL-template.md`
|
||||
- `agent-ops/skills/common/plan/SKILL.md`
|
||||
- `agent-ops/skills/common/code-review/SKILL.md`
|
||||
- `agent-ops/skills/common/finalize-task-routing/SKILL.md`
|
||||
- `agent-ops/skills/common/plan/templates/review-stub-template.md`
|
||||
- `agent-test/local/rules.md`
|
||||
- `agent-test/local/testing-smoke.md`
|
||||
- `agent-roadmap/current.md`
|
||||
- `agent-roadmap/priority-queue.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/PHASE.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/agent-comparison-benchmark-pipeline.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md`
|
||||
- `agent-spec/index.md`
|
||||
- `agent-spec/input/openai-compatible-surface.md`
|
||||
- `agent-spec/runtime/provider-pool-config-refresh.md`
|
||||
- `agent-contract/index.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
- `Makefile`
|
||||
- `.gitignore`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- S02 requires a skill to accept a manifest path, validate preconditions, and route supported execution into the benchmark pipeline.
|
||||
- The SDD separates caller adapters and report rendering into later Epics; this skill may recognize those intents but cannot simulate them.
|
||||
- Evidence Map S02 requires an executable skill/preflight contract and skill-to-CLI parity evidence.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- Local, credential-free documentation and contract tests only.
|
||||
- `.agent-ops-source` is absent, so the owned destination is `agent-ops/skills/project/` and the routing rule belongs in `agent-ops/rules/project/rules.md`.
|
||||
- No duplicate `iop-agent-comparison-benchmark` skill exists. Use the create-skill template/frontmatter exactly, then validate it.
|
||||
- Predecessors `01` through `04` have no completion records; implementation must wait.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
No project skill or rule recognizes benchmark validate/run/resume/status/report-readiness intent, enforces exact capability gates, or proves its documented commands match the CLI.
|
||||
|
||||
### Symbol References
|
||||
|
||||
Resolve final CLI help, exit codes, artifact names, and capability strings from completed predecessors. The skill must not invent options or implement state transitions in prose.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
The skill and its project routing rule are one agent-responsibility boundary. Contract tests must land with them so documentation cannot drift from the executable surface.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
Excluded: executable pipeline changes, caller adapters, credential discovery, scoring, web validation, Markdown report generation, subagents, dispatchers, roadmap mutation, and Git operations.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- evaluation_mode: `isolated-reassessment`
|
||||
- finalizer: `finalize-task-policy.sh`, mode `pair`
|
||||
- build closures all `true`; scores `1/0/0/0/1`; base/route `local-fit`; grade `G02`; catalog `worker/local/G02`; filename `PLAN-local-G02.md`.
|
||||
- review closures all `true`; scores `1/0/0/1/1`; route `official-review`; grade `G03`; catalog `review/cloud/G03`; filename `CODE_REVIEW-cloud-G03.md`.
|
||||
- large_indivisible_context `false`; risks `boundary_contract`, `variant_product`; rework `0`; evidence integrity failure `false`; capability gap none.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Run create-skill preflight/template validation and create the project benchmark skill with exact validate/run/resume/status/report-readiness triggers, CLI delegation, safety rules, and capability gates.
|
||||
- [ ] Route the benchmark request family in project rules and add deterministic skill/frontmatter/routing/CLI-help contract tests.
|
||||
- [ ] Resolve predecessors `01` through `04`, then run skill contract, CLI help, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Create the project benchmark operator skill
|
||||
|
||||
**Problem:** Agents have no supported operational entrypoint and could bypass validation, mutate the fixed testbed, leak credentials, or claim unavailable adapter/report behavior.
|
||||
|
||||
**Solution:** Follow the create-skill preflight and template to add `iop-agent-comparison-benchmark` under the project skill root. Give it explicit triggers for benchmark manifest validation, run, resume, status, and report-readiness requests. Require manifest path for validate/run, exact harness-issued run id for resume/status, fixed clean `../iop-s2` provenance, contained output, fresh-session/cache policy, and CLI help/capability preflight. Delegate every supported stateful operation verbatim to `scripts/agent_comparison_benchmark.py`; do not expose the internal workspace API or reproduce pipeline policy in prose.
|
||||
|
||||
Until later-Epic commands exist, caller execution must return `capability-unavailable: caller-adapter` and report/output requests must return `capability-unavailable: report-output`. Stop without fallback, fabricated evidence, ad-hoc provider calls, subagents, or orchestration dispatchers.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
No project skill routes benchmark pipeline operations or capability gates.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```text
|
||||
iop-agent-comparison-benchmark delegates supported CLI operations and fails closed for caller-adapter/report-output.
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md` from the repository template with valid frontmatter, triggers, preflight, supported commands, artifacts, safety, and stop conditions.
|
||||
- [ ] Update `agent-ops/rules/project/rules.md` with one narrow routing rule covering the benchmark pipeline request family, including report readiness.
|
||||
|
||||
**Test Strategy:** Parse frontmatter and required sections; assert the skill contains the exact two capability strings, no public `prepare`, no dispatcher/provider fallback, and only commands exposed by real CLI help.
|
||||
|
||||
**Verification:** `python3 -m unittest scripts.agent_benchmark.skill_contract_test` exits 0.
|
||||
|
||||
### [API-2] Lock the skill to the executable surface
|
||||
|
||||
**Problem:** Skill prose can drift into a second, unsafe implementation or advertise later-Epic commands before they exist.
|
||||
|
||||
**Solution:** Add a credential-free contract test that checks the create-skill template/frontmatter invariants, project rule routing, documented command forms against `--help`, supported/unsupported capability matrix, absence of the internal workspace API from user routing, and forbidden dispatcher/secret/fallback language. Invoke only help and tracked text; never run a stateful benchmark command.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
No executable check binds benchmark skill instructions to the CLI.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
self.assert_skill_commands_match_cli_help()
|
||||
self.assert_capability("report-output", available=False)
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_benchmark/skill_contract_test.py` covering template/frontmatter, routing, CLI parity, capability gates, and forbidden behavior.
|
||||
- [ ] Record exact output in `CODE_REVIEW-cloud-G03.md`.
|
||||
|
||||
**Test Strategy:** Read tracked files and invoke only `python3 scripts/agent_comparison_benchmark.py --help`; no workspace, provider, browser, or process lifecycle is started.
|
||||
|
||||
**Verification:** Skill contract, CLI help, aggregate tests, and patch integrity pass.
|
||||
|
||||
## Dependencies and Execution Order
|
||||
|
||||
1. Required predecessors `01`, `02`, `03`, and `04` are encoded by `05+01,02,03,04_...`.
|
||||
2. At planning time none has exactly one valid active/archive `complete.log`; runtime must wait.
|
||||
3. Before implementation, require exactly one allowed completion for every predecessor. Missing or multiple matches fail closed.
|
||||
4. Reread router/create-skill, rerun preflight, inspect completed CLI help, create from template, add routing/tests, then verify.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Items |
|
||||
|------|-------|
|
||||
| `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md` | API-1 |
|
||||
| `agent-ops/rules/project/rules.md` | API-1 |
|
||||
| `scripts/agent_benchmark/skill_contract_test.py` | API-2 |
|
||||
| `agent-task/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/CODE_REVIEW-cloud-G03.md` | API-2 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
1. `python3 -c 'from pathlib import Path; g="m-agent-comparison-benchmark-pipeline"; ids=("01","02","03","04"); a=Path("agent-task")/g; r=Path("agent-task/archive"); found={i:sorted([*a.glob(f"{i}_*/complete.log"),*a.glob(f"{i}+*/complete.log"),*r.glob(f"*/*/{g}/{i}_*/complete.log"),*r.glob(f"*/*/{g}/{i}+*/complete.log")],key=str) for i in ids}; bad={i:[str(p) for p in ps] for i,ps in found.items() if len(ps)!=1}; assert not bad,bad; print("\n".join(str(found[i][0]) for i in ids))'`
|
||||
- Expected: exactly one allowed completion path for each predecessor before implementation.
|
||||
2. `python3 -m unittest scripts.agent_benchmark.skill_contract_test`
|
||||
- Expected: template/frontmatter, routing, CLI parity, safety, and capability-gate cases pass.
|
||||
3. `python3 scripts/agent_comparison_benchmark.py --help`
|
||||
- Expected: exit 0 and only completed-pipeline subcommands are documented by the skill.
|
||||
4. `make test-agent-comparison-benchmark`
|
||||
- Expected: all credential-free benchmark tests pass.
|
||||
5. `git diff --check`
|
||||
- Expected: exit 0 with no whitespace errors.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,129 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill plan=0 tag=API milestone-task=benchmark-skill -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-09
|
||||
task=m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill, plan=0, tag=API
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files and verify that output in `Verification Results` matches code.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals.
|
||||
2. Archive `CODE_REVIEW-cloud-G03.md` → `code_review_cloud_G03_0.log` and `PLAN-local-G02.md` → `plan_local_G02_0.log`.
|
||||
3. If PASS, write `complete.log` and move active task directory to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/`. If WARN/FAIL, fully write the next filesystem state required by the code-review skill.
|
||||
4. If PASS, preserve `milestone-task=benchmark-skill` in `complete.log` and report it for runtime aggregation; roadmap evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check applicable `Review-Only Checklist` items at the final `.log` location before reporting.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|--------|
|
||||
| API-1 Create the project benchmark skill | [ ] |
|
||||
| API-2 Lock documentation to the executable surface | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Create the project benchmark skill with exact preflight, command, safety, and capability-gating behavior.
|
||||
- [ ] Route matching project requests to the skill and add deterministic skill/CLI contract tests.
|
||||
- [ ] Run predecessor, skill contract, CLI help, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Archive active `CODE_REVIEW-*-G??.md` to `code_review_cloud_G03_0.log`.
|
||||
- [ ] Archive active `PLAN-*-G??.md` to `plan_local_G02_0.log`.
|
||||
- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write `complete.log` based on `agent-ops/skills/common/code-review/templates/complete-log-template.md` and leave no active `.md` files.
|
||||
- [ ] If PASS, move active task directory `agent-task/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/` to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/` and update this checklist at the final archive path.
|
||||
- [ ] If PASS, preserve and report `milestone-task=benchmark-skill` for runtime aggregation, without modifying roadmap or directly calling `update-roadmap`.
|
||||
- [ ] If PASS for split work, remove empty active parent only when no sibling remains.
|
||||
- [ ] If WARN/FAIL, write the next filesystem state matching code-review verdict and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record deviations and rationale._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record implementation decisions._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- The skill uses the completed CLI as the only stateful execution boundary.
|
||||
- Manifest/testbed/output preflight and no-secret rules match the approved SDD.
|
||||
- Routing is narrow and no orchestration dispatcher or ad-hoc provider command is introduced.
|
||||
- Missing caller-adapter/report capabilities are reported and stopped, never simulated.
|
||||
|
||||
## Verification Results
|
||||
|
||||
### `test -f agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/complete.log && test -f agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/complete.log && test -f agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/complete.log && test -f agent-task/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/complete.log`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 -m unittest scripts.agent_benchmark.skill_contract_test`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 scripts/agent_comparison_benchmark.py --help`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `make test-agent-comparison-benchmark`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `git diff --check`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results | Fixed headings/commands | Implementing agent fills actual output only |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,137 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill plan=1 tag=API milestone-task=benchmark-skill -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record the exact blocker, attempted commands/output, and resume condition only in implementation-owned evidence fields.
|
||||
> Do not ask the user directly, present choices, call user-input tools, create control-plane stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only, even after compaction/resume.
|
||||
> Follow the ownership table at the bottom of this file for which sections you own.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-09
|
||||
task=m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill, plan=1, tag=API
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/plan_local_G02_0.log` and `agent-task/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/code_review_cloud_G03_0.log`.
|
||||
- Review state: the prior pair was unimplemented and had no official verdict; it was archived by the requested Epic self-review replan.
|
||||
- Self-review defect: its predecessor checks named only active task paths, but PASS predecessors are normally moved under `agent-task/archive/YYYY/MM/`.
|
||||
- Scope carried forward: `benchmark-skill`, SDD S02, implementation files, and credential-free verification remain unchanged; only dependency resolution and paired evidence are corrected.
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** The finalization steps below are review-agent only. Implementing agents must not execute this section.
|
||||
|
||||
Compare implementation of each item against source files and verify that output in `Verification Results` matches code.
|
||||
Review completion means the following steps are finished:
|
||||
|
||||
1. Append verdict and `review_rework_count` / `evidence_integrity_failure` routing signals.
|
||||
2. Archive `CODE_REVIEW-cloud-G03.md` → `code_review_cloud_G03_1.log` and `PLAN-local-G02.md` → `plan_local_G02_1.log`.
|
||||
3. If PASS, write `complete.log` and move active task directory to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/`. If WARN/FAIL, fully write the next filesystem state required by the code-review skill.
|
||||
4. If PASS, preserve `milestone-task=benchmark-skill` in `complete.log` and report it for runtime aggregation; roadmap evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check applicable `Review-Only Checklist` items at the final `.log` location before reporting.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|--------|
|
||||
| API-1 Create the project benchmark skill | [ ] |
|
||||
| API-2 Lock documentation to the executable surface | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Create the project benchmark skill with exact preflight, command, safety, and capability-gating behavior.
|
||||
- [ ] Route matching project requests to the skill and add deterministic skill/CLI contract tests.
|
||||
- [ ] Resolve all indexed predecessors from the allowed active/archive candidates, then run skill contract, CLI help, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** This checklist is used only by the review agent.
|
||||
> Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one verdict of `PASS`, `WARN`, or `FAIL` and verified `review_rework_count`, `evidence_integrity_failure` to `Code Review Result`.
|
||||
- [ ] Verify that verdict, `Dimension Assessment`, and Required/Suggested/Nit classifications match.
|
||||
- [ ] Archive active `CODE_REVIEW-*-G??.md` to `code_review_cloud_G03_1.log`.
|
||||
- [ ] Archive active `PLAN-*-G??.md` to `plan_local_G02_1.log`.
|
||||
- [ ] Verify that the Agent-Ops managed block in `.gitignore` unignores `agent-task/**/*.md` and `agent-task/**/*.log` and ignores `agent-roadmap/current.md`.
|
||||
- [ ] If PASS, write `complete.log` based on `agent-ops/skills/common/code-review/templates/complete-log-template.md` and leave no active `.md` files.
|
||||
- [ ] If PASS, move active task directory `agent-task/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/` to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/` and update this checklist at the final archive path.
|
||||
- [ ] If PASS, preserve and report `milestone-task=benchmark-skill` for runtime aggregation, without modifying roadmap or directly calling `update-roadmap`.
|
||||
- [ ] If PASS for split work, remove empty active parent only when no sibling remains.
|
||||
- [ ] If WARN/FAIL, write the next filesystem state matching code-review verdict and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record deviations and rationale._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record implementation decisions._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- The skill uses the completed CLI as the only stateful execution boundary.
|
||||
- Manifest/testbed/output preflight and no-secret rules match the approved SDD.
|
||||
- Routing is narrow and no orchestration dispatcher or ad-hoc provider command is introduced.
|
||||
- Predecessor indices `01` through `04` each resolve to exactly one active or monthly-archive `complete.log`; missing or ambiguous matches fail before implementation.
|
||||
- Missing caller-adapter/report capabilities are reported and stopped, never simulated.
|
||||
|
||||
## Verification Results
|
||||
|
||||
### `python3 -c 'from pathlib import Path; g="m-agent-comparison-benchmark-pipeline"; ids=("01","02","03","04"); a=Path("agent-task")/g; r=Path("agent-task/archive"); found={i:sorted([*a.glob(f"{i}_*/complete.log"),*a.glob(f"{i}+*/complete.log"),*r.glob(f"*/*/{g}/{i}_*/complete.log"),*r.glob(f"*/*/{g}/{i}+*/complete.log")],key=str) for i in ids}; bad={i:[str(p) for p in ps] for i,ps in found.items() if len(ps)!=1}; assert not bad,bad; print("\n".join(str(found[i][0]) for i in ids))'`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 -m unittest scripts.agent_benchmark.skill_contract_test`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 scripts/agent_comparison_benchmark.py --help`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `make test-agent-comparison-benchmark`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `git diff --check`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementing agent must not modify or execute these |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementing agent checks `[ ]` → `[x]` only |
|
||||
| Review-Only Checklist | Review agent only | Implementing agent must not modify or check this section |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholder text with actual content |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results | Fixed headings/commands | Implementing agent fills actual output only |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,136 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill plan=2 tag=API milestone-task=benchmark-skill -->
|
||||
|
||||
# Code Review Reference - API
|
||||
|
||||
> **[IMPLEMENTING AGENT — READ FIRST] Filling in this file is the mandatory final step of implementation.**
|
||||
> The task is NOT complete until every implementation-owned section below is filled in.
|
||||
> Complete the `Implementation Checklist`; the final checklist item is mandatory before saving.
|
||||
> Fill implementation-owned sections, then stop with active files in place and report ready for review.
|
||||
> Execute the plan's selected root cause, scope, files, and dependency decisions as written. Do not choose another owner, narrow/expand the write boundary, or replace a fix with another verification attempt.
|
||||
> If implementation is blocked, record exact evidence and the resume condition only in implementation-owned fields.
|
||||
> Do not ask the user, call user-input tools, create stop files, or classify the next state.
|
||||
> Finalization (`Code Review Result`, log rename, `complete.log`, archive moves, `Review-Only Checklist`) is review-agent-only.
|
||||
> Follow the ownership table at the bottom of this file.
|
||||
|
||||
## Overview
|
||||
|
||||
date=2026-08-09
|
||||
task=m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill, plan=2, tag=API
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/plan_local_G02_1.log` and `agent-task/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/code_review_cloud_G03_1.log` (generation 1 retains generation 0 history).
|
||||
- Review state: unimplemented, no official verdict, replaced through explicit plan `write` mode.
|
||||
- Self-review defects: the prior plan did not explicitly route report-intent requests to a stable fail-closed response and did not encode the create-skill preflight/template checks as implementation requirements.
|
||||
- Scope carried forward: `benchmark-skill`, S02, project routing, CLI parity tests, predecessor resolution, and credential-free verification.
|
||||
|
||||
## For the Review Agent
|
||||
|
||||
> **[REVIEW AGENT ONLY]** Implementing agents must not execute this section.
|
||||
|
||||
Compare every item to source and verify pasted command output.
|
||||
|
||||
1. Append verdict and verified routing signals.
|
||||
2. Archive this review to `code_review_cloud_G03_2.log` and the plan to `plan_local_G02_2.log`.
|
||||
3. On PASS, create `complete.log` and move the task directory to its monthly group archive; otherwise write the required next state.
|
||||
4. On PASS, preserve/report `milestone-task=benchmark-skill`; roadmap evaluation belongs to `sync-milestone-workstate`.
|
||||
5. Check review-only items at the final log location.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Item Completion
|
||||
|
||||
| Item | Status |
|
||||
|------|---------|
|
||||
| API-1 Create the project benchmark operator skill | [ ] |
|
||||
| API-2 Lock the skill to the executable surface | [ ] |
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Run create-skill preflight/template validation and create the project benchmark skill with exact request triggers, CLI delegation, safety rules, and capability gates.
|
||||
- [ ] Route the benchmark request family in project rules and add deterministic skill/frontmatter/routing/CLI-help contract tests.
|
||||
- [ ] Resolve predecessors `01` through `04`, then run skill contract, CLI help, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
## Review-Only Checklist
|
||||
|
||||
> **[REVIEW AGENT ONLY]** Implementing agents must not modify or check this section.
|
||||
|
||||
- [ ] Append one `PASS`, `WARN`, or `FAIL` verdict plus verified `review_rework_count` and `evidence_integrity_failure`.
|
||||
- [ ] Verify verdict, dimensions, and Required/Suggested/Nit classifications agree.
|
||||
- [ ] Archive active review to `code_review_cloud_G03_2.log`.
|
||||
- [ ] Archive active plan to `plan_local_G02_2.log`.
|
||||
- [ ] Verify `.gitignore` managed task/roadmap rules.
|
||||
- [ ] If PASS, write canonical `complete.log` and leave no active `.md` files.
|
||||
- [ ] If PASS, move to `agent-task/archive/YYYY/MM/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/` and update this checklist there.
|
||||
- [ ] If PASS, preserve/report `milestone-task=benchmark-skill` without directly changing roadmap.
|
||||
- [ ] If PASS for split work, remove empty parent or justify remaining siblings.
|
||||
- [ ] If WARN/FAIL, materialize the required next state and do not write `complete.log`.
|
||||
|
||||
## Deviations from Plan
|
||||
|
||||
_Record any deviations from the plan and the rationale here._
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
_Record key design decisions here._
|
||||
|
||||
## Reviewer Checkpoints
|
||||
|
||||
- Router/create-skill preflight, destination ownership, duplicate check, template, and frontmatter validation were followed.
|
||||
- Project routing narrowly recognizes validate/prepare/run/resume/status/report-readiness benchmark intent.
|
||||
- Supported stateful work delegates only to the deterministic CLI; the skill contains no second implementation or dispatcher.
|
||||
- Missing caller and report capabilities return exact `capability-unavailable: caller-adapter` and `capability-unavailable: report-output` results without fallback.
|
||||
- Contract tests bind frontmatter/routing/documented commands to real CLI help and forbid secret/provider/dispatcher behavior.
|
||||
|
||||
## Verification Results
|
||||
|
||||
### Predecessor completion check from `PLAN-local-G02.md`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 -m unittest scripts.agent_benchmark.skill_contract_test`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `python3 scripts/agent_comparison_benchmark.py --help`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `make test-agent-comparison-benchmark`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
### `git diff --check`
|
||||
|
||||
```text
|
||||
<actual output>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
> **[IMPLEMENTING AGENT — BEFORE SAVING] Have you filled in every implementation-owned section?**
|
||||
> If anything is blank, go back and fill it in before saving this file.
|
||||
> Leave review-agent-only sections unchanged.
|
||||
|
||||
## Section Ownership
|
||||
|
||||
| Section | Owner | Note |
|
||||
|---------|-------|------|
|
||||
| Header comment, Overview, Review Agent Instructions | Fixed at stub creation | Implementer must not modify or execute these |
|
||||
| Archive Evidence Snapshot | Fixed at stub creation from plan when present | Read only cited evidence when required |
|
||||
| Implementation Item Completion (item names) | Fixed at stub creation | Implementer checks status only |
|
||||
| Implementation Checklist (item text/order) | Fixed at stub creation from plan | Implementer checks status only |
|
||||
| Review-Only Checklist | Review agent only | Implementer must not modify |
|
||||
| Deviations from Plan, Key Design Decisions | Implementing agent | Replace placeholders with evidence |
|
||||
| Reviewer Checkpoints | Fixed at stub creation | Pre-filled from plan |
|
||||
| Verification Results (section headings + commands) | Fixed at stub creation | Fill output only; changes require deviation |
|
||||
| Code Review Result | Review agent appends | Not included in stub |
|
||||
|
|
@ -0,0 +1,164 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill plan=0 tag=API milestone-task=benchmark-skill -->
|
||||
|
||||
# Benchmark Operator Skill
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections of `CODE_REVIEW-cloud-G03.md` is the mandatory final implementation step. Before editing skill files, reread `agent-ops/skills/common/router.md` and follow the routed create-skill instructions. Run every verification command and paste actual output. Do not finalize, archive, write `complete.log`, ask the user, or change predecessor ownership.
|
||||
|
||||
## Background
|
||||
|
||||
Once the deterministic manifest/workspace/lifecycle/attempt pipeline exists, repository agents need a project-owned operational entrypoint. This packet documents and tests that entrypoint; executable product behavior remains in the completed CLI, not in prose or an orchestration dispatcher.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `AGENTS.md`
|
||||
- `agent-ops/rules/project/rules.md`
|
||||
- `agent-ops/rules/project/domain/testing/rules.md`
|
||||
- `agent-ops/skills/common/router.md`
|
||||
- `agent-ops/skills/common/create-skill/SKILL.md`
|
||||
- `agent-test/local/rules.md`
|
||||
- `agent-test/local/testing-smoke.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/agent-comparison-benchmark-pipeline.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md`
|
||||
- `agent-spec/input/openai-compatible-surface.md`
|
||||
- `agent-spec/runtime/provider-pool-config-refresh.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
- `scripts/e2e-single-request-claude.sh`
|
||||
- `Makefile`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- Approved/released SDD scenario S02 requires a skill to accept a manifest path, validate required inputs/preconditions, and route execution into the benchmark pipeline.
|
||||
- Evidence Map row S02 directly becomes API-1's skill/preflight contract, API-2's skill-to-CLI parity test, and the contract/help Final Verification commands while preserving no-secret, fixed-testbed, fresh-session, bounded-lifecycle, and immutable-evidence policy.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- Local, credential-free documentation/contract tests only.
|
||||
- Handoff: no separate verification handoff was supplied; the approved SDD, repository create-skill workflow, local test rules, and final predecessor CLI are the verification sources.
|
||||
- The skill must invoke deterministic `scripts/agent_comparison_benchmark.py` commands and must not launch subagents, orchestration dispatchers, providers, browsers, or report generation itself.
|
||||
- Caller adapters and final report output remain later Epics. Their absence must be surfaced as unavailable capability, not hidden or worked around.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
No project skill or routing rule currently exposes manifest validation, preparation, run, resume, or status for this benchmark.
|
||||
|
||||
### Symbol References
|
||||
|
||||
Resolve exact CLI help, exit behavior, and artifact names from all four completed predecessors. Documentation and contract tests must match those implemented symbols verbatim.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
Creating a project skill changes the agent responsibility/routing boundary, so it is not direct-small even though its file count is small. Predecessors `01_benchmark_manifest`, `02+01_isolated_workspace`, `03+01_run_lifecycle`, and `04+01,02,03_repeat_attempt` are each missing their active or archived `complete.log` at planning time.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
Excluded: new executable pipeline logic, caller adapters, credential discovery, scoring, web validation, Markdown report rendering, roadmap mutation, dispatcher use, and Git commit/push.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- evaluation_mode: `first-pass`
|
||||
- finalizer: `finalize-task-policy.sh`, mode `pair`
|
||||
- build closures all `true`; scores `1/0/0/0/1`; route `local-fit`; grade `G02`; catalog `worker/local/G02`; filename `PLAN-local-G02.md`.
|
||||
- review closures all `true`; scores `1/0/0/1/1`; route `official-review`; grade `G03`; catalog `review/cloud/G03`; filename `CODE_REVIEW-cloud-G03.md`.
|
||||
- risks: `boundary_contract`, `variant_product`; large_indivisible_context `false`; rework `0`; evidence failure `false`; capability gap none.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Create the project benchmark skill with exact preflight, command, safety, and capability-gating behavior.
|
||||
- [ ] Route matching project requests to the skill and add deterministic skill/CLI contract tests.
|
||||
- [ ] Run predecessor, skill contract, CLI help, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Create the project benchmark skill
|
||||
|
||||
**Problem:** Agents have no supported operational entrypoint and could bypass validation, mutate the testbed, leak credentials, or claim unsupported report/adapter behavior.
|
||||
|
||||
**Solution:** Create `iop-agent-comparison-benchmark` through the repository create-skill workflow. Give it narrow triggers for benchmark manifest validate/prepare/run/resume/status requests. Require the manifest path, approved fixed dev testbed, clean source, output containment, and explicit capability preflight. Delegate all stateful work to the deterministic CLI. For caller execution or final report requests before those later-Epic capabilities exist, report the exact unavailable capability and stop; never simulate output or fall back to ad-hoc provider commands. Explicitly forbid raw secrets and orchestration dispatchers.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
No project skill routes benchmark pipeline operations.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```text
|
||||
iop-agent-comparison-benchmark validates preconditions and delegates supported operations to scripts/agent_comparison_benchmark.py.
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md` with valid frontmatter, triggers, preflight, supported commands, artifact expectations, safety, and stop conditions.
|
||||
- [ ] Update `agent-ops/rules/project/rules.md` with one explicit routing rule for benchmark pipeline operations.
|
||||
|
||||
**Test Strategy:** Parse frontmatter and required sections; assert every documented command exists in real CLI help and unsupported later-Epic capabilities are gated.
|
||||
|
||||
**Verification:** `python3 -m unittest scripts.agent_benchmark.skill_contract_test` exits 0.
|
||||
|
||||
### [API-2] Lock documentation to the executable surface
|
||||
|
||||
**Problem:** A skill can drift from its CLI and become an unsafe second implementation of pipeline policy.
|
||||
|
||||
**Solution:** Add a credential-free contract test that extracts skill command forms and compares them with CLI subcommands/options. Assert the project rule maps only the intended request family, no dispatcher command is referenced, no secret-bearing option is documented, and run/resume capability failure language is present until adapters arrive.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
Skill prose and executable CLI behavior can drift without detection.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
self.assert_skill_commands_match_cli_help()
|
||||
self.assert_later_epic_capabilities_fail_closed()
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_benchmark/skill_contract_test.py` covering frontmatter, routing, command surface, forbidden behavior, and capability gates.
|
||||
- [ ] Record exact verification output in `CODE_REVIEW-cloud-G03.md`.
|
||||
|
||||
**Test Strategy:** Inspect tracked text and invoke only `--help`; do not run any stateful benchmark subcommand.
|
||||
|
||||
**Verification:** Skill contract, CLI help, aggregate tests, and patch integrity pass.
|
||||
|
||||
## Dependencies and Execution Order
|
||||
|
||||
1. Required predecessors:
|
||||
- `agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/complete.log`
|
||||
- `agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/complete.log`
|
||||
- `agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/complete.log`
|
||||
- `agent-task/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/complete.log`
|
||||
2. At planning time all completion records are absent; do not implement until every path exists.
|
||||
3. Reread the router/create-skill instructions, inspect final CLI help, implement API-1, then API-2.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Items |
|
||||
|------|-------|
|
||||
| `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md` | API-1 |
|
||||
| `agent-ops/rules/project/rules.md` | API-1 |
|
||||
| `scripts/agent_benchmark/skill_contract_test.py` | API-2 |
|
||||
| `agent-task/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/CODE_REVIEW-cloud-G03.md` | API-2 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
1. `test -f agent-task/m-agent-comparison-benchmark-pipeline/01_benchmark_manifest/complete.log && test -f agent-task/m-agent-comparison-benchmark-pipeline/02+01_isolated_workspace/complete.log && test -f agent-task/m-agent-comparison-benchmark-pipeline/03+01_run_lifecycle/complete.log && test -f agent-task/m-agent-comparison-benchmark-pipeline/04+01,02,03_repeat_attempt/complete.log`
|
||||
- Expected: exit 0 before implementation begins.
|
||||
2. `python3 -m unittest scripts.agent_benchmark.skill_contract_test`
|
||||
- Expected: frontmatter, routing, CLI command, safety, and capability-gate cases pass.
|
||||
3. `python3 scripts/agent_comparison_benchmark.py --help`
|
||||
- Expected: exit 0 and documented completed-pipeline subcommands are present.
|
||||
4. `make test-agent-comparison-benchmark`
|
||||
- Expected: all credential-free benchmark tests pass.
|
||||
5. `git diff --check`
|
||||
- Expected: exit 0 with no whitespace errors.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,169 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill plan=1 tag=API milestone-task=benchmark-skill -->
|
||||
|
||||
# Benchmark Operator Skill
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections of `CODE_REVIEW-cloud-G03.md` is the mandatory final implementation step. Before editing skill files, reread `agent-ops/skills/common/router.md` and follow the routed create-skill instructions. Run every verification command and paste actual output. Do not finalize, archive, write `complete.log`, ask the user, or change predecessor ownership.
|
||||
|
||||
## Background
|
||||
|
||||
Once the deterministic manifest/workspace/lifecycle/attempt pipeline exists, repository agents need a project-owned operational entrypoint. This packet documents and tests that entrypoint; executable product behavior remains in the completed CLI, not in prose or an orchestration dispatcher.
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/plan_local_G02_0.log` and `agent-task/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/code_review_cloud_G03_0.log`.
|
||||
- Review state: the prior pair was unimplemented and had no official verdict; it was archived by the requested Epic self-review replan.
|
||||
- Self-review defect: its predecessor checks named only active task paths, but PASS predecessors are normally moved under `agent-task/archive/YYYY/MM/`.
|
||||
- Scope carried forward: `benchmark-skill`, SDD S02, implementation files, and credential-free verification remain unchanged; only dependency resolution and paired evidence are corrected.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `AGENTS.md`
|
||||
- `agent-ops/rules/project/rules.md`
|
||||
- `agent-ops/rules/project/domain/testing/rules.md`
|
||||
- `agent-ops/skills/common/router.md`
|
||||
- `agent-ops/skills/common/create-skill/SKILL.md`
|
||||
- `agent-test/local/rules.md`
|
||||
- `agent-test/local/testing-smoke.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/agent-comparison-benchmark-pipeline.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md`
|
||||
- `agent-spec/input/openai-compatible-surface.md`
|
||||
- `agent-spec/runtime/provider-pool-config-refresh.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
- `agent-ops/skills/common/plan/SKILL.md`
|
||||
- `scripts/e2e-single-request-claude.sh`
|
||||
- `Makefile`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- Approved/released SDD scenario S02 requires a skill to accept a manifest path, validate required inputs/preconditions, and route execution into the benchmark pipeline.
|
||||
- Evidence Map row S02 directly becomes API-1's skill/preflight contract, API-2's skill-to-CLI parity test, and the contract/help Final Verification commands while preserving no-secret, fixed-testbed, fresh-session, bounded-lifecycle, and immutable-evidence policy.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- Local, credential-free documentation/contract tests only.
|
||||
- Handoff: no separate verification handoff was supplied; the approved SDD, repository create-skill workflow, local test rules, and final predecessor CLI are the verification sources.
|
||||
- The skill must invoke deterministic `scripts/agent_comparison_benchmark.py` commands and must not launch subagents, orchestration dispatchers, providers, browsers, or report generation itself.
|
||||
- Caller adapters and final report output remain later Epics. Their absence must be surfaced as unavailable capability, not hidden or worked around.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
No project skill or routing rule currently exposes manifest validation, preparation, run, resume, or status for this benchmark.
|
||||
|
||||
### Symbol References
|
||||
|
||||
Resolve exact CLI help, exit behavior, and artifact names from all four completed predecessors. Documentation and contract tests must match those implemented symbols verbatim.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
Creating a project skill changes the agent responsibility/routing boundary, so it is not direct-small even though its file count is small. Predecessors `01_benchmark_manifest`, `02+01_isolated_workspace`, `03+01_run_lifecycle`, and `04+01,02,03_repeat_attempt` are each missing their active or archived `complete.log` at planning time.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
Excluded: new executable pipeline logic, caller adapters, credential discovery, scoring, web validation, Markdown report rendering, roadmap mutation, dispatcher use, and Git commit/push.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- evaluation_mode: `isolated-reassessment`
|
||||
- finalizer: `finalize-task-policy.sh`, mode `pair`
|
||||
- build closures all `true`; scores `1/0/0/0/1`; route `local-fit`; grade `G02`; catalog `worker/local/G02`; filename `PLAN-local-G02.md`.
|
||||
- review closures all `true`; scores `1/0/0/1/1`; route `official-review`; grade `G03`; catalog `review/cloud/G03`; filename `CODE_REVIEW-cloud-G03.md`.
|
||||
- risks: `boundary_contract`, `variant_product`; large_indivisible_context `false`; rework `0`; evidence failure `false`; capability gap none.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Create the project benchmark skill with exact preflight, command, safety, and capability-gating behavior.
|
||||
- [ ] Route matching project requests to the skill and add deterministic skill/CLI contract tests.
|
||||
- [ ] Resolve all indexed predecessors from the allowed active/archive candidates, then run skill contract, CLI help, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Create the project benchmark skill
|
||||
|
||||
**Problem:** Agents have no supported operational entrypoint and could bypass validation, mutate the testbed, leak credentials, or claim unsupported report/adapter behavior.
|
||||
|
||||
**Solution:** Create `iop-agent-comparison-benchmark` through the repository create-skill workflow. Give it narrow triggers for benchmark manifest validate/prepare/run/resume/status requests. Require the manifest path, approved fixed dev testbed, clean source, output containment, and explicit capability preflight. Delegate all stateful work to the deterministic CLI. For caller execution or final report requests before those later-Epic capabilities exist, report the exact unavailable capability and stop; never simulate output or fall back to ad-hoc provider commands. Explicitly forbid raw secrets and orchestration dispatchers.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
No project skill routes benchmark pipeline operations.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```text
|
||||
iop-agent-comparison-benchmark validates preconditions and delegates supported operations to scripts/agent_comparison_benchmark.py.
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md` with valid frontmatter, triggers, preflight, supported commands, artifact expectations, safety, and stop conditions.
|
||||
- [ ] Update `agent-ops/rules/project/rules.md` with one explicit routing rule for benchmark pipeline operations.
|
||||
|
||||
**Test Strategy:** Parse frontmatter and required sections; assert every documented command exists in real CLI help and unsupported later-Epic capabilities are gated.
|
||||
|
||||
**Verification:** `python3 -m unittest scripts.agent_benchmark.skill_contract_test` exits 0.
|
||||
|
||||
### [API-2] Lock documentation to the executable surface
|
||||
|
||||
**Problem:** A skill can drift from its CLI and become an unsafe second implementation of pipeline policy.
|
||||
|
||||
**Solution:** Add a credential-free contract test that extracts skill command forms and compares them with CLI subcommands/options. Assert the project rule maps only the intended request family, no dispatcher command is referenced, no secret-bearing option is documented, and run/resume capability failure language is present until adapters arrive.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
Skill prose and executable CLI behavior can drift without detection.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
self.assert_skill_commands_match_cli_help()
|
||||
self.assert_later_epic_capabilities_fail_closed()
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_benchmark/skill_contract_test.py` covering frontmatter, routing, command surface, forbidden behavior, and capability gates.
|
||||
- [ ] Record exact verification output in `CODE_REVIEW-cloud-G03.md`.
|
||||
|
||||
**Test Strategy:** Inspect tracked text and invoke only `--help`; do not run any stateful benchmark subcommand.
|
||||
|
||||
**Verification:** Skill contract, CLI help, aggregate tests, and patch integrity pass.
|
||||
|
||||
## Dependencies and Execution Order
|
||||
|
||||
1. Required predecessor indices are `01_benchmark_manifest`, `02+01_isolated_workspace`, `03+01_run_lifecycle`, and `04+01,02,03_repeat_attempt`, encoded by the `05+01,02,03,04_...` task directory.
|
||||
2. At planning time none has an allowed active/archive `complete.log` candidate. Runtime scheduling must wait; completed predecessors may be under the active task group or `agent-task/archive/YYYY/MM/`.
|
||||
3. Before implementation, require exactly one matching completion record for each index across the allowed `NN_*/complete.log` and `NN+*/complete.log` locations. Missing or multiple matches fail closed and must not be guessed.
|
||||
4. Reread the router/create-skill instructions, inspect final CLI help, implement API-1, then API-2.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Items |
|
||||
|------|-------|
|
||||
| `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md` | API-1 |
|
||||
| `agent-ops/rules/project/rules.md` | API-1 |
|
||||
| `scripts/agent_benchmark/skill_contract_test.py` | API-2 |
|
||||
| `agent-task/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/CODE_REVIEW-cloud-G03.md` | API-2 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
1. `python3 -c 'from pathlib import Path; g="m-agent-comparison-benchmark-pipeline"; ids=("01","02","03","04"); a=Path("agent-task")/g; r=Path("agent-task/archive"); found={i:sorted([*a.glob(f"{i}_*/complete.log"),*a.glob(f"{i}+*/complete.log"),*r.glob(f"*/*/{g}/{i}_*/complete.log"),*r.glob(f"*/*/{g}/{i}+*/complete.log")],key=str) for i in ids}; bad={i:[str(p) for p in ps] for i,ps in found.items() if len(ps)!=1}; assert not bad,bad; print("\n".join(str(found[i][0]) for i in ids))'`
|
||||
- Expected: exit 0 and print exactly one active or archived completion path for each predecessor index before implementation begins.
|
||||
2. `python3 -m unittest scripts.agent_benchmark.skill_contract_test`
|
||||
- Expected: frontmatter, routing, CLI command, safety, and capability-gate cases pass.
|
||||
3. `python3 scripts/agent_comparison_benchmark.py --help`
|
||||
- Expected: exit 0 and documented completed-pipeline subcommands are present.
|
||||
4. `make test-agent-comparison-benchmark`
|
||||
- Expected: all credential-free benchmark tests pass.
|
||||
5. `git diff --check`
|
||||
- Expected: exit 0 with no whitespace errors.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
|
|
@ -0,0 +1,180 @@
|
|||
<!-- task=m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill plan=2 tag=API milestone-task=benchmark-skill -->
|
||||
|
||||
# Benchmark Operator Skill
|
||||
|
||||
## For the Implementing Agent
|
||||
|
||||
Filling the implementation-owned sections of `CODE_REVIEW-cloud-G03.md` is mandatory. Resolve predecessors `01` through `04`, reread the router and create-skill instructions, run their preflight/template flow, execute every verification command, paste actual output, and leave the active pair for official review. Do not finalize, archive, write `complete.log`, ask the user, or change predecessor ownership.
|
||||
|
||||
## Background
|
||||
|
||||
The deterministic pipeline needs a project-owned operator skill that recognizes supported benchmark requests and delegates to the CLI. The skill is routing/documentation, not a second implementation or orchestration dispatcher. Caller adapters and report rendering remain later Epics and must fail closed.
|
||||
|
||||
## Archive Evidence Snapshot
|
||||
|
||||
- Prior artifacts: `agent-task/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/plan_local_G02_1.log` and `agent-task/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/code_review_cloud_G03_1.log` (generation 1 retains generation 0 history).
|
||||
- Review state: unimplemented, no official verdict, replaced through explicit plan `write` mode.
|
||||
- Self-review defects: the prior plan did not explicitly route report-intent requests to a stable fail-closed response and did not encode the create-skill preflight/template checks as implementation requirements.
|
||||
- Scope carried forward: `benchmark-skill`, S02, project routing, CLI parity tests, predecessor resolution, and credential-free verification.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Files Read
|
||||
|
||||
- `AGENTS.md`
|
||||
- `agent-ops/rules/project/rules.md`
|
||||
- `agent-ops/rules/project/domain/testing/rules.md`
|
||||
- `agent-ops/rules/common/rules-roadmap.md`
|
||||
- `agent-ops/skills/common/router.md`
|
||||
- `agent-ops/skills/common/create-skill/SKILL.md`
|
||||
- `agent-ops/skills/common/create-skill/templates/SKILL-template.md`
|
||||
- `agent-ops/skills/common/plan/SKILL.md`
|
||||
- `agent-ops/skills/common/finalize-task-routing/SKILL.md`
|
||||
- `agent-test/local/rules.md`
|
||||
- `agent-test/local/testing-smoke.md`
|
||||
- `agent-roadmap/current.md`
|
||||
- `agent-roadmap/priority-queue.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/PHASE.md`
|
||||
- `agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/agent-comparison-benchmark-pipeline.md`
|
||||
- `agent-roadmap/sdd/knowledge-tool-optimization-extension/agent-comparison-benchmark-pipeline/SDD.md`
|
||||
- `agent-spec/index.md`
|
||||
- `agent-spec/input/openai-compatible-surface.md`
|
||||
- `agent-spec/runtime/provider-pool-config-refresh.md`
|
||||
- `agent-contract/index.md`
|
||||
- `agent-contract/outer/anthropic-compatible-api.md`
|
||||
- `agent-contract/outer/openai-compatible-api.md`
|
||||
- `agent-contract/inner/edge-config-runtime-refresh.md`
|
||||
- `Makefile`
|
||||
- `.gitignore`
|
||||
|
||||
### SDD Criteria
|
||||
|
||||
- S02 requires a skill to accept a manifest path, validate preconditions, and route supported execution into the benchmark pipeline.
|
||||
- The SDD separates caller adapters and report rendering into later Epics; this skill may recognize those intents but cannot simulate them.
|
||||
- Evidence Map S02 requires an executable skill/preflight contract and skill-to-CLI parity evidence.
|
||||
|
||||
### Verification Context
|
||||
|
||||
- Local, credential-free documentation and contract tests only.
|
||||
- `.agent-ops-source` is absent, so the owned destination is `agent-ops/skills/project/` and the routing rule belongs in `agent-ops/rules/project/rules.md`.
|
||||
- No duplicate `iop-agent-comparison-benchmark` skill exists. Use the create-skill template/frontmatter exactly, then validate it.
|
||||
- Predecessors `01` through `04` have no completion records; implementation must wait.
|
||||
|
||||
### Test Coverage Gaps
|
||||
|
||||
No project skill or rule recognizes benchmark validate/prepare/run/resume/status/report intent, enforces exact capability gates, or proves its documented commands match the CLI.
|
||||
|
||||
### Symbol References
|
||||
|
||||
Resolve final CLI help, exit codes, artifact names, and capability strings from completed predecessors. The skill must not invent options or implement state transitions in prose.
|
||||
|
||||
### Split Judgment
|
||||
|
||||
The skill and its project routing rule are one agent-responsibility boundary. Contract tests must land with them so documentation cannot drift from the executable surface.
|
||||
|
||||
### Scope Rationale
|
||||
|
||||
Excluded: executable pipeline changes, caller adapters, credential discovery, scoring, web validation, Markdown report generation, subagents, dispatchers, roadmap mutation, and Git operations.
|
||||
|
||||
### Final Routing
|
||||
|
||||
- evaluation_mode: `isolated-reassessment`
|
||||
- finalizer: `finalize-task-policy.sh`, mode `pair`
|
||||
- build closures all `true`; scores `1/0/0/0/1`; base/route `local-fit`; grade `G02`; catalog `worker/local/G02`; filename `PLAN-local-G02.md`.
|
||||
- review closures all `true`; scores `1/0/0/1/1`; route `official-review`; grade `G03`; catalog `review/cloud/G03`; filename `CODE_REVIEW-cloud-G03.md`.
|
||||
- large_indivisible_context `false`; risks `boundary_contract`, `variant_product`; rework `0`; evidence integrity failure `false`; capability gap none.
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
- [ ] Run create-skill preflight/template validation and create the project benchmark skill with exact request triggers, CLI delegation, safety rules, and capability gates.
|
||||
- [ ] Route the benchmark request family in project rules and add deterministic skill/frontmatter/routing/CLI-help contract tests.
|
||||
- [ ] Resolve predecessors `01` through `04`, then run skill contract, CLI help, aggregate, and patch-integrity verification.
|
||||
- [ ] Fill implementation-owned sections in CODE_REVIEW-*-G??.md with actual implementation notes and verification output.
|
||||
|
||||
### [API-1] Create the project benchmark operator skill
|
||||
|
||||
**Problem:** Agents have no supported operational entrypoint and could bypass validation, mutate the fixed testbed, leak credentials, or claim unavailable adapter/report behavior.
|
||||
|
||||
**Solution:** Follow the create-skill preflight and template to add `iop-agent-comparison-benchmark` under the project skill root. Give it explicit triggers for benchmark manifest validation, workspace preparation, run, resume, status, and report-readiness requests. Require manifest path, fixed clean `../iop-s2` provenance, contained output, fresh-session/cache policy, and CLI capability preflight. Delegate every supported stateful operation verbatim to `scripts/agent_comparison_benchmark.py`; do not reproduce pipeline policy in prose.
|
||||
|
||||
Until later-Epic commands exist, caller execution must return `capability-unavailable: caller-adapter` and report/output requests must return `capability-unavailable: report-output`. Stop without fallback, fabricated evidence, ad-hoc provider calls, subagents, or orchestration dispatchers.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
No project skill routes benchmark pipeline operations or capability gates.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```text
|
||||
iop-agent-comparison-benchmark delegates supported CLI operations and fails closed for caller-adapter/report-output.
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md` from the repository template with valid frontmatter, triggers, preflight, supported commands, artifacts, safety, and stop conditions.
|
||||
- [ ] Update `agent-ops/rules/project/rules.md` with one narrow routing rule covering the benchmark pipeline request family, including report readiness.
|
||||
|
||||
**Test Strategy:** Parse frontmatter and required sections; assert the skill contains the exact two capability strings, no dispatcher/provider fallback, and only commands exposed by real CLI help.
|
||||
|
||||
**Verification:** `python3 -m unittest scripts.agent_benchmark.skill_contract_test` exits 0.
|
||||
|
||||
### [API-2] Lock the skill to the executable surface
|
||||
|
||||
**Problem:** Skill prose can drift into a second, unsafe implementation or advertise later-Epic commands before they exist.
|
||||
|
||||
**Solution:** Add a credential-free contract test that checks the create-skill template/frontmatter invariants, project rule routing, documented command forms against `--help`, supported/unsupported capability matrix, and forbidden dispatcher/secret/fallback language. Invoke only help and tracked text; never run a stateful benchmark command.
|
||||
|
||||
Before:
|
||||
|
||||
```text
|
||||
No executable check binds benchmark skill instructions to the CLI.
|
||||
```
|
||||
|
||||
After:
|
||||
|
||||
```python
|
||||
self.assert_skill_commands_match_cli_help()
|
||||
self.assert_capability("report-output", available=False)
|
||||
```
|
||||
|
||||
**Modified Files and Checklist:**
|
||||
|
||||
- [ ] Add `scripts/agent_benchmark/skill_contract_test.py` covering template/frontmatter, routing, CLI parity, capability gates, and forbidden behavior.
|
||||
- [ ] Record exact output in `CODE_REVIEW-cloud-G03.md`.
|
||||
|
||||
**Test Strategy:** Read tracked files and invoke only `python3 scripts/agent_comparison_benchmark.py --help`; no workspace, provider, browser, or process lifecycle is started.
|
||||
|
||||
**Verification:** Skill contract, CLI help, aggregate tests, and patch integrity pass.
|
||||
|
||||
## Dependencies and Execution Order
|
||||
|
||||
1. Required predecessors `01`, `02`, `03`, and `04` are encoded by `05+01,02,03,04_...`.
|
||||
2. At planning time none has exactly one valid active/archive `complete.log`; runtime must wait.
|
||||
3. Before implementation, require exactly one allowed completion for every predecessor. Missing or multiple matches fail closed.
|
||||
4. Reread router/create-skill, rerun preflight, inspect completed CLI help, create from template, add routing/tests, then verify.
|
||||
|
||||
## Modified Files Summary
|
||||
|
||||
| File | Items |
|
||||
|------|-------|
|
||||
| `agent-ops/skills/project/iop-agent-comparison-benchmark/SKILL.md` | API-1 |
|
||||
| `agent-ops/rules/project/rules.md` | API-1 |
|
||||
| `scripts/agent_benchmark/skill_contract_test.py` | API-2 |
|
||||
| `agent-task/m-agent-comparison-benchmark-pipeline/05+01,02,03,04_benchmark_skill/CODE_REVIEW-cloud-G03.md` | API-2 |
|
||||
|
||||
## Final Verification
|
||||
|
||||
1. `python3 -c 'from pathlib import Path; g="m-agent-comparison-benchmark-pipeline"; ids=("01","02","03","04"); a=Path("agent-task")/g; r=Path("agent-task/archive"); found={i:sorted([*a.glob(f"{i}_*/complete.log"),*a.glob(f"{i}+*/complete.log"),*r.glob(f"*/*/{g}/{i}_*/complete.log"),*r.glob(f"*/*/{g}/{i}+*/complete.log")],key=str) for i in ids}; bad={i:[str(p) for p in ps] for i,ps in found.items() if len(ps)!=1}; assert not bad,bad; print("\n".join(str(found[i][0]) for i in ids))'`
|
||||
- Expected: exactly one allowed completion path for each predecessor before implementation.
|
||||
2. `python3 -m unittest scripts.agent_benchmark.skill_contract_test`
|
||||
- Expected: template/frontmatter, routing, CLI parity, safety, and capability-gate cases pass.
|
||||
3. `python3 scripts/agent_comparison_benchmark.py --help`
|
||||
- Expected: exit 0 and only completed-pipeline subcommands are documented by the skill.
|
||||
4. `make test-agent-comparison-benchmark`
|
||||
- Expected: all credential-free benchmark tests pass.
|
||||
5. `git diff --check`
|
||||
- Expected: exit 0 with no whitespace errors.
|
||||
|
||||
After completing all code changes, fill implementation-owned sections in `CODE_REVIEW-*-G??.md`.
|
||||
Loading…
Reference in a new issue