iop/scripts/fixtures/agent-comparison-benchmark-report.expected.md
toki 029ff0d2c8 feat(benchmark): 비교 파이프라인을 완성한다
동일한 IOP 경유 과업을 caller와 model 설정만 바꿔 재현하고, 실패를 포함한 실행·검증·채점 근거를 보존할 수 있어야 한다.
2026-08-12 01:44:26 +09:00

8.2 KiB

Agent comparison benchmark report

Run identity

field value
run_id run-20260811T010203Z-123456abcdef
manifest_digest sha256:39715007db41882abcfe5a3fd5f8e3cc4bcadce0ef31c17df4f3347825184359
pipeline_version 2

Immutable conditions

field value
environment dev
fixture landing-v1 (sha256:4b5c9dcfe799d21f86a4462a2ada35275d72d5fe919b2adb069eeb3f4ac72fcb)
rubric landing-quality-v1
session_policy fresh
setup_cache_policy isolated
evaluator codex/judge-model/xhigh

Execution preflight

sequence status results
unavailable 0

Attempt outcomes

cell repetition attempt execution terminal web scoring total rank
cell-sentinel 1 1 success success passed scored 99 1
cell-sentinel 1 2 success success passed scored 99 1
cell-sentinel 1 3 failed failed not_run unscored
cell-sentinel 1 4 success success passed scoring_failed
cell-sentinel 1 5 success success passed blocked

Quality score breakdown

cell/repetition/attempt category score max
cell-sentinel/r1/a1 task_fidelity 24 25
cell-sentinel/r1/a1 visual_hierarchy 25 25
cell-sentinel/r1/a1 responsive_composition 20 20
cell-sentinel/r1/a1 typography_readability 15 15
cell-sentinel/r1/a1 polish_consistency 15 15
cell-sentinel/r1/a2 task_fidelity 24 25
cell-sentinel/r1/a2 visual_hierarchy 25 25
cell-sentinel/r1/a2 responsive_composition 20 20
cell-sentinel/r1/a2 typography_readability 15 15
cell-sentinel/r1/a2 polish_consistency 15 15

Timing and token evidence

cell/repetition/attempt time observations token observations
cell-sentinel/r1/a1 total_duration=1 ns; clock=harness_monotonic; source=harness; submitted_at,first_output_at=unavailable; reason=not_observed; source=harness; first_write_observed_at,first_write_mtime=unavailable; reason=not_observed; source=workspace_poll cache_write_tokens,cached_input_tokens,input_tokens,model_calls,model_duration,output_tokens,queue_duration,reasoning_tokens,tool_calls,tool_duration,total_duration,total_tokens=unavailable; reason=not_reported; source=harness
cell-sentinel/r1/a2 total_duration=1 ns; clock=harness_monotonic; source=harness; submitted_at,first_output_at=unavailable; reason=not_observed; source=harness; first_write_observed_at,first_write_mtime=unavailable; reason=not_observed; source=workspace_poll cache_write_tokens,cached_input_tokens,input_tokens,model_calls,model_duration,output_tokens,queue_duration,reasoning_tokens,tool_calls,tool_duration,total_duration,total_tokens=unavailable; reason=not_reported; source=harness
cell-sentinel/r1/a3 total_duration=1 ns; clock=harness_monotonic; source=harness; submitted_at,first_output_at=unavailable; reason=not_observed; source=harness; first_write_observed_at,first_write_mtime=unavailable; reason=not_observed; source=workspace_poll cache_write_tokens,cached_input_tokens,input_tokens,model_calls,model_duration,output_tokens,queue_duration,reasoning_tokens,tool_calls,tool_duration,total_duration,total_tokens=unavailable; reason=not_reported; source=harness
cell-sentinel/r1/a4 total_duration=1 ns; clock=harness_monotonic; source=harness; submitted_at,first_output_at=unavailable; reason=not_observed; source=harness; first_write_observed_at,first_write_mtime=unavailable; reason=not_observed; source=workspace_poll cache_write_tokens,cached_input_tokens,input_tokens,model_calls,model_duration,output_tokens,queue_duration,reasoning_tokens,tool_calls,tool_duration,total_duration,total_tokens=unavailable; reason=not_reported; source=harness
cell-sentinel/r1/a5 total_duration=1 ns; clock=harness_monotonic; source=harness; submitted_at,first_output_at=unavailable; reason=not_observed; source=harness; first_write_observed_at,first_write_mtime=unavailable; reason=not_observed; source=workspace_poll cache_write_tokens,cached_input_tokens,input_tokens,model_calls,model_duration,output_tokens,queue_duration,reasoning_tokens,tool_calls,tool_duration,total_duration,total_tokens=unavailable; reason=not_reported; source=harness

Web validation and scoring provenance

cell/repetition/attempt web gates screenshots score_id evaluator scoring condition
cell-sentinel/r1/a1 generated_files=pass, static_safety=pass, images=pass, network=pass, console=pass, responsive=pass, accessibility=pass screenshot-desktop.png, screenshot-mobile.png score-000001 codex/judge-route/judge-model/xhigh recorded
cell-sentinel/r1/a2 generated_files=pass, static_safety=pass, images=pass, network=pass, console=pass, responsive=pass, accessibility=pass screenshot-desktop.png, screenshot-mobile.png score-000001 codex/judge-route/judge-model/xhigh recorded
cell-sentinel/r1/a3 generated_files=fail, static_safety=fail, images=fail, network=fail, console=fail, responsive=fail, accessibility=fail unavailable unavailable lifecycle_failed
cell-sentinel/r1/a4 generated_files=pass, static_safety=pass, images=pass, network=pass, console=pass, responsive=pass, accessibility=pass screenshot-desktop.png, screenshot-mobile.png score-000001 codex/judge-route/judge-model/xhigh invalid_worksheet
cell-sentinel/r1/a5 generated_files=pass, static_safety=pass, images=pass, network=pass, console=pass, responsive=pass, accessibility=pass screenshot-desktop.png, screenshot-mobile.png unavailable evaluator_preflight_blocked

Limitations

  • Values marked unavailable retain the producing source and reason; they are not inferred as zero.
  • Automatic web gates establish eligibility only and contribute no quality points.
  • Equal scored totals share a competition rank; unscored and scoring-failed attempts do not receive a rank.

Raw evidence index

contained pointer
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw