iop/scripts/fixtures/agent-comparison-benchmark-report.expected.md
toki 8f00606c03 fix(benchmark): 결과 경계를 독립 축으로 분리한다
제품 결과와 harness·process·artifact 실패가 하나의 성공 값으로 덮이지 않도록 durable evidence와 모든 소비자 계약을 함께 마이그레이션한다.
2026-08-12 21:01:51 +09:00

8.4 KiB

Agent comparison benchmark report

Run identity

field value
run_id run-20260811T010203Z-123456abcdef
manifest_digest sha256:39715007db41882abcfe5a3fd5f8e3cc4bcadce0ef31c17df4f3347825184359
pipeline_version 2

Immutable conditions

field value
environment dev
fixture landing-v1 (sha256:4b5c9dcfe799d21f86a4462a2ada35275d72d5fe919b2adb069eeb3f4ac72fcb)
rubric landing-quality-v1
session_policy fresh
setup_cache_policy isolated
evaluator codex/judge-model/xhigh

Execution preflight

sequence status results
unavailable 0

Attempt outcomes

cell repetition attempt controller product harness process artifact scoring total rank
cell-sentinel 1 1 completed succeeded passed exited passed scored 99 1
cell-sentinel 1 2 completed succeeded passed exited passed scored 99 1
cell-sentinel 1 3 completed failed passed exited failed unscored
cell-sentinel 1 4 completed succeeded passed exited passed scoring_failed
cell-sentinel 1 5 completed succeeded passed exited passed blocked

Quality score breakdown

cell/repetition/attempt category score max
cell-sentinel/r1/a1 task_fidelity 24 25
cell-sentinel/r1/a1 visual_hierarchy 25 25
cell-sentinel/r1/a1 responsive_composition 20 20
cell-sentinel/r1/a1 typography_readability 15 15
cell-sentinel/r1/a1 polish_consistency 15 15
cell-sentinel/r1/a2 task_fidelity 24 25
cell-sentinel/r1/a2 visual_hierarchy 25 25
cell-sentinel/r1/a2 responsive_composition 20 20
cell-sentinel/r1/a2 typography_readability 15 15
cell-sentinel/r1/a2 polish_consistency 15 15

Timing and token evidence

cell/repetition/attempt time observations token observations
cell-sentinel/r1/a1 total_duration=1 ns; clock=harness_monotonic; source=harness; submitted_at,first_output_at=unavailable; reason=not_observed; source=harness; first_write_observed_at,first_write_mtime=unavailable; reason=not_observed; source=workspace_poll cache_write_tokens,cached_input_tokens,input_tokens,model_calls,model_duration,output_tokens,queue_duration,reasoning_tokens,tool_calls,tool_duration,total_duration,total_tokens=unavailable; reason=not_reported; source=harness
cell-sentinel/r1/a2 total_duration=1 ns; clock=harness_monotonic; source=harness; submitted_at,first_output_at=unavailable; reason=not_observed; source=harness; first_write_observed_at,first_write_mtime=unavailable; reason=not_observed; source=workspace_poll cache_write_tokens,cached_input_tokens,input_tokens,model_calls,model_duration,output_tokens,queue_duration,reasoning_tokens,tool_calls,tool_duration,total_duration,total_tokens=unavailable; reason=not_reported; source=harness
cell-sentinel/r1/a3 total_duration=1 ns; clock=harness_monotonic; source=harness; submitted_at,first_output_at=unavailable; reason=not_observed; source=harness; first_write_observed_at,first_write_mtime=unavailable; reason=not_observed; source=workspace_poll cache_write_tokens,cached_input_tokens,input_tokens,model_calls,model_duration,output_tokens,queue_duration,reasoning_tokens,tool_calls,tool_duration,total_duration,total_tokens=unavailable; reason=not_reported; source=harness
cell-sentinel/r1/a4 total_duration=1 ns; clock=harness_monotonic; source=harness; submitted_at,first_output_at=unavailable; reason=not_observed; source=harness; first_write_observed_at,first_write_mtime=unavailable; reason=not_observed; source=workspace_poll cache_write_tokens,cached_input_tokens,input_tokens,model_calls,model_duration,output_tokens,queue_duration,reasoning_tokens,tool_calls,tool_duration,total_duration,total_tokens=unavailable; reason=not_reported; source=harness
cell-sentinel/r1/a5 total_duration=1 ns; clock=harness_monotonic; source=harness; submitted_at,first_output_at=unavailable; reason=not_observed; source=harness; first_write_observed_at,first_write_mtime=unavailable; reason=not_observed; source=workspace_poll cache_write_tokens,cached_input_tokens,input_tokens,model_calls,model_duration,output_tokens,queue_duration,reasoning_tokens,tool_calls,tool_duration,total_duration,total_tokens=unavailable; reason=not_reported; source=harness

Web validation and scoring provenance

cell/repetition/attempt web gates screenshots score_id evaluator scoring condition
cell-sentinel/r1/a1 generated_files=pass, static_safety=pass, images=pass, network=pass, console=pass, responsive=pass, accessibility=pass screenshot-desktop.png, screenshot-mobile.png score-000001 codex/judge-route/judge-model/xhigh recorded
cell-sentinel/r1/a2 generated_files=pass, static_safety=pass, images=pass, network=pass, console=pass, responsive=pass, accessibility=pass screenshot-desktop.png, screenshot-mobile.png score-000001 codex/judge-route/judge-model/xhigh recorded
cell-sentinel/r1/a3 generated_files=pass, static_safety=pass, images=fail, network=fail, console=fail, responsive=fail, accessibility=fail unavailable unavailable product_failed, process_nonzero_exit, web_failed, web_reason_render_not_run, gate_images, gate_network, gate_console, gate_responsive, gate_accessibility
cell-sentinel/r1/a4 generated_files=pass, static_safety=pass, images=pass, network=pass, console=pass, responsive=pass, accessibility=pass screenshot-desktop.png, screenshot-mobile.png score-000001 codex/judge-route/judge-model/xhigh invalid_worksheet
cell-sentinel/r1/a5 generated_files=pass, static_safety=pass, images=pass, network=pass, console=pass, responsive=pass, accessibility=pass screenshot-desktop.png, screenshot-mobile.png unavailable evaluator_preflight_blocked

Limitations

  • Values marked unavailable retain the producing source and reason; they are not inferred as zero.
  • Automatic web gates establish eligibility only and contribute no quality points.
  • Equal scored totals share a competition rank; unscored and scoring-failed attempts do not receive a rank.

Raw evidence index

contained pointer
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw
raw