iop/agent-contract/inner/execution-runtime.md
toki f9442edfef feat(runtime): provider liveness 복구를 완성한다
장시간 무응답 attempt를 안전하게 fence하고 provider health와 분리 관측해야 중복 출력 없이 기존 recovery budget으로 재실행할 수 있다.
2026-08-06 08:49:59 +09:00

18 KiB

Provider Execution Runtime Contract

Contract metadata

  • id: iop.execution-runtime
  • boundary: inner
  • status: active
  • source evidence:
    • packages/go/execution/types.go
    • packages/go/execution/liveness.go
    • packages/go/execution/registry.go
    • packages/go/execution/emitter.go
    • packages/go/execution/failure.go
    • apps/node/internal/node/runtime_bridge.go
    • apps/node/internal/node/health_probe.go
    • apps/node/internal/node/command_handler.go
    • apps/node/internal/node/liveness_watchdog.go
    • apps/node/internal/transport/session.go
    • apps/edge/internal/service/model_queue_release.go
    • apps/edge/internal/service/node_command.go
    • apps/edge/internal/openai/stream_gate_runtime.go
    • apps/edge/internal/openai/stream_gate_stall_recovery_test.go

Scope

The execution package defines host-neutral provider primitives. It owns provider registration, lifecycle, execution events, typed failures, cancellation, token usage, and the three provider commands capabilities, transport_status, and ollama_api.

session_id is an opaque correlation value. It does not select, create, resume, or terminate a process. Repeated requests with the same value are independent executions. Cancellation targets a non-empty run_id only.

Requirements

  • Providers implement the narrow Provider interface and may expose optional lifecycle, command, or tunnel capabilities.
  • Event emitters preserve order and publish exactly one terminal event.
  • Registry lookup uses provider identity and returns typed failures for missing or unavailable providers.
  • Callers must reject commands outside the closed provider-command allowlist before provider lookup.
  • Token usage remains observation data attached to execution or tunnel results.
  • DefaultResponseStallTimeoutMS = 300000 is the documented default. ResolveStallTimeoutMS(ms) validates then maps zero to the default; safe positive values pass through, while negative or overflow values return an error.
  • ClassifyRuntimeEvent returns start for EventTypeStart, progress for non-empty delta/message or non-terminal usage, terminal for complete/error/cancelled (before usage check), and none for empty/unknown events.
  • ClassifyProviderTunnelFrame returns progress for response_start (with or without headers) and non-empty body, terminal for end/error (before payload check), progress for usage, and none for empty/unknown frames.
  • ValidateStallTimeoutMS(ms) rejects negative values and values exceeding maxSafeStallTimeoutMS; zero is allowed (use default).
  • NodeProviderConf.EffectiveResponseStallTimeoutMS() returns the effective timeout for a provider candidate.
  • RunRequest.ResponseStallTimeoutMS and ProviderTunnelRequest.ResponseStallTimeoutMS carry the selected provider's effective timeout; zero on the wire means the Node applies the documented default.
  • The Node wire boundary normalizes zero to 300000 and rejects negative or overflow values before router/provider invocation.
  • response_stalled is a stable typed failure. Node transport mappers (runEventToProto and tunnelFrameToProto) populate the optional wire ExecutionFailure message only for FailureCodeResponseStalled, attaching a defensive clone of allowlisted metadata keys (failure_code, provider_health, liveness_classification, idle_duration_ms, run_id, attempt_id, attempt_fence, adapter, target, and health_observation_seq); nil and non-stalled failures leave wire ExecutionFailure absent while preserving legacy error string fields (RunEvent.Error / ProviderTunnelFrame.Error). Caller metadata cannot override these values, and no raw payload, credential, or recovery_eligible signal is admitted.
  • The Node watchdog starts from attempt admission, resets only on the documented progress dispositions, stops on provider terminal, and emits one typed stall terminal. It does not retry providers or infer recovery eligibility. Retryable=true means only that the local provider ownership fence was confirmed within the bounded close grace.
  • After the watchdog claims a stall it joins two independent bounded outcomes without extending either serially — the fixed close-grace fence and the exact-target health probe — then assembles exactly one allowlisted terminal. The joined liveness_classification/provider_health pair is exactly request_stalled/available, provider_unhealthy/unavailable, or health_unknown/unknown (fail-closed default). Provider availability observed here is evidence only: it never resets progress, changes the fence, revives output, or authorizes retry, and late provider output stays fenced.
  • health_observation_seq is a connection-scoped monotonic sequence sourced from the transport Session. A new connection starts at zero, so the first finalized observation is one; normalized and tunnel observations on the same connection share the source and receive unique, increasing values under concurrency. Internal or unbound execution paths omit the key entirely and never encode a process-global generation.
  • ProviderPoolDispatchRequest carries two request-local recovery-hint fields: AvoidProviderID (non-empty to prefer a runtime-eligible alternate over the avoided provider) and AllowAvoidedProviderFallback (explicit permission to retain the avoided provider when no alternate exists and it remains runtime eligible). The queue applies identical avoidance filtering to both initial and queued re-resolution. Zero values preserve current selection behavior. This is selection policy only: it does not create a retry loop, reserve a slot, change provider priority, persist the hints, or count retries. The fallback permission is always derived from exact probe-backed available evidence by the caller (never from current overlay state).
  • A Node capabilities command performs the same bounded exact-target ProbeHealth operation. Its stable result evidence is the requested adapter instance key (adapter_key), exact target, fail-closed normalized provider_status, and the next health_observation_seq from that same transport Session. Probe errors, unsupported probing, and adapter/instance/target mismatches report unknown; raw capability status is not recovery evidence.
  • Edge accepts a typed stall observation for provider-wide projection only after authoritative reception (node_id, connection_generation) matches the tracked immutable dispatch lease (node_id, connection_generation, provider_id, adapter, target), the local attempt fence is confirmed, and the observation sequence is strictly newer. A current terminal still releases its lease exactly once when health evidence is absent, malformed, mismatched, or stale; a reception-owner mismatch changes neither overlay nor lease state.
  • Every validated current bound stall is annotated with Edge-owned provider_id, the validated provider_health, and recovery_handoff=confirmed, including an out-of-order terminal whose health projection is sequence-stale. Only a fresh unavailable observation lowers the generation-scoped runtime overlay. The token proves reception, lease binding, and local-fence handoff only; it is never recovery_eligible and never authorizes retry.
  • Every supported OpenAI Chat/Responses normalized or tunnel request enters one request-local StreamGate runtime, which is the sole liveness owner even when configured semantic filtering is disabled. That runtime may consume the confirmed handoff as a raw-free response_stalled provider error while its endpoint adapters preserve the disabled-semantic native status, headers, JSON/SSE/tunnel order, validation, usage, cancellation, and terminal behavior. It retains only the stable failure code, confirmed-handoff token, and available|unavailable|unknown health classification; Node/provider messages and arbitrary metadata are not copied. Exact replay additionally requires the existing uncommitted, uncancelled, side-effect-safe, snapshot-backed, shared-budget gate. A confirmed old terminal closes its Edge transport without another CancelRun; pool re-admission consumes the provider once as AvoidProviderID, with same-provider fallback only for exact available evidence.
  • The runtime overlay is keyed by (node_id, connection_generation, provider_id) and remains separate from configuration health. It excludes the provider from effective admission and projects it unavailable in status snapshots. Recovery requires a later CAPABILITIES result for the same current adapter/target mapping with strictly higher sequence and exact normalized available; malformed, ambiguous, stale-generation, unknown, and unavailable results are no-ops.

Health probe contract

The execution package owns the stable, fail-closed probe outcome vocabulary consumed by Node terminal assembly. It is the typed three-way boundary between an inconclusive probe and a definitive provider-health classification; nothing else maps provider probe results to health.

  • ProviderHealth is the stable normalized value: request_stalled, provider_unhealthy, or health_unknown (fail-closed default).
  • LivenessClassification is the stable observable category a probe outcome reduces through: available, unavailable, timeout, error, unsupported, unknown, and identity_mismatch.
  • ProbeOutcome is the typed, target-aware input; ClassifyProbeOutcome reduces it to a classification and NormalizeProbeOutcome maps it to health. The mapping is exactly: available → request_stalled; a validated matching unavailable result → provider_unhealthy; every error, timeout, unsupported adapter, unknown status, empty/mismatched adapter or target, and instance mismatch → health_unknown.
  • A returned error takes precedence over any reported status, so endpoint construction, request/network, non-success HTTP, and decode failures can never be confused with a positive exact-target-absent result.
  • The Node probe coordinator (ProbeHealth) roots its own five-second bounded context from the background, re-checks that deadline/cancellation after the probe returns, validates exact adapter and target identity (including a pinned instance key when set), and feeds only the typed normalizer. It never copies arbitrary provider metadata.
  • ResolveProbeFunc returns nil for an adapter that does not implement ProviderProber; a nil hook makes ProbeHealth fail closed to health_unknown without invoking any endpoint.

Probe completion is evidence only. The probe itself must never reset original request progress, change the attempt fence, authorize retry, sequence the watchdog terminal, directly mutate the Edge overlay, or infer recovery. Node owns the stall-terminal join and the connection-scoped health_observation_seq. Edge owns reception-generation and immutable-lease validation, the separate runtime overlay, candidate exclusion, snapshot projection, and exact later CAPABILITIES recovery. The ingress recovery host remains the sole owner of commit, cancellation, side-effect, budget, candidate, and replay eligibility decisions.

Prohibited ownership

The package must not own interactive shells, persistent processes, terminal emulation, working-directory mutation, resumable conversations, local quota probing, or arbitrary host command execution. It must not import application-internal packages or generated transport types.

Operational evidence projections

The Node and Edge owners expose bounded operational projections derived exclusively from the established stall terminal, health-overlay, and recovery decisions documented above. These projections never widen the Node↔Edge wire protocol: they carry no new frame, field, ordering rule, or retry semantic, and they are emitted only after the authoritative decision is finalized.

Node stall observations (owner: Node process-global)

  • iop_node_response_stalls_total (counter): labels execution_path, provider_health, liveness_classification, attempt_fence. Every claimed stall increments exactly one series.
  • iop_node_response_stall_duration_seconds (histogram): same four labels. Samples the idle duration in seconds.
  • Dedicated structured log node_response_stall_observation: fields execution_path, provider_health, liveness_classification, attempt_fence, idle_duration_ms.
  • Label values are closed and low-cardinality: execution_path ∈ {normalized, provider_tunnel, unknown}; provider_health ∈ {available, unavailable, unknown}; liveness_classification ∈ {request_stalled, provider_unhealthy, health_unknown}; attempt_fence ∈ {confirmed, unconfirmed, unknown}.
  • Prohibited from metric labels and general logs: raw prompt/response, credential, caller metadata, recovery_eligible. High-cardinality inputs normalize to unknown.
  • Observer failure is fire-and-forget and never suppresses the terminal.
  • Source: apps/node/internal/node/liveness_observability.go; test: apps/node/internal/node/liveness_observability_test.go::TestNodeLivenessObservability.

Edge provider-health overlay observations (owner: Edge service queue process-global)

  • iop_edge_provider_health_evidence_total (counter): labels source, evidence_health, decision. Records authoritative overlay decisions.
  • iop_edge_provider_health_transitions_total (counter): labels from_health, to_health. Records overlay state transitions.
  • Dedicated structured log edge_provider_health_observation: fields source, evidence_health, decision, from_health, to_health, state_changed.
  • Label values are closed: source ∈ {stall, probe, unknown}; evidence_health ∈ {available, unavailable, unknown}; decision ∈ {applied, rejected_stale, rejected_binding, rejected_ambiguous, inconclusive}; from_health/to_health ∈ {available, unavailable, unknown}.
  • Prohibited from metric labels and general logs: provider, node, run, session, adapter, target, payload, or credential values.
  • Edge delivery is synchronous after decision/release/pump and after the queue lock is released; observer latency can delay handler return but cannot retain the lock or change the finalized transition.
  • Source: apps/edge/internal/service/provider_health_observability.go; test: apps/edge/internal/service/provider_health_observability_test.go::TestProviderHealthObservability and TestProviderHealthObservabilityDoesNotExposeSentinels.

Edge OpenAI recovery observations (owner: Edge OpenAI server request-local wrapper with process-global collectors)

  • iop_edge_liveness_recovery_eligibility_total (counter): labels execution_path, provider_health, commit_state, eligibility. Records eligibility decisions per liveness cycle.
  • iop_edge_liveness_recovery_results_total (counter): labels execution_path, provider_health, recovery_result. Records at most one final result per liveness cycle.
  • Dedicated structured log edge_liveness_recovery_observation: fields phase, execution_path, provider_health, commit_state, eligibility, recovery_result.
  • Label values are closed: execution_path ∈ {normalized, provider_tunnel, unknown}; provider_health ∈ {available, unavailable, unknown}; commit_state ∈ {transport_uncommitted, stream_open, terminal_committed, unknown}; eligibility ∈ {eligible, no_owner, post_commit, unconfirmed_fence, caller_cancelled, tool_side_effect, budget_exhausted, no_candidate, same_provider_forbidden, other}; recovery_result ∈ {redispatched, plan_rejected, abort_failed, rebuild_failed, dispatch_failed, not_selected, terminal, other}.
  • Prohibited from metric labels and general logs: correlation, attempt, run, session, model, provider, node, plan, shared_attempt_id, credential, or slot identifiers.
  • Each request owns one fresh wrapper; the collectors are process-global and registered once at package init.
  • phase is the bounded request-local cycle phase: idle before any eligible observation, eligible_pending after an eligible eligibility decision until the cycle resolves (redispatched, plan_rejected, abort_failed, rebuild_failed, dispatch_failed, not_selected, or terminal). Only these two values appear in the lifecycle; every other row carries one of them.
  • Empty eligibility and recovery_result rows belong to the lifecycle transitions that do not record a metric row: private filter rows that are not filter_evaluated, a second eligibility while eligible_pending, provider errors the liveness filter did not treat as a stall, and non-ExactReplay recovery observations that fall outside the private cycle. They are documented here so the safe-log field vocabulary is complete and not read as implying a missing classification.
  • Current immutable observations yield provider_health=unknown because the predecessor's private filter_evaluated observation does not carry provider health — health lives only in the request-local recovery state bridge, never in the immutable timeline. The closed classifier reserves available and unavailable for future health-bearing observations without claiming either is currently emitted.
  • Source: apps/edge/internal/openai/liveness_recovery_observability.go; test: apps/edge/internal/openai/liveness_recovery_observability_test.go::TestOpenAILivenessObservationSink and TestOpenAILivenessRecoveryObservability.

Fresh health recovery in provider snapshots

A recovered provider appears in the existing Edge provider snapshot overlay as status=available, health=available, with effective capacity restored to configured values. The snapshot reflects the same (node_id, connection_generation, provider_id) key used by the runtime overlay. A newer connection generation does not inherit the old overlay.

Leakage boundary

Operational projections exclude raw payloads, credentials, caller-controlled identities, and any unbounded identifier from metric labels and general structured logs. The exclusion applies to metric labels and general logs only; valid typed terminal metadata (e.g. run_id, adapter, target on the allowlisted stall metadata map) remains on the wire as already required by the typed terminal contract.

Verification

  • go test -count=1 ./packages/go/execution
  • go test -race -count=1 ./packages/go/execution
  • go vet ./packages/go/execution