iop/docs/edge-local-dev-guide.md
toki 39fa1da55b fix(benchmark): 원샷 비교 실행 실패를 해소한다
호출기별 격리 쓰기 계약과 Gemini ingress, 단일 요청 stage 처리를 맞춰 실제 9-cell 비교가 생성물을 남길 수 있게 한다.
2026-08-12 14:37:17 +09:00

16 KiB

Edge Local Quickstart

이 문서는 Edge를 로컬에서 빌드하고 실행해 Node를 연결하고, OpenAI-compatible smoke까지 확인하는 최소 절차다.

설정 기준은 edge.yaml이다. 주소는 환경 변수로 흩뿌리지 않는다.

1. Package

make build

2. Edge 설정

rm -rf "$HOME/iop-edge-test"
mkdir -p "$HOME/iop-edge-test"
tar -xzf build/packages/iop-edge-<platform>.tar.gz -C "$HOME/iop-edge-test"
cd "$HOME/iop-edge-test/iop-edge-<platform>"
./iop-edge config init

edge.yaml에서 아래 값만 맞춘다. 모델만 필요하면 바꾼다.

Provider-first 구성 (권장)

Provider-first는 models[]로 라우팅 키를 정의하고 nodes[].providers[]로 provider 후보를 선언한다. 두 구조를 models[].providers[provider_id] 키로 연결한다.

models:
  # Top-level catalog: canonical routing key → provider-pool mapping.
  # models[].id is the external model id; providers maps provider id → served model.
  - id: "gemma4:26b"
    providers:
      ollama-gemma: "gemma4:26b"

edge:
  id: "edge-1"
  name: "Edge 1"

server:
  listen: "0.0.0.0:19090"

bootstrap:
  listen: "0.0.0.0:18080"
  artifact_dir: "artifacts"

refresh:
  enabled: true
  listen: "127.0.0.1:19093"

# Legacy adapter/target fallback (backward compat only).
# New deploys should use models[] + nodes[].providers[].
openai:
  enabled: true
  listen: "0.0.0.0:18081"
  provider_id: "ollama-gemma"
  adapter: "ollama"
  target: "gemma4:26b"

nodes:
  - id: "node-ollama-1"
    providers:
      - id: "ollama-gemma"
        type: "ollama"
        category: "local_inference"
        base_url: "http://127.0.0.1:11434"
        models:
          - "gemma4:26b"
        capacity: 2

models[]를 사용할 경우 openai.adapter/openai.target은 하향 호환 fallback이다. 실제 routing은 nodes[].providers[].id를 기준으로 provider-pool에서 선택된다.

확인:

./iop-edge --config edge.yaml env
./iop-edge --config edge.yaml config check

3. Node 등록

./iop-edge --config edge.yaml node register node-ollama-1 \
  --adapter ollama \
  --ollama-base-url http://127.0.0.1:11434

출력된 OS별 bootstrap 명령 원문을 보관한다. 명령 원문에는 실제 token이 포함되므로 tracked 문서에는 기록하지 않는다.

4. Edge 실행

mkdir -p logs run
nohup ./iop-edge --config edge.yaml serve > logs/edge.stdout.log 2>&1 &
echo $! > run/iop-edge.pid

확인:

# Control Plane에 연결된 Edge 상태 조회
curl -fsS http://<control-plane-host>:18000/edges

5. Node 실행

3단계에서 출력된 bootstrap 명령을 Node host에서 그대로 실행한다. Linux/macOS는 generated curl | bash 명령을 사용하고, Windows native PowerShell은 generated .ps1 bootstrap과 Start-IopNode 함수를 사용한다.

Node 연결 확인:

curl -fsS http://<control-plane-host>:18000/edges/<edge-id>/status

6. 기본 Smoke

curl -fsS http://<edge-host>:18081/v1/models

./iop-edge --config edge.yaml smoke openai \
  --model gemma4:26b \
  --base-url http://<edge-host>:18081 \
  --timeout 60s

외부 OpenAI-compatible client base URL:

http://<edge-host>:18081/v1

7. Thinking/Reasoning 제어 smoke

/v1/chat/completions 요청은 OpenAI-compatible field를 기본으로 사용한다. Provider-pool pure passthrough는 selected provider가 지원하는 OpenAI-compatible 표준 field와 provider extension field를 보존해야 하며, chat_template_kwargs 같은 provider-native option을 IOP allowlist로 막지 않는다. think, reasoning_effort, thinking_token_budget, include_reasoning은 normalized backend 또는 provider별 차이를 보완하기 위한 IOP 확장 field다.

  • dev GX10 laguna-s:2.1 공식 sampling baseline은 Laguna S 2.1 모델 카드의 vLLM recipe를 따른다: temperature=0.7, top_p=0.95, model generation_configtop_k=20.
  • GX10 공식 근거: https://huggingface.co/poolside/Laguna-S-2.1
  • dev OneXPlayer/RTX5090 ornith:35b sampling baseline은 temperature=0.6, top_p=0.95, top_k=20을 유지한다. 공식 예제에 없는 non-neutral repeat_penalty는 임의로 추가하지 않는다. RTX5090 saved recipe의 --repeat-penalty 1.0--min-p 0.00은 출력을 바꾸지 않는 neutral runtime serialization로만 유지한다.
  • Ornith 공식 근거: https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B
  • 이 값은 caller가 sampling field를 생략했을 때 쓰는 provider 기본값이다. caller가 명시한 sampling field가 있으면 요청값이 우선한다.
  • dev GX10 vLLM은 poolside/Laguna-S-2.1-NVFP4와 quantization-matched Laguna-S-2.1-DFlash-NVFP4를 사용하고 --tool-call-parser poolside_v1, --reasoning-parser poolside_v1, --chat-template /run/iop/laguna-s-2.1-thinking.jinja, --default-chat-template-kwargs '{"enable_thinking":true}', --override-generation-config '{"temperature":0.7,"top_p":0.95}'로 설정한다.
  • GX10 host의 /home/toki/iop-gx10-vllm/laguna-s-2.1-thinking.jinja는 container에 read-only bind한다. stock generation prefix <think>가 첫 생성 토큰 </think>를 유도해 reasoning이 비는 현상이 재현됐으므로, dev template은 <think>\n을 사용한다.
  • dev OneXPlayer/RTX5090 Lemonade는 해당 Ornith recipe의 llamacpp_args--temp 0.6 --top-p 0.95 --top-k 20을 저장한다. RTX5090 수동 toggle은 이 sampling 값과 함께 preserve_thinking=true, unified KV, Q8 KV, neutral min-p/repeat-penalty를 exact profile로 검증한다.
  • Laguna reasoning은 provider-native reasoning field로 나가며 Pi openai-completions가 이를 thinking_delta로 소비한다. Pi Laguna profile은 thinking level을 chat_template_kwargs.enable_thinking으로 전달하고 preserve_thinking=true로 이전 assistant reasoning을 유지한다.
  • 출력 smoke는 같은 요청을 Pi highoff로 대조한다. high에서는 thinking_start/thinking_delta/thinking_end와 최종 text가, off에서는 thinking event 0개와 최종 text가 나와야 한다. agentic multi-turn에서는 tool-call 전후 reasoning, tool result, 최종 text까지 확인한다.
  • 현재 dev-corp provider-pool device mapping은 gemma4:26b -> Mac Studio provider capacity 5, ornith:35b -> DGX Spark 01/02 provider 합산 capacity 8이다. 세부 endpoint와 runtime args는 agent-test/inventory-dev-corp.yaml을 기준으로 한다.
  • dev-corp provider-pool 안정 smoke: 기본 smoke에서는 think, reasoning_effort, thinking_token_budget을 생략하고 현재 provider 기본값을 유지한다. Provider-native passthrough를 검증할 때는 selected provider가 직접 지원하는 field를 그대로 보낸다.
  • 현재 dev-corp public capacity smoke 표준은 public OpenAI-compatible base https://digitalplatform.iop.ai.kr/v1/chat/completions에서 ornith:35b 9/6 동시 요청과 gemma4:26b 9/6 동시 요청을 각각 확인하는 방식이다.
  • 일반 표준 caller의 stream=false 측정은 /v1/chat/completions에서 요청 파라미터만으로 확인한다.
  • include_reasoning=false는 non-provider normalized route의 hide 동작 기준이다. dev-corp provider-pool pure passthrough에서는 client가 reasoning field를 선택적으로 무시/제거한다.
  • think=false 또는 reasoning_effort=none은 hide-only 옵션이 아니라 thinking disable 요청이다. Provider-native field가 있는 경우 해당 field를 우선 사용해 passthrough 보존을 검증한다.
  • Provider-pool passthrough 파라미터의 세부 계약과 금지/허용 범위는 agent-contract/outer/openai-compatible-api.md를 기준으로 한다.

예시 (dev-corp gemma4:26b provider-pool non-stream 측정, think 생략):

아래 18081 포트는 local Edge 예시다. dev-corp public smoke에서는 base URL을 https://digitalplatform.iop.ai.kr/v1로 바꾼다. Direct Edge listener http://digitalplatform.iop.ai.kr:18086/v1도 동작하지만 사용자-facing 기본값은 포트 없는 public URL이다.

curl -fsS http://<edge-host>:18081/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -H 'Authorization: Bearer <token>' \
  -d '{"model":"gemma4:26b","messages":[{"role":"user","content":"hello"}],"stream":false}'

예시 (non-provider normalized route에서 response reasoning_content 숨김):

curl -fsS http://<edge-host>:18081/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"<model-alias>","messages":[{"role":"user","content":"hello"}],"include_reasoning":false}'

8. Raw text tool-call boundary smoke

긴 agent prompt와 tools[] 요청은 provider가 native tool_calls 대신 assistant content에 raw 블록을 담아 응답할 수 있다. 계약 기준은 agent-contract/outer/openai-compatible-api.md의 Chat Completions raw text tool-call 정규화 정책이다.

요청 payload는 tracked 문서에 secret 없이 남기고, token은 원격 환경 변수에서 주입한다. 실제 token 원문을 명령/로그/보고에 남기지 않는다.

판정 기준:

  • 응답 body와 SSE delta 어디에도 raw tool block 원문이 성공 content로 남지 않는다.
  • 요청 tools[]에 있는 valid tool-call은 message.tool_calls 또는 stream delta.tool_callsfinish_reason: "tool_calls"로 정규화된다.
  • 요청 tools[]에 없는 unknown tool hallucination이나 malformed 블록은 success content가 아니라 tool_validation_error로 끝난다.

evidence는 tracked 문서가 아니라 ignored run 위치(agent-test/runs/**) 또는 code-review output path에 저장한다.

9. Managed credential plane and TLS startup

Managed mode is a separate startup profile, not a live toggle. Keep every certificate, private key, at-rest keyring, issuer key, recipient key, IOP token, and provider credential outside the checkout in an operator-owned directory with restrictive permissions. Tracked YAML contains file paths only.

The complete chain must be configured before any process starts:

  • Control Plane credential HTTPS: server certificate/key for the credential API; the caller validates its CA and server name.
  • Control Plane to Edge: mutual TLS with CA-signed workload identities. The Control Plane expects role=edge and the hello edge_id; Edge expects the configured Control Plane role/name and server name.
  • Edge to Node: mutual TLS with exact Edge/Node role and name expectations.
  • Lease crypto: Control Plane mounts the at-rest keyring and issuer private key; Node mounts its recipient private key and the issuer public key. Edge receives no decryption key.
  • OpenAI/Anthropic ingress: Edge serves TLS whenever managed mode is enabled.

Representative Edge settings:

credential_plane:
  enabled: true
  lease_ttl_seconds: 30
  lease_cache_size: 256

tls:
  enabled: true
  cert: /etc/iop/secrets/edge.crt
  key: /etc/iop/secrets/edge.key
  ca: /etc/iop/secrets/credential-ca.crt
  peer_role: node

control_plane:
  enabled: true
  wire_addr: cp.internal:18002
  tls:
    enabled: true
    cert: /etc/iop/secrets/edge.crt
    key: /etc/iop/secrets/edge.key
    ca: /etc/iop/secrets/credential-ca.crt
    server_name: cp.internal
    peer_role: control-plane

openai:
  enabled: true
  tls:
    enabled: true
    cert: /etc/iop/secrets/edge-http.crt
    key: /etc/iop/secrets/edge-http.key

Managed validation rejects legacy principal mappings, openai.bearer_token, openai.provider_auth, provider credential headers/environment/arguments, endpoint user-info, or any plaintext hop. Use config check before serve; do not weaken validation to mix legacy and managed sources.

Safe slot lifecycle

  1. Run principal bootstrap locally on the Control Plane host. Capture the one-time token only in a protected secret channel; never paste it into YAML, shell history, logs, review artifacts, or chat.
  2. Call the dedicated credential HTTPS listener with that bearer token. Create accepts provider material only as application/octet-stream plus the vendor/kind/alias headers. Create the route separately with slot id, profile id, upstream model, and resource selector.
  3. Confirm the authenticated principal sees only its projected route id/alias from Edge model discovery. A managed request must bind to that route and must never fall back to a legacy model, provider, or another same-model slot.
  4. Rotate with the current IOP-Expected-Revision. Confirm the next successful provider attempt reports the new safe credential_revision; the former plaintext must not appear in SQLite, logs, metrics, or captured output.
  5. Disable for reversible suspension or revoke for permanent invalidation. Confirm a later full request fails, no Node lease consumption or upstream counter advances, and a same-model slot is not selected as fallback.
  6. For lease-expiry confirmation, wait past the configured TTL and verify an old envelope cannot be replayed. A later authorized request must acquire a fresh lease against the current projection and revisions.

The repository qualification exercises CA-signed identities, no-cert/wrong-peer failures, two same-model slots, ciphertext persistence, rotation attribution, and post-revoke no-fallback:

credential_smoke_parent="$(mktemp -d /config/workspace/iop-credential-slot-guide.XXXXXX)"
TMPDIR="$credential_smoke_parent" make test-credential-slot-smoke
rmdir "$credential_smoke_parent"

The deterministic Messages qualification succeeds alongside Chat: the Control Plane canonicalizes built-in lowercase API-key header names (for example x-api-key to X-Api-Key) before signing the lease scope, so both managed profiles reach Node/upstream exactly once with their exact header semantics. A lease failure fails closed before dispatch and never falls back to caller auth or another slot; treat a Chat-only result or any fallback as a qualification failure.

Official agy route smoke

공식 agy 1.1.12 API-key provider는 upstream Gemini key가 아니라 관리형 IOP principal token을 사용한다. dev operator가 이미 발급한 token과 CA 파일을 보호된 token/ 아래에 둔 경우 값을 명령행에 직접 쓰지 않고 다음처럼 읽는다.

read -r IOP_BENCH_TOKEN < token/.iop-bench
export GEMINI_API_KEY="$IOP_BENCH_TOKEN"
export SSL_CERT_FILE="$PWD/token/iop-dev-ca.pem"
export NODE_EXTRA_CA_CERTS="$PWD/token/iop-dev-ca.pem"

GOOGLE_GEMINI_BASE_URL="https://<edge-host>:<https-port>/gemini/<direct-route-id>" \
  agy --sandbox --output-format stream-json --model 'Gemini 3.6 Flash' \
  --print 'Reply only with OK. Do not use tools or modify files.'

GOOGLE_GEMINI_BASE_URL="https://<edge-host>:<https-port>/gemini/<hybrid-preset-id>" \
  agy --sandbox --output-format stream-json --model 'Gemini 3.6 Flash' \
  --print 'Inspect README.md and report only its first Markdown heading. Do not modify files.'

unset IOP_BENCH_TOKEN GEMINI_API_KEY SSL_CERT_FILE NODE_EXTRA_CA_CERTS

--effort는 API-key provider 호출에 넣지 않는다. direct와 hybrid 모두 JSONL의 마지막 record가 event=result, 중첩 result.status=SUCCESS 한 건이어야 한다. hybrid는 plan/work/review가 포함되므로 direct보다 오래 걸릴 수 있으며, caller timeout을 이유로 같은 scored attempt를 재실행하지 않는다. Gemini ingress와 공식 event 구조의 상세 계약은 agent-contract/outer/gemini-compatible-api.md를 기준으로 한다.

benchmark 전체 preflight에서는 Claude/agy/Codex에 같은 IOP principal을 secret environment reference로 연결하고 IOP_BENCH_CONFIG_OBSERVATION_ENV가 가리키는 operator-owned route/binding observation을 함께 제공한다. 사설 CA 환경은 각 isolated caller child에 SSL_CERT_FILENODE_EXTRA_CA_CERTS로 전달된다. preflight가 모든 cell을 ready로 판정하기 전에는 scored run을 시작하지 않는다.

Incident redaction check

Before retaining logs or evidence, reject any artifact containing an IOP bearer token, provider credential, slot alias, lease id, certificate private key, keyring material, recipient/issuer private key, target URL with credentials, prompt, or response body. Public metrics may contain only stable safe references such as credential_slot_ref and credential_revision; request/run/session/attempt/node ids and raw payloads are not credential-attribution labels.