diff --git a/agent-ops/skills/common/code-review/SKILL.md b/agent-ops/skills/common/code-review/SKILL.md index f4ae99b1..bfbb45c8 100644 --- a/agent-ops/skills/common/code-review/SKILL.md +++ b/agent-ops/skills/common/code-review/SKILL.md @@ -164,7 +164,7 @@ The diff is the starting point, not the boundary. Follow behavior and API connec Review scope control: - Use the plan's commands and checkpoints as the primary evidence. Add one focused, possibly table-driven reproducer only when needed to prove a suspected blocking defect; do not build speculative exhaustive probe matrices. -- **NO OVERENGINEERING** — Do not require or propose anything beyond the user request and correctness. +- **NO OVERENGINEERING** — Do not require or propose anything beyond the requested work and correctness. - Execute the applicable plan verification commands and any focused reproducer needed for the verdict. Treat implementation-owned output as a handoff and comparison source, not as a substitute for fresh reviewer verification. If recorded output is absent or insufficient but the command is available and safe in the current authorized environment, run it and repair `Verification Results` before classifying findings. If a check fails, collect enough source/runtime data to establish the root cause and one implementable fix; never emit a diagnostic-only finding that asks the next worker to investigate or choose among alternatives. - In a follow-up review, keep Required findings within the current plan, inherited Required findings, direct regressions from the fix, and concrete violations of the original SDD or contract acceptance criteria. Exclude unrelated pre-existing work from the verdict and Required/Suggested/Nit counts; mention it only in the final report as an out-of-scope task candidate. - Before adding a new Required that the current plan did not state, cite the exact original plan/SDD/contract criterion it violates or provide a concrete failing case. Do not require a preferred test shape when existing deterministic evidence proves the same behavior. diff --git a/agent-ops/skills/common/create-roadmap/SKILL.md b/agent-ops/skills/common/create-roadmap/SKILL.md index a4da1791..7a455c6c 100644 --- a/agent-ops/skills/common/create-roadmap/SKILL.md +++ b/agent-ops/skills/common/create-roadmap/SKILL.md @@ -85,7 +85,7 @@ agent-roadmap/ ## 작성 규칙 -- **NO OVERENGINEERING** — Do not add anything beyond the user request and required behavior. +- **과설계 금지:** 사용자 요청과 동작에 반드시 필요한 것만 만든다. 그 외에는 추가하거나 다른 항목으로 남기지 않는다. - 기본 작성 언어는 한국어다. - 상태 표기는 `[스케치]`, `[계획]`, `[진행중]`, `[검토중]`, `[완료]`, `[보류]`, `[폐기]` 중 하나만 사용한다. - `[스케치]`는 방향성, 문제의식, 후보 범위, 미정 질문을 기록하는 컨셉 상태다. 구현 가능한 계획이 아니므로 구현 계획 생성 대상이 아니다. diff --git a/agent-ops/skills/common/plan/SKILL.md b/agent-ops/skills/common/plan/SKILL.md index 9e799ed2..62033417 100644 --- a/agent-ops/skills/common/plan/SKILL.md +++ b/agent-ops/skills/common/plan/SKILL.md @@ -195,7 +195,7 @@ Before choosing plan files or task directory names, apply the split decision pol Complete all items below before creating active plan/review files. Work through them in order; do not proceed to the next step until every checkbox is done. Keep the user request as the scope anchor and reconcile derived acceptance conditions before the split decision; do not create a separate routing summary. In `prepare-follow-up`, treat the reviewer's closed finding packet as the decision authority: repository reads validate its consistency and supply implementation mechanics, but do not reopen root cause or solution selection. If required evidence, root cause, or a selected fix is missing or contradicted, return `needs_evidence` to code-review so the reviewer corrects it in the same review pass; never pass investigation or alternatives to the worker. The only allowed file edits before writing plan/review files are local `agent-roadmap/current.md` creation or `.gitignore` block repair needed for roadmap routing. - [ ] **Resolve verification context** — because implementation plans include verification, consume supplied `verification_context` when present and confirm its source paths, commands, expected results, preconditions, constraints, gaps, and confidence still apply. On first pass, derive missing facts from repository manifests, scripts, workflows, domain rules, related tests, user-provided environment facts, and safe read-only probes. In `prepare-follow-up`, require the reviewer to have collected every fact needed for diagnosis and fix selection; derive only mechanical command/path details, and return `needs_evidence` rather than performing missing review analysis. Record which facts came from the handoff and which came from repository-native validation. A missing optional first-pass handoff is not a user-review blocker. -- [ ] **NO OVERENGINEERING** — Do not add anything beyond the user request and required behavior. +- [ ] **NO OVERENGINEERING** — Plan only what the user requested and what is strictly required to make it work. Do not add anything else. - [ ] **Read all source files in full** — read every source file the change will touch, whole file. No partial reads. - [ ] **Preflight external verification** — when any required verification leaves the current checkout, including remote runner, field/bootstrap, external provider, Docker/code-server, emulator/device, or shared long-running runtime, confirm or derive a read-only preflight before writing final verification commands. Record runner, repo root/workdir, branch/HEAD/dirty state, source sync status, binary/artifact paths, command help/version output needed by the verification, config path, runtime identity, ports/process state, external hosts, and OS/arch assumptions. If the preflight shows stale artifacts, dirty/divergent checkout, wrong identity, missing command, closed ports, host OS mismatch, or unsynced source, add an explicit setup/sync/rebuild step or report the blocker. - [ ] **Read all test files in full** — read every test file that exercises the changed behavior, including files identified by the verification context and repository test layout. diff --git a/agent-ops/skills/common/update-roadmap/SKILL.md b/agent-ops/skills/common/update-roadmap/SKILL.md index b0cdfbbe..448db525 100644 --- a/agent-ops/skills/common/update-roadmap/SKILL.md +++ b/agent-ops/skills/common/update-roadmap/SKILL.md @@ -250,7 +250,7 @@ agent-roadmap/ | 작업 컨텍스트/TODO | 에이전트가 확정할 수 없는 결정 또는 조사/확인이 먼저 필요해 기능 Task로 확정하기 어렵다 | - 먼저 요청 내용의 규모를 판정한다. 배치 위치를 찾기 전에 `phase`, `milestone`, `epic`, `task`, `subtask`, `context` 중 가장 작은 충분한 단위를 고른다. -- **NO OVERENGINEERING** — Do not add or retain anything beyond the user request and required behavior. +- **과설계 금지:** 사용자 요청과 필수 동작 외에는 추가하거나 남기지 않는다. - 요청이 방향성, 문제의식, 컨셉, 운영 원칙 수준이고 기능 Task나 실행 범위가 아직 부족하면 새 항목의 상태는 `[스케치]`로 둔다. - `[스케치]` Phase/Milestone을 만들 때는 `승격 조건`에 `[계획]`으로 전환하기 위해 필요한 정의, 결정, 경계, 후속 구현 Milestone 후보를 체크리스트로 남긴다. - 가장 작은 충분한 단위 원칙을 따른다. 애매하면 새 Phase나 새 Milestone으로 키우지 말고, 기존 Milestone의 Epic/Task에 넣을 수 있는지 먼저 확인한다. diff --git a/agent-ops/skills/project/dev-runtime-deploy/SKILL.md b/agent-ops/skills/project/dev-runtime-deploy/SKILL.md index 0212f305..e71431e8 100644 --- a/agent-ops/skills/project/dev-runtime-deploy/SKILL.md +++ b/agent-ops/skills/project/dev-runtime-deploy/SKILL.md @@ -143,13 +143,13 @@ dev-runtime provider pool을 `dev` 기준 git-flow release로 배포한다. `dev 9. **Run the OpenAI-compatible capacity smoke** - Qualify `/v1/responses` and `/v1/chat/completions` separately. Do not fail solely because the unimplemented legacy `/v1/completions` endpoint is absent. - Use a short request that asks for a structured 700-1200-token answer and explicitly bounds provider-native thinking. Compute the emitted JSON Unicode rune count, `runes/4 + runes/16` estimate, and context class. Do not reuse long-context or repetition fixtures as a normal-capacity oracle. - - In managed mode, authenticate `GET /v1/credentials/routes` with the same principal token used at OpenAI ingress. Require one active exact route alias, then intersect its `resource_selector`, profile, and upstream model with one healthy connected provider snapshot. - - Run `scripts/e2e-openai-managed-capacity-smoke.sh` for one endpoint and route at a time. Require the emitted request to classify as `normal`, then send the selected provider's `capacity` plus one request. If it classifies as `long`, fail this normal-capacity gate and use the separate long-context admission smoke with `long_context_capacity`; never add capacity from a provider excluded by the authenticated route. - - Keep current projected Ornith cases separate: qualify `ornith:35b` only through `onexplayer-lemonade`, and qualify `ornith-fast` only through `rtx5090-lemonade`. Run Chat and Responses independently for each alias. Do not mutate routes, capacities, or timeout values to make the smoke pass. + - In managed mode, authenticate `GET /v1/credentials/routes` with the same principal token used at OpenAI ingress. Require one active exact route alias. An explicit `resource_selector` must intersect with one exact healthy provider snapshot; `default` intersects the route profile/upstream model and model group with every matching healthy provider snapshot. + - Run `scripts/e2e-openai-managed-capacity-smoke.sh` for one endpoint and route at a time. Require the emitted request to classify as `normal`. For an explicit selector, send the selected provider's `capacity` plus one request. For a `default` pool, send the sum of every eligible matching provider's `capacity` plus one and require each provider's observed peak to reach its own capacity. If it classifies as `long`, fail this normal-capacity gate and use the separate long-context admission smoke with `long_context_capacity`; never add capacity from a provider excluded by the authenticated route. + - Keep current projected Ornith cases separate: qualify the `ornith:35b` default pool through GX10 `4` + OneXPlayer `3` with RTX excluded, and qualify `ornith-fast` only through `rtx5090-lemonade`. Run Chat and Responses independently for each alias. Do not mutate routes, capacities, or timeout values to make the smoke pass. - Require HTTP 200 for every request, exactly one endpoint-native success terminal and one `[DONE]` per stream, selected-provider peak `in_flight` equal to eligible capacity, `queued >= 1`, and selected final counters `0/0`. Treat a missing route match, changed capacity, wrong context class, missing status sample, or incomplete terminal as a fail-closed release blocker. - Every invocation must allocate a distinct mode-0700 directory under ignored `agent-test/runs/**`. Accept only request/result/status files bound to that invocation's manifest and timestamps. Preserve raw routes, request bodies, response bodies, curl errors, token/header values, route ids, slot ids, prompts, and output only inside that directory; retain only the script's allowlisted sanitized summary as task or tracked evidence. - - For unprojected or legacy provider pools, retain the existing model-group capacity rule. Current baselines are Laguna `laguna-s:2.1` on GX10 capacity `4` (five requests) and Qwen `qwen3.6:35b` on mac-mlx-vllm capacity `2` (three requests). - - Allow Qwen and Laguna reasoning/thinking text as normal output. For Laguna, compare Pi `high` thinking events with `off` zero-thinking events and final text. For agent/tool-call qualification, test forced, automatic, streaming, and multi-turn tool calls separately from capacity qualification. + - For unprojected provider pools, retain the existing model-group capacity rule. The current unprojected baseline is Qwen `qwen3.6:35b` on mac-mlx-vllm capacity `2` (three requests). Stopped Laguna is rollback-only and has no active provider capacity. + - Allow Qwen and Ornith reasoning/thinking text as normal output. For Ornith, compare Pi `high` thinking events with `off` zero-thinking events and final text. For agent/tool-call qualification, test forced, automatic, streaming, and multi-turn tool calls separately from capacity qualification. 10. **git-flow release finish와 tag 반영** - 빌드 전·후 테스트, 배포 후 연결 검증, `/v1/responses`와 `/v1/chat/completions` capacity smoke가 모두 성공했는지 다시 확인한다. Managed route가 있으면 `scripts/e2e-openai-managed-capacity-smoke.sh`의 네 Ornith route/endpoint case와 current-run provenance gate가 모두 통과해야 한다. diff --git a/agent-roadmap/phase/knowledge-tool-optimization-extension/PHASE.md b/agent-roadmap/phase/knowledge-tool-optimization-extension/PHASE.md index 07c8c3cf..11470d0d 100644 --- a/agent-roadmap/phase/knowledge-tool-optimization-extension/PHASE.md +++ b/agent-roadmap/phase/knowledge-tool-optimization-extension/PHASE.md @@ -73,6 +73,10 @@ Phase를 가로지르는 실제 다음 작업 선택은 [전역 마일스톤 실 - 경로: [[surface-01] Inference API Surface와 실행 Lifecycle 책임 경계 리팩터링](milestones/inference-api-surface-execution-lifecycle-refactor.md) - 요약: `iop-s0`의 bench-02 결과와 제품 delta가 `dev`에 정합화된 뒤 OpenAI Chat/Responses, Anthropic Messages, Gemini의 wire 계약은 surface별로 유지하고, 기존 provider service 경계를 재사용하며 cross-surface handler 재진입과 단계별 lifecycle 소유권만 정리한다. +- [스케치] [gate-01] StreamGate 경량 통과 모드 + - 경로: [[gate-01] StreamGate 경량 통과 모드](milestones/stream-gate-operating-modes.md) + - 요약: 전역 StreamGate의 공통 응답 수명주기는 유지하면서 의미 필터, 응답 보류와 자동 복구 없이 최소 확인과 동기 추적만 수행하는 경량 통과 모드를 추가한다. 현재 후속 요청 후보 거부 현상은 별도 핸드오프로 보존한다. + - [스케치] [output-03] OpenAI-compatible Runtime Output Integrity Filter - 경로: [[output-03] OpenAI-compatible Runtime Output Integrity Filter](milestones/openai-compatible-runtime-output-integrity-filter.md) - 요약: terminal assistant 응답이 content, valid tool call, 명시 허용 structured/error finish 중 하나를 만족해야 한다는 runtime invariant를 정의하고, empty terminal, reasoning-only, incomplete tool-call syntax 같은 deterministic violation을 공통 filter pipeline과 bounded retry 정책으로 묶는다. diff --git a/agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/stream-gate-operating-modes.md b/agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/stream-gate-operating-modes.md new file mode 100644 index 00000000..d71e83f5 --- /dev/null +++ b/agent-roadmap/phase/knowledge-tool-optimization-extension/milestones/stream-gate-operating-modes.md @@ -0,0 +1,73 @@ +# Milestone: [gate-01] StreamGate 경량 통과 모드 + +## 위치 + +- Roadmap: [ROADMAP.md](../../../ROADMAP.md) +- Phase: [PHASE.md](../PHASE.md) + +## 목표 + +전역 StreamGate의 응답 수명주기는 유지하면서 의미 필터, 응답 보류와 자동 복구 없이 최소 확인과 동기 추적만 수행하는 경량 통과 모드를 추가한다. 기존 필터 동작은 그대로 유지하고 설정으로 두 동작을 명확히 구분한다. + +## 상태 + +[스케치] + +## 승격 조건 + +- [ ] 경량 통과 모드의 최소 확인 범위와 관측 항목을 확정한다. +- [ ] 기존 필터 동작과 경량 통과 모드를 구분하는 설정 형태와 기본값을 확정한다. +- [ ] 설정 변경이 포함되므로 `[계획]` 승격 시 SDD 필요 여부를 다시 판정한다. + +## 구현 잠금 + +- 상태: 잠금 +- SDD: 불필요 +- SDD 문서: 없음 +- SDD 사유: 현재는 모드 의미와 경계를 정하는 스케치이며, 설정·상태 전이 구현 범위가 확정되지 않았다. `[계획]` 승격 시 SDD 필요 여부를 다시 판정한다. +- 잠금 해제 조건: 승격 조건 충족 +- 결정 필요: 아래 목록 + - 경량 통과에서 남길 최소 동기 관측 항목 + - 설정 이름과 기본값 + +## 범위 + +- 전역 StreamGate의 경량 통과 설정 +- 최소 확인과 동기 추적 +- 의미 필터, provider filter capability 승인, 응답 보류와 자동 복구 없이 기존 release 경로로 전달하는 동작 +- 기존 필터 동작과 경량 통과 동작의 기본 회귀 검증 + +## 기능 + +### Epic: [light-pass] 경량 통과 모드 + +기존 공통 StreamGate 수명주기를 유지하면서 요청을 지연시키는 필터 동작만 제외한다. + +- [ ] [mode-config] 기존 필터 동작과 경량 통과 동작을 구분하는 설정을 추가하고 요청 시작 시 선택된 값을 고정한다. +- [ ] [pass-through] 경량 통과에서는 최소 확인과 동기 추적만 수행하고 의미 필터, filter capability 승인, 응답 보류, 자동 복구와 provider 전환을 실행하지 않는다. +- [ ] [regression] 도구 호출 전후의 후속 요청에서 event 순서, terminal, usage와 cancel이 보존되고 필터 대기나 자동 복구가 발생하지 않는지 검증한다. + +## 완료 리뷰 + +- 상태: 없음 +- 요청일: 없음 +- 완료 근거: 없음 +- 검토 항목: 없음 +- 리뷰 코멘트: 없음 + +## 범위 제외 + +- 개별 output filter의 탐지 규칙과 품질 개선 +- [StreamGate 후속 요청 후보 거부 문제](../../../../HANDOFF-stream-gate-followup-request-candidate-rejection.md)의 원인을 이 Milestone에서 미리 확정하거나 특정 수정안으로 고정하는 일 +- 공통 StreamGate runtime 밖에 별도의 endpoint별 응답 파이프라인을 추가하는 일 +- caller 또는 agent 제품별로 모드를 자동 변경하는 정책 +- 완전 비활성 모드와 추가 감시 단계 +- 장시간 원문 감시와 비동기 분석 + +## 작업 컨텍스트 + +- 관련 경로: `packages/go/streamgate`, `packages/go/config`, `apps/edge/internal/openai`, `apps/edge/internal/service`, `configs` +- 표준선: 경량 통과는 공통 StreamGate 수명주기를 유지하고 최소한의 동기 추적 외에는 응답을 보류하거나 자동 복구하지 않는다. +- 실행 순서와 차단 관계: [전역 마일스톤 실행 순서](../../../priority-queue.md) +- 관련 Milestone: [OpenAI-compatible 출력 검증 필터](openai-compatible-output-validation-filters.md), [Inference API Surface와 실행 Lifecycle 책임 경계 리팩터링](inference-api-surface-execution-lifecycle-refactor.md) +- 확인 필요: 승격 조건과 `구현 잠금 > 결정 필요` 항목 diff --git a/agent-roadmap/priority-queue.md b/agent-roadmap/priority-queue.md index 65617631..ab3410b3 100644 --- a/agent-roadmap/priority-queue.md +++ b/agent-roadmap/priority-queue.md @@ -21,6 +21,11 @@ 1. [[surface-01] Inference API Surface와 실행 Lifecycle 책임 경계 리팩터링](phase/knowledge-tool-optimization-extension/milestones/inference-api-surface-execution-lifecycle-refactor.md) OpenAI, Anthropic, Gemini wire 계약은 분리하고 기존 provider service 경계를 재사용하면서 cross-surface handler 재진입과 단계별 lifecycle 소유권을 정리한다. 구현 시작 조건은 workspace lock과 Milestone의 제품-delta 정합성 gate에서 관리한다. +### gate + +1. [[gate-01] StreamGate 경량 통과 모드](phase/knowledge-tool-optimization-extension/milestones/stream-gate-operating-modes.md) + 공통 StreamGate 수명주기를 유지하면서 의미 필터, 응답 보류와 자동 복구 없이 최소 확인과 동기 추적만 수행하는 경량 통과 모드를 추가한다. + ### output 1. [[output-01] OpenAI-compatible 출력 검증 필터](phase/knowledge-tool-optimization-extension/milestones/openai-compatible-output-validation-filters.md) diff --git a/agent-test/dev/edge-smoke.md b/agent-test/dev/edge-smoke.md index e80a0526..0d83f014 100644 --- a/agent-test/dev/edge-smoke.md +++ b/agent-test/dev/edge-smoke.md @@ -3,7 +3,7 @@ test_env: dev test_profile: edge-smoke domain: edge verification_type: smoke -last_rule_updated_at: 2026-08-13 +last_rule_updated_at: 2026-08-15 --- # edge-smoke dev 테스트 @@ -53,9 +53,9 @@ Claude Anthropic-compatible 단일 요청 Agent 실행을 검증할 때는 Claud - bootstrap HTTP: `http://toki-labs.com:18082` - Edge OpenAI-compatible base URL: `http://toki-labs.com:18083/v1` - Edge-Node TCP transport: `toki-labs.com:18084` -- active model aliases: `laguna-s:2.1`(GX10), `ornith:35b`/`ornith-fast`(OneXPlayer/RTX5090), `qwen3.6:35b`(mac-mlx-vllm) +- active model aliases: `ornith:35b`(GX10/OneXPlayer, raw capacity `7`), `ornith-fast`(RTX5090), `qwen3.6:35b`(mac-mlx-vllm). RTX5090은 `ornith:35b` 후보가 아니며 `laguna-s:2.1`은 stopped rollback 자산이라 active provider capacity가 없다. - host Pi dispatcher profile: `agent-test/inventory-agent.yaml`의 `environments.dev.agents.pi` 기준. 현재 기본 provider/model/thinking level은 `iop` / `glm-5.2` / `high`이고, dispatcher local-model route는 `iop/ornith:35b` / `high`다. 두 모델 모두 IOP Edge `http://toki-labs.com:18083/v1`을 사용하며 Pi local direct providers는 제거된 상태다. credential 원문 대신 같은 inventory의 SOPS `token_ref`만 기준으로 삼는다. -- provider/model separation: Laguna는 GX10 `poolside_v1`, Qwen은 mac `qwen`/`qwen3`, Ornith는 Lemonade runtime profile을 사용한다. dev-corp `gemma4:26b`와 stopped DiffusionGemma 설정을 이들 profile에 섞지 않는다. +- provider/model separation: GX10 Ornith는 `qwen3_xml`/`qwen3`, mac Qwen은 `qwen`/`qwen3`, Lemonade Ornith는 저장된 runtime profile을 사용한다. stopped Laguna의 `poolside_v1` profile이나 dev-corp `gemma4` profile을 이들 active provider에 섞지 않는다. 노드 후보: @@ -75,24 +75,24 @@ Claude Anthropic-compatible 단일 요청 Agent 실행을 검증할 때는 Claud - GX10 vLLM node: `gx10-vllm-node` / `gx10-vllm` - SSH/user: `ssh toki@192.168.0.91` - provider endpoint: `http://192.168.0.91:8001/v1` - - served model: `laguna-s:2.1` + - served model: `ornith:35b` - capacity baseline: `4` - priority baseline: `1` - - runtime baseline: custom image `iop-vllm-laguna-s21:v0.25.1-cu130`, model `poolside/Laguna-S-2.1-NVFP4` revision `07614121b31898586430f189d27a25a0be310843`, `--served-model-name laguna-s:2.1`, `--max-model-len 262144`, `--max-num-seqs 4`, `--gpu-memory-utilization 0.70`, `--enable-prefix-caching`, `--enable-auto-tool-choice --tool-call-parser poolside_v1`, `--reasoning-parser poolside_v1`, `--chat-template /run/iop/laguna-s-2.1-thinking.jinja`, `--default-chat-template-kwargs '{"enable_thinking":true}'`, `--override-generation-config '{"temperature":0.7,"top_p":0.95}'` - - thinking template baseline: host `/home/toki/iop-gx10-vllm/laguna-s-2.1-thinking.jinja`를 container에 read-only bind하고 generation prefix를 `\n`으로 둔다. stock `` prefix가 첫 토큰 ``를 유도해 reasoning이 비는 현상을 방지한다. - - speculative decoding baseline: `poolside/Laguna-S-2.1-DFlash-NVFP4` revision `723794750422b3efbf3a7b3af76dffb4ba035943`, `method=dflash`, `num_speculative_tokens=7` - - long-context admission baseline: observed GPU KV cache `373711` tokens, `long_context_capacity=1` (vLLM max concurrency `1.43x` for a 262144-token request) - - reasoning baseline: provider-native `reasoning` delta를 Edge passthrough가 보존하고 Pi가 `thinking_delta`로 소비한다. Pi Laguna compat는 thinking level을 `chat_template_kwargs.enable_thinking`으로 전달하고 `preserve_thinking=true`를 유지한다. + - runtime baseline: image `vllm/vllm-openai:nightly-aarch64`, model `deepreinforce-ai/Ornith-1.0-35B-FP8` revision `1ab57ce0b44950e498a88756f40ad1ed4d0f30ca`, `--served-model-name ornith:35b`, `--max-model-len 262144`, `--max-num-seqs 4`, `--gpu-memory-utilization 0.50`, `--enable-prefix-caching`, `--enable-auto-tool-choice --tool-call-parser qwen3_xml`, `--reasoning-parser qwen3`, `--trust-remote-code`, `--language-model-only`, `--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20}'` + - long-context admission baseline: conservative `total_context_tokens=1048576`, `long_context_capacity=4`; current startup reports GPU KV cache `1173993` tokens and full-context concurrency `4.48x`. + - runtime ownership: active container는 `iop-vllm-ornith35b-fp8`; stopped `iop-vllm-laguna-s21`은 명시적 rollback 전용이며 provider capacity에 포함하지 않는다. + - reasoning baseline: `chat_template_kwargs.enable_thinking`으로 model-native reasoning을 제어하고 `qwen3_xml`/`qwen3` parser를 사용한다. - workspace: `/home/toki/iop-gx10-vllm` - OneXPlayer Lemonade node: `onexplayer-lemonade-node` / `onexplayer-lemonade` - SSH/user: `ssh r0bin@192.168.0.59` - 접속 기준: 현재 작업 호스트에서 직접 SSH - - provider endpoint: `http://192.168.0.59:13305/v1` - - served model: `Ornith-1.0-35B-GGUF-llamacpp-tp1-Q5_K_M` - - capacity baseline: `1` (agent 장문 요청의 provider 독점 실행 기준) + - Node provider endpoint: `http://127.0.0.1:8001/v1` (OneX host-local llama backend) + - Lemonade lifecycle endpoint: `http://192.168.0.59:13305/v1` + - served model: `ornith:35b` (`--alias`); backend model은 `Ornith-1.0-35B-GGUF-llamacpp-tp1-Q5_K_M` + - capacity baseline: `3` - priority baseline: `2` - - load baseline: checkpoint `LordNeel/Ornith-1.0-35B-GGUF-llamacpp-tp1:Q5_K_M`, backend `vulkan`, ctx size `262144`, `llamacpp_args="--spec-type none -np 1 -cb -fa on -b 4096 -ub 1024 --kv-unified --temp 0.6 --top-p 0.95 --top-k 20"`, `save_options=true` - - long-context admission baseline: `total_context_tokens=262144`, `long_context_capacity=1` (IOP의 provider 직렬 admission과 llama-server의 `-np 1`을 일치시켜 agent turn별 prefix cache가 서로 다른 slot에 분산되지 않게 한다. `/slots`는 slot 1개와 `n_ctx=262144`를 보고한다) + - load baseline: checkpoint `LordNeel/Ornith-1.0-35B-GGUF-llamacpp-tp1:Q5_K_M`, backend `vulkan`, ctx size `524288`, `llamacpp_args="--spec-type none --alias ornith:35b -cb -fa on -b 4096 -ub 1024 --kv-unified --temp 0.6 --top-p 0.95 --top-k 20"`, `save_options=true` + - long-context admission baseline: `total_context_tokens=524288`, `long_context_capacity=2`; 고정 `-np` 없이 `/slots`가 자동 slot 4개와 각 `n_ctx=262144`를 보고하고 Edge normal capacity는 `3`으로 제한한다. - workspace: `C:/Users/r0bin/iop-field` - RTX5090 Lemonade node: `rtx5090-lemonade-node` / `rtx5090-lemonade` - SSH/user: `ssh iop-dev-rtx5090` @@ -106,7 +106,7 @@ Claude Anthropic-compatible 단일 요청 Agent 실행을 검증할 때는 Claud - process baseline: IOP 관련 Windows 부팅 자동 실행은 비활성이고, 수동 `remote-llm-toggle.ps1`이 Lemonade Server, Ornith, IOP Node를 순차 UP/DOWN한다. Startup shortcut, Run entry, Task Scheduler, Windows service는 IOP dev 배포가 생성·재생성하지 않는다. - workspace: `C:/Users/r0bin/iop-field` -GX10은 Laguna S 2.1 NVFP4 + DFlash NVFP4를 사용한다. OneXPlayer와 RTX5090 Lemonade는 Ornith Q5 GGUF를 사용하고 runtime speculative decoding은 `--spec-type none`으로 끈다. provider family별 parser/template을 섞지 않는다. +GX10은 Ornith FP8 vLLM을 사용하고 OneXPlayer와 RTX5090 Lemonade는 Ornith Q5 GGUF를 사용한다. stopped Laguna 컨테이너와 `poolside_v1` profile은 rollback 전용이며 active provider에 섞지 않는다. OneXPlayer Lemonade Node는 원격 runner나 Edge host에서 다시 SSH하거나 proxy process로 띄우지 않는다. 현재 작업 호스트에서 OneXPlayer Windows host에 `ssh r0bin@192.168.0.59`로 직접 접속한 뒤 generated PowerShell bootstrap을 실행한다. @@ -114,7 +114,7 @@ RTX5090 Lemonade Node도 원격 runner나 Edge host를 경유하지 않는다. mac-mlx-vllm provider는 mac-codex-node 소속 resource로, Edge host와 같은 macOS host에서 vllm-mlx process로 실행한다. vllm-mlx API는 외부에 직접 노출하지 않고 `127.0.0.1:8002`에 bind하며, Edge OpenAI-compatible adapter가 local bearer header로 호출한다. 운영 파일은 `/Users/toki/agent-work/iop-mlx-vllm/vllm-mlx.pid`, `logs/vllm-mlx.stdout.log`, `logs/vllm-mlx.stderr.log`를 기준으로 한다. Docker와 macOS 여유 메모리를 고려해 capacity는 `2`를 기본선으로 유지한다. -GX10 Laguna 공식 vLLM baseline은 `temperature=0.7`, `top_p=0.95`, model `generation_config`의 `top_k=20`, `tool_call_parser=poolside_v1`, `reasoning_parser=poolside_v1`, `enable_thinking=true`다. 현재 dev는 reasoning 출력을 위해 generation prefix를 `\n`으로 보정한 로컬 템플릿을 사용한다. Edge/Pi smoke는 `high`에서 thinking event, `off`에서 thinking event 0개와 최종 text를 대조하고 agentic multi-turn에서는 tool-call 전후 reasoning과 최종 text까지 확인한다. OneXPlayer/RTX5090 Ornith sampling은 기존 `temperature=0.6`, `top_p=0.95`, `top_k=20` 기준을 유지한다. +GX10/OneXPlayer/RTX5090 Ornith sampling은 `temperature=0.6`, `top_p=0.95`, `top_k=20` 기준을 유지한다. GX10은 `qwen3_xml`/`qwen3` parser와 `chat_template_kwargs.enable_thinking`을 사용하고, Edge/Pi smoke는 `high` reasoning과 `off` zero-reasoning 및 tool-call을 각각 검증한다. ## 명령 @@ -133,7 +133,7 @@ GX10 Laguna 공식 vLLM baseline은 `temperature=0.7`, `top_p=0.95`, model `gene - OpenAI-compatible 경계를 바꾼 경우 dev `18083` 기준 `iop-edge smoke openai` 또는 동등한 `/healthz`, `/v1/models`, `/v1/responses` 확인으로 edge service와 node adapter 경로 수렴을 확인한다. - runtime `edge.yaml`에서 `openai.enabled=true`이면 `openai.provider_id`가 현재 CLI/provider identity와 일치하는지 `config check`로 검증한다. 이 값이 빠진 candidate는 배포하지 않는다. - OpenAI-compatible tool-call 경계(`apps/edge/internal/openai/**`의 text tool-call 합성/validation)를 바꾼 경우 dev `18083` `/v1/chat/completions`에 Pi/Cline형 `tools[]` 요청을 non-stream/stream 각각 최소 1회 보내 raw text tool-call boundary smoke를 수행한다. token은 원격 환경 변수에서 주입하고 원문을 명령/로그/보고에 남기지 않는다. -- dev provider-pool tool-call drift를 확인할 때는 active model group별 provider만 섞는다. Laguna group은 `gx10-vllm` direct와 Edge `laguna-s:2.1`, Ornith group은 `onexplayer-lemonade`/`rtx5090-lemonade` direct와 Edge `ornith:35b`, Qwen group은 `mac-mlx-vllm` direct와 Edge `qwen3.6:35b`를 각각 검증한다. +- dev provider-pool tool-call drift를 확인할 때는 active model group별 provider만 섞는다. `ornith:35b` group은 `gx10-vllm`/`onexplayer-lemonade` direct와 Edge alias를 검증하고, RTX5090은 `ornith-fast`에서만 검증한다. Qwen group은 `mac-mlx-vllm` direct와 Edge `qwen3.6:35b`를 사용하며 stopped Laguna는 active drift 표본에 포함하지 않는다. - bootstrap/artifact 경계를 바꾼 경우 dev artifact/base URL 후보 `18082`가 local/test `18080` field baseline을 덮어쓰지 않는지 확인한다. ## 보조 검증 @@ -150,11 +150,12 @@ GX10 Laguna 공식 vLLM baseline은 `temperature=0.7`, `top_p=0.95`, model `gene - raw text tool-call boundary smoke: Pi/Cline형 `tools[]` 요청에서 응답 body와 SSE delta 어디에도 ``, `{{`, `<|mask_end|>` 원문이 성공 content로 남지 않는다. 요청 `tools[]`에 있는 valid raw text tool-call은 `message.tool_calls`(또는 stream `delta.tool_calls`)와 `finish_reason: "tool_calls"`로 정규화되고, unknown tool hallucination이나 malformed 블록은 성공 content가 아니라 non-stream `tool_validation_error`(HTTP 502) 또는 SSE `tool_validation_error` 이벤트로 끝난다. non-stream/strict buffered stream은 bounded retry 후 차단을 확인한다. 세부 payload와 계약은 `docs/edge-local-dev-guide.md`의 Raw text tool-call boundary smoke와 `agent-contract/outer/openai-compatible-api.md`를 따른다. - raw boundary smoke evidence는 tracked 문서가 아니라 ignored run 위치(`agent-test/runs/**`)나 code-review output path에 저장하고, 저장물에도 token 원문을 남기지 않는다. - provider-pool dispatch는 `in_flight >= capacity`인 provider를 후보에서 제외하고, 남은 후보 중 가장 낮은 `in_flight` 레벨을 먼저 선택한다. 같은 `in_flight` 레벨 안에서는 낮은 priority 값과 rotation으로 분산한다. -- dev-runtime managed capacity smoke는 `scripts/e2e-openai-managed-capacity-smoke.sh`를 사용한다. 같은 principal token으로 authenticated route alias를 먼저 조회하고, exact active route의 `resource_selector`, profile, upstream model과 일치하는 healthy provider snapshot 하나만 eligible pool로 본다. `ornith:35b`는 `onexplayer-lemonade=1`, `ornith-fast`는 `rtx5090-lemonade=1`로 각각 분리하며, 현재 projected route에서 제외된 두 provider의 capacity를 더하지 않는다. Chat/Responses를 alias별로 따로 실행해 각각 2개 요청에서 selected `in_flight=1`, `queued>=1`을 확인한다. -- managed smoke는 실제 emitted request JSON의 Unicode rune 수와 `runes/4 + runes/16` estimate를 계산하고 `long_context_threshold_tokens`와 대조한다. normal qualification은 `capacity`, 별도 long-context 시나리오는 `long_context_capacity`를 사용한다. 장문 반복/liveness payload로 normal capacity를 증명하거나 normal 요청으로 long slot을 증명하지 않는다. unprojected/legacy Laguna `laguna-s:2.1`과 Qwen `qwen3.6:35b`는 기존 model-group capacity `4`/`2` 기준을 유지한다. +- dev-runtime managed capacity smoke는 같은 principal token으로 authenticated route alias를 먼저 조회한다. explicit `resource_selector`는 정확히 그 provider 하나만, `default` selector는 route profile/upstream과 model group이 모두 일치하는 healthy provider만 eligible pool로 본다. 현재 `ornith:35b`는 GX10 `4` + OneXPlayer `3`, `ornith-fast`는 RTX5090 `1`이며 다른 provider capacity를 더하지 않는다. Chat/Responses를 alias별로 실행하고 eligible total `capacity + 1` 요청에서 provider별 exact peak와 queue를 확인한다. +- managed smoke는 실제 emitted request JSON의 Unicode rune 수와 `runes/4 + runes/16` estimate를 계산하고 `long_context_threshold_tokens`와 대조한다. normal qualification은 `capacity`, 별도 long-context 시나리오는 `long_context_capacity`를 사용한다. 장문 반복/liveness payload로 normal capacity를 증명하거나 normal 요청으로 long slot을 증명하지 않는다. stopped Laguna는 capacity가 없고, unprojected Qwen `qwen3.6:35b`만 기존 model-group capacity `2` 기준을 유지한다. - capacity smoke 완료 후 대상 provider의 `in_flight=0`, `queued=0` 회복을 확인한다. - long-context admission 시나리오(normal 10-way, mixed long/normal, all-long-slot-full)와 최종 회복 근거는 `agent-test/dev/long-context-admission-smoke.md`와 `scripts/e2e-long-context-admission-smoke.sh`를 사용한다. - `/v1/responses`와 `/v1/chat/completions`는 각각 짧은 입력에서 700~1200 token 수준의 구조화된 응답을 유도하고 provider-native thinking을 제한하며 50~100ms 간격으로 status를 polling한다. 각 요청의 HTTP 200, endpoint-native success terminal 정확히 1개, `[DONE]` 정확히 1개, selected provider의 eligible capacity 미초과, exact peak와 queue, 최종 0 회복을 별도 증거로 남긴다. +- 2026-08-15 `ornith:35b` default pool Chat 검증은 8/8 HTTP 200과 정상 terminal, GX10 `peak in_flight=4`, OneXPlayer `peak in_flight=3`, 합산 peak `7`, queue peak `2`, RTX peak `0`, final `0/0`으로 통과했다. 별도 Responses 검증은 5/5 HTTP 200과 각 `response.completed`는 확인됐지만 `[DONE]` terminal이 누락되어 fail-closed로 미통과 상태다. 이 미통과를 Responses capacity 완료 근거로 사용하지 않는다. - managed smoke invocation마다 ignored `agent-test/runs/**` 아래 mode `0700` unique directory와 immutable run id를 생성한다. manifest가 소유하고 현재 run/dispatch 이후 생성된 request/result/status만 판정하며 stale/missing/foreign-run body는 실패한다. tracked/review evidence에는 sanitized summary의 run id, script hash, alias, selected provider, endpoint, computed shape/class, terminal count, peak/queue/final counter, outcome만 허용하고 token/header, route/slot id, prompt, response body와 출력은 기록하지 않는다. - Qwen provider-pool smoke는 thinking/reasoning 텍스트가 포함될 수 있다. 추론 출력 자체를 실패로 보지 말고 HTTP 성공, model alias, final marker 포함 여부, provider node log/run count 증가로 판정한다. 응답 전체가 특정 token과 정확히 같은지 비교하는 strict exact-match는 이 profile의 기본 판정으로 쓰지 않는다. - Qwen provider를 agent/tool-call 용도로 검증할 때는 일반 chat smoke와 별도로 forced tool call, auto tool call, streaming `delta.tool_calls`, multi-turn tool result 후 최종 답변을 확인한다. raw native marker나 reasoning text가 assistant content로 새면 해당 model/runtime의 parser/template profile 미확정으로 보고한다. diff --git a/agent-test/dev/node-smoke.md b/agent-test/dev/node-smoke.md index 1b7cb206..896386b0 100644 --- a/agent-test/dev/node-smoke.md +++ b/agent-test/dev/node-smoke.md @@ -3,7 +3,7 @@ test_env: dev test_profile: node-smoke domain: node verification_type: smoke -last_rule_updated_at: 2026-08-02 +last_rule_updated_at: 2026-08-15 --- # node-smoke dev 테스트 @@ -60,24 +60,24 @@ dev-runtime의 실제 4-node 연결을 점검할 때는 원격 runner `ssh toki@ - GX10 vLLM node: `gx10-vllm-node` / `gx10-vllm` - SSH/user: `ssh toki@192.168.0.91` - provider endpoint: `http://192.168.0.91:8001/v1` - - served model: `laguna-s:2.1` + - served model: `ornith:35b` - capacity baseline: `4` - priority baseline: `1` - - runtime baseline: custom image `iop-vllm-laguna-s21:v0.25.1-cu130`, model `poolside/Laguna-S-2.1-NVFP4` revision `07614121b31898586430f189d27a25a0be310843`, `--served-model-name laguna-s:2.1`, `--max-model-len 262144`, `--max-num-seqs 4`, `--gpu-memory-utilization 0.70`, `--enable-prefix-caching`, `--enable-auto-tool-choice --tool-call-parser poolside_v1`, `--reasoning-parser poolside_v1`, `--chat-template /run/iop/laguna-s-2.1-thinking.jinja`, `--default-chat-template-kwargs '{"enable_thinking":true}'`, `--override-generation-config '{"temperature":0.7,"top_p":0.95}'` - - thinking template baseline: host `/home/toki/iop-gx10-vllm/laguna-s-2.1-thinking.jinja`를 container `/run/iop/laguna-s-2.1-thinking.jinja`에 read-only bind하고, generation prefix를 `\n`으로 둔다. stock `` prefix는 첫 생성 토큰으로 ``를 내보내 reasoning이 비는 현상이 재현됐다. - - speculative decoding baseline: quantization-matched `poolside/Laguna-S-2.1-DFlash-NVFP4` revision `723794750422b3efbf3a7b3af76dffb4ba035943`, `method=dflash`, `num_speculative_tokens=7` - - long-context admission baseline: observed GPU KV cache `373711` tokens, `long_context_capacity=1` (vLLM max concurrency `1.43x` for a 262144-token request) - - reasoning baseline: provider-native stream field is `reasoning`; Pi `openai-completions` accepts `reasoning_content`, `reasoning`, and `reasoning_text`. Pi Laguna model compat sends `chat_template_kwargs.enable_thinking` from the selected thinking level and keeps `preserve_thinking=true`. + - runtime baseline: image `vllm/vllm-openai:nightly-aarch64`, model `deepreinforce-ai/Ornith-1.0-35B-FP8` revision `1ab57ce0b44950e498a88756f40ad1ed4d0f30ca`, `--served-model-name ornith:35b`, `--max-model-len 262144`, `--max-num-seqs 4`, `--gpu-memory-utilization 0.50`, `--enable-prefix-caching`, `--enable-auto-tool-choice --tool-call-parser qwen3_xml`, `--reasoning-parser qwen3`, `--trust-remote-code`, `--language-model-only`, `--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20}'` + - long-context admission baseline: conservative `total_context_tokens=1048576`, `long_context_capacity=4`; current startup reports GPU KV cache `1173993` tokens and full-context concurrency `4.48x` for a 262144-token request. + - runtime ownership: `iop-vllm-ornith35b-fp8` is active on host port `8001`; stopped `iop-vllm-laguna-s21` is retained only for explicit rollback and contributes no provider capacity. + - reasoning baseline: `chat_template_kwargs.enable_thinking` toggles model-native reasoning; tool/reasoning parser는 `qwen3_xml`/`qwen3`를 사용한다. - workspace: `/home/toki/iop-gx10-vllm` - OneXPlayer Lemonade node: `onexplayer-lemonade-node` / `onexplayer-lemonade` - SSH/user: `ssh r0bin@192.168.0.59` - 접속 기준: 현재 작업 호스트에서 직접 SSH - - provider endpoint: `http://192.168.0.59:13305/v1` - - served model: `Ornith-1.0-35B-GGUF-llamacpp-tp1-Q5_K_M` - - capacity baseline: `1` (agent 장문 요청의 provider 독점 실행 기준) + - Node provider endpoint: `http://127.0.0.1:8001/v1` (OneX host-local llama backend) + - Lemonade lifecycle endpoint: `http://192.168.0.59:13305/v1` + - served model: `ornith:35b` (`--alias`); backend model은 `Ornith-1.0-35B-GGUF-llamacpp-tp1-Q5_K_M` + - capacity baseline: `3` - priority baseline: `2` - - load baseline: checkpoint `LordNeel/Ornith-1.0-35B-GGUF-llamacpp-tp1:Q5_K_M`, backend `vulkan`, ctx size `262144`, `llamacpp_args="--spec-type none -np 1 -cb -fa on -b 4096 -ub 1024 --kv-unified --temp 0.6 --top-p 0.95 --top-k 20"`, `save_options=true` - - long-context admission baseline: `total_context_tokens=262144`, `long_context_capacity=1` (IOP의 provider 직렬 admission과 llama-server의 `-np 1`을 일치시켜 agent turn별 prefix cache가 서로 다른 slot에 분산되지 않게 한다. `/slots`는 slot 1개와 `n_ctx=262144`를 보고한다) + - load baseline: checkpoint `LordNeel/Ornith-1.0-35B-GGUF-llamacpp-tp1:Q5_K_M`, backend `vulkan`, ctx size `524288`, `llamacpp_args="--spec-type none --alias ornith:35b -cb -fa on -b 4096 -ub 1024 --kv-unified --temp 0.6 --top-p 0.95 --top-k 20"`, `save_options=true` + - long-context admission baseline: `total_context_tokens=524288`, `long_context_capacity=2`. 고정 `-np`는 제거하며 `/slots`는 자동 slot 4개와 각 `n_ctx=262144`를 보고한다. IOP normal capacity는 보수적으로 `3`만 광고한다. - workspace: `C:/Users/r0bin/iop-field` - RTX5090 Lemonade node: `rtx5090-lemonade-node` / `rtx5090-lemonade` - SSH/user: `ssh iop-dev-rtx5090` @@ -92,7 +92,7 @@ dev-runtime의 실제 4-node 연결을 점검할 때는 원격 runner `ssh toki@ - process baseline: IOP Node와 Lemonade Server의 Windows 부팅 자동 실행은 비활성이다. 수동 운영 owner는 `C:/Users/r0bin/iop-field/remote-llm-toggle.ps1`이며 UP은 Lemonade Server -> Ornith load -> Node, DOWN은 Node -> Ornith unload -> Lemonade Server 순서다. IOP dev 배포는 Startup shortcut, Run entry, Task Scheduler, Windows service를 생성하지 않는다. - workspace: `C:/Users/r0bin/iop-field` -GX10은 Laguna S 2.1 NVFP4 + DFlash NVFP4를 사용한다. OneXPlayer와 RTX5090 Lemonade는 Ornith Q5 GGUF를 사용하고 runtime speculative decoding은 `--spec-type none`으로 끈다. provider family별 parser/template을 섞지 않는다. +GX10은 Ornith FP8 vLLM을 사용하고 OneXPlayer와 RTX5090 Lemonade는 Ornith Q5 GGUF를 사용한다. GX10의 stopped Laguna 컨테이너는 rollback 자산일 뿐 active provider나 capacity로 계산하지 않는다. provider runtime별 parser/template을 섞지 않는다. GX10은 Linux/ARM64 bootstrap, OneXPlayer와 RTX5090은 Windows native PowerShell bootstrap을 기본으로 한다. 두 Windows node는 현재 작업 호스트에서 직접 접속해 세팅하며, 원격 runner나 Edge host에서 proxy 실행하지 않는다. @@ -106,9 +106,7 @@ Edge config payload가 거부되어 Node가 `internal config error`로 종료되 RTX5090은 `ssh iop-dev-rtx5090`으로 batch 접속을 먼저 확인하고 Node process, toggle status의 `edge_connected`, Control Plane의 connected node 상태를 함께 확인한다. `RemoteLLM_mode.ahk`는 모니터 profile 적용 후 수동 toggle script를 한 번만 실행하며 Windows 부팅, Task Scheduler, 직접 Lemonade load 명령에 연결하지 않는다. 자동화 검증에서는 상태에 따라 반전되는 기본 Toggle 대신 `-Action Status|Up|Down`을 명시하고, `ready`는 Node의 Edge TCP `18084` 연결까지 포함해야 한다. Lemonade 복구 시에는 `0.0.0.0:13305`, 정확한 Ornith profile, llama slot `1`/`n_ctx=262144`를 확인한다. 비밀번호나 개인키 경로는 문서와 실행 로그에 기록하지 않는다. -GX10 Laguna 공식 vLLM baseline은 `temperature=0.7`, `top_p=0.95`, model `generation_config`의 `top_k=20`, `tool_call_parser=poolside_v1`, `reasoning_parser=poolside_v1`, `enable_thinking=true`다. 현재 dev는 stock generation prefix가 think block을 즉시 닫는 현상을 막기 위해 `\n` 로컬 템플릿을 추가한다. Pi의 `high` thinking은 `chat_template_kwargs.enable_thinking=true`로 전달하고 이전 reasoning block은 `preserve_thinking=true`로 재사용한다. think smoke는 `high`에서 `thinking_start`/`thinking_delta`/`thinking_end`, `off`에서 thinking event 0개와 최종 text를 확인하고, agentic multi-turn에서는 tool-call 전후 reasoning과 최종 text까지 확인한다. - -OneXPlayer/RTX5090 Ornith 공식 sampling baseline은 `temperature=0.6`, `top_p=0.95`, `top_k=20`이다. caller의 명시값이 provider 기본값보다 우선한다. +GX10/OneXPlayer/RTX5090 Ornith 공식 sampling baseline은 `temperature=0.6`, `top_p=0.95`, `top_k=20`이다. GX10 vLLM은 `qwen3_xml`/`qwen3` parser와 `chat_template_kwargs.enable_thinking`을 사용하고, Lemonade provider는 각 저장 recipe를 유지한다. caller의 명시값이 provider 기본값보다 우선한다. Qwen provider를 agent/tool-call 용도로 검증할 때는 일반 chat smoke와 별도로 forced tool call, auto tool call, streaming `delta.tool_calls`, multi-turn tool result 후 최종 답변을 확인한다. raw native marker나 reasoning text가 assistant content로 새면 해당 model/runtime의 parser/template profile 미확정으로 보고한다. Qwen runtime에는 Qwen 전용 parser/template 검증값만 사용한다. dev-corp Gemma 계열의 `tool_call_parser=gemma4`, `reasoning_parser=gemma4`, Gemma4 chat template/profile을 Qwen provider에 복사하지 않는다. diff --git a/agent-test/inventory-dev.yaml b/agent-test/inventory-dev.yaml index c89eece3..9dd39499 100644 --- a/agent-test/inventory-dev.yaml +++ b/agent-test/inventory-dev.yaml @@ -2,7 +2,7 @@ inventory_id: inventory-dev common_inventory: agent-test/inventory.yaml test_env: dev profile: dev-runtime-provider-pool -last_updated_at: "2026-08-13" +last_updated_at: "2026-08-15" source: remote_runner: @@ -66,7 +66,7 @@ build: node_windows_amd64: build/dev-runtime/bin/iop-node-windows-amd64.exe model: - alias: laguna-s:2.1 + alias: ornith:35b aliases: "claude-sonnet-5": observed_at: "2026-08-10" @@ -241,22 +241,20 @@ model: "ornith:35b": status: active_edge_model_group display_name: Ornith 1.0 35B - capacity_total: 2 + capacity_total: 7 providers: + - id: gx10-vllm + served_model: ornith:35b - id: onexplayer-lemonade - served_model: Ornith-1.0-35B-GGUF-llamacpp-tp1-Q5_K_M - - id: rtx5090-lemonade - served_model: Ornith-1.0-35B-GGUF-llamacpp-tp1-Q5_K_M + served_model: ornith:35b "laguna-s:2.1": - observed_at: "2026-07-24" - status: active_edge_model_group + observed_at: "2026-08-15" + status: stopped_rollback_only_no_active_provider_capacity display_name: Poolside Laguna S 2.1 context_window: 262144 default_max_tokens: 65536 - capacity_total: 4 - providers: - - id: gx10-vllm - served_model: laguna-s:2.1 + capacity_total: 0 + providers: [] ornith-fast: observed_at: "2026-07-18" status: active_edge_model_group @@ -288,7 +286,7 @@ model: - gx10-vllm-node - onexplayer-lemonade-node - rtx5090-lemonade-node - known_risk: ornith-fast and ornith:35b each admit against an independent model-group capacity counter, so one simultaneous request through each alias can target the same RTX5090 provider until provider-owned shared capacity is implemented. + isolation_note: RTX5090 is mapped only to ornith-fast and is not a candidate in the ornith:35b model group. provider_capacity_total: 2 provider_capacity_status: verified_with_edge_routing_and_provider_metrics_2026_07_24 context_window: 262144 @@ -305,10 +303,54 @@ model: default_max_tokens: 32768 min_max_tokens: 32768 active_edge_model_group: + observed_at: "2026-08-15" + id: ornith:35b + display_name: Ornith 1.0 35B + status: active_iop_edge_group + context_window: 262144 + default_max_tokens: 32768 + min_max_tokens: 16384 + provider_capacity_total: 7 + providers: + gx10-vllm: + runtime_type: vllm + endpoint: http://192.168.0.91:8001/v1 + served_model: ornith:35b + upstream_model: deepreinforce-ai/Ornith-1.0-35B-FP8 + quantization: fp8 + capacity: 4 + priority: 1 + total_context_tokens: 1048576 + long_context_capacity: 4 + onexplayer-lemonade: + runtime_type: lemonade + endpoint: http://127.0.0.1:8001/v1 + lifecycle_endpoint: http://192.168.0.59:13305/v1 + served_model: ornith:35b + backend_model: Ornith-1.0-35B-GGUF-llamacpp-tp1-Q5_K_M + quantization: Q5_K_M + capacity: 3 + priority: 2 + total_context_tokens: 524288 + long_context_capacity: 2 + route_note: The active ornith:35b credential route uses resource_selector default and canonical upstream ornith:35b. Its eligible pool is GX10 capacity 4 plus OneXPlayer capacity 3; RTX is isolated to ornith-fast. + smoke: + gx10_direct_chat_completions: passed + gx10_direct_responses: passed + gx10_direct_forced_tool_call: passed + gx10_provider_snapshot: healthy_capacity_4_long_capacity_4 + onex_runtime_restore: passed_ctx_524288_auto_4_slots_each_262144 + onex_provider_snapshot: healthy_capacity_3_long_capacity_2 + iop_route_projection: active_default_pool_revision_6 + iop_single_chat: passed_exact_gx10_iop_ok + iop_managed_chat_capacity: passed_http_200_5_of_5_peak_in_flight_4_peak_queued_1_final_0_0 + iop_managed_responses_capacity: failed_all_5_http_200_and_response_completed_but_done_terminal_missing + iop_managed_pool_chat_capacity: passed_8_of_8_http_200_and_terminal_gx_peak_4_onex_peak_3_pool_peak_7_queue_peak_2_rtx_peak_0_final_0_0 + previous_laguna_group_observation: observed_at: "2026-07-24" id: laguna-s:2.1 display_name: Poolside Laguna S 2.1 - status: active_iop_edge_group + status: historical_stopped_on_gx10_replaced_by_ornith_2026_08_15 openai_base_url_public: http://toki-labs.com:18083/v1 context_window: 262144 default_max_tokens: 65536 @@ -421,7 +463,7 @@ model: onexplayer-lemonade: runtime_type: lemonade endpoint: http://192.168.0.59:13305/v1 - served_model: Ornith-1.0-35B-GGUF-llamacpp-tp1-Q5_K_M + served_model: ornith:35b response_model: ornith-1.0-35b-Q5_K_M.gguf upstream_model: LordNeel/Ornith-1.0-35B-GGUF-llamacpp-tp1 upstream_url: https://huggingface.co/LordNeel/Ornith-1.0-35B-GGUF-llamacpp-tp1 @@ -441,7 +483,7 @@ model: rtx5090-lemonade: runtime_type: lemonade endpoint: http://192.168.0.111:13305/v1 - served_model: Ornith-1.0-35B-GGUF-llamacpp-tp1-Q5_K_M + served_model: ornith:35b response_model: ornith-1.0-35b-Q5_K_M.gguf upstream_model: LordNeel/Ornith-1.0-35B-GGUF-llamacpp-tp1 upstream_url: https://huggingface.co/LordNeel/Ornith-1.0-35B-GGUF-llamacpp-tp1 @@ -639,7 +681,7 @@ model: wall_sec: 44.01 total_completion_tokens: 3072 reasoning_policy: model_native_thinking - laguna_official_sampling: + previous_laguna_official_sampling_observation: source: https://huggingface.co/poolside/Laguna-S-2.1 source_section: vLLM and Controlling reasoning temperature: 0.7 @@ -672,8 +714,8 @@ model: thinking_format: chat-template enable_thinking_source: thinking.enabled preserve_thinking: true - previous_ornith_official_sampling_observation: - status: historical_gx10_replaced_by_laguna_s_2_1 + ornith_official_sampling: + status: active_gx10_restored_2026_08_15 source: https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B source_section: Quickstart and Chat Completions API examples temperature: 0.6 @@ -723,7 +765,8 @@ model: avoid: literal think tags and longer no-quote prohibitions avoid_reason: compared variants leaked reasoning into final content or ended with finish_reason=length full_capacity_smoke: - scope: full Ornith capacity is GX10 4 plus OneXPlayer 3 plus RTX5090 1 + scope: historical pre-alias-isolation Ornith qualification across GX10 4 plus OneXPlayer 3 plus RTX5090 1; this is not the current ornith:35b route capacity + current_route_note: ornith:35b now pools only GX10 4 plus OneXPlayer 3; RTX5090 is isolated to ornith-fast concurrent_requests: 9 chat_completions: http_200: 9 @@ -749,20 +792,18 @@ model: default_chat_template_kwargs: enable_thinking: true gx10_vllm: - model: laguna-s:2.1 - tool_call_parser: poolside_v1 - reasoning_parser: poolside_v1 - chat_template_host_path: /home/toki/iop-gx10-vllm/laguna-s-2.1-thinking.jinja - chat_template_container_path: /run/iop/laguna-s-2.1-thinking.jinja - thinking_generation_prefix: "\n" + model: ornith:35b + upstream_model: deepreinforce-ai/Ornith-1.0-35B-FP8 + tool_call_parser: qwen3_xml + reasoning_parser: qwen3 default_chat_template_kwargs: enable_thinking: true pi_chat_template_kwargs: enable_thinking: source: thinking.enabled preserve_thinking: true - tool_call_parser_status: Pi read-tool multi-turn smoke passed with reasoning before the tool call and after the tool result - separation_note: Keep Laguna poolside_v1, Qwen qwen/qwen3, and dev-corp Gemma parser/template profiles separate. + tool_call_parser_status: direct forced-tool smoke passed; Pi reasoning and multi-turn tool qualification remains governed by the current Ornith smoke record + separation_note: Keep active GX10 Ornith qwen3_xml/qwen3, mac Qwen qwen/qwen3, stopped Laguna poolside_v1, and dev-corp Gemma parser/template profiles separate. content_completion_policy: enforce_min_max_tokens: true reason: prevent reasoning-only completions from exhausting caller max_tokens before final content @@ -770,7 +811,8 @@ model: capacity_gate: providers at or above capacity are excluded from dispatch candidates selection_order: lower in_flight level wins among providers with available capacity same_in_flight_tiebreak: lower numeric priority first, then deterministic rotation - capacity_smoke: + previous_laguna_capacity_smoke: + status: historical_before_gx10_ornith_swap model: laguna-s:2.1 endpoints: - /v1/responses @@ -786,7 +828,7 @@ model: prompt_policy: long_reasoning_allowed exact_output_match: false latest_provider_snapshot: - observed_at: "2026-07-24" + observed_at: "2026-08-15" total_capacity: 10 all_nodes_connected: true providers: @@ -795,9 +837,9 @@ model: priority: 1 in_flight: 0 queued: 0 - long_context_capacity: 1 + long_context_capacity: 4 health: healthy - served_model: laguna-s:2.1 + served_model: ornith:35b mac-mlx-vllm: capacity: 2 priority: 2 @@ -813,7 +855,7 @@ model: queued: 0 long_context_capacity: 2 health: healthy - served_model: Ornith-1.0-35B-GGUF-llamacpp-tp1-Q5_K_M + served_model: ornith:35b rtx5090-lemonade: capacity: 1 priority: 0 @@ -822,7 +864,8 @@ model: long_context_capacity: 1 health: healthy served_model: Ornith-1.0-35B-GGUF-llamacpp-tp1-Q5_K_M - laguna_agent_tool_smoke: + previous_laguna_agent_tool_smoke: + status: historical_before_gx10_ornith_swap observed_at: "2026-07-24" model: laguna-s:2.1 request_shape: Pi openai-completions streaming with read tool @@ -873,7 +916,8 @@ model: mac-mlx-vllm: max_in_flight: 2 max_queued: 1 - final_content_smoke: + previous_laguna_final_content_smoke: + status: historical_before_gx10_ornith_swap model: laguna-s:2.1 endpoint: /v1/chat/completions concurrent_requests: 5 @@ -1040,72 +1084,67 @@ nodes: id: gx10-vllm type: vllm endpoint: http://192.168.0.91:8001/v1 - served_model: laguna-s:2.1 + served_model: ornith:35b capacity: 4 priority: 1 # Long-context admission policy (maps to edge.yaml nodes[].providers[]). - # Laguna NVFP4 runtime reports 373,570 GPU KV tokens and 1.43x maximum - # concurrency for full 262144-token requests, so only one long slot is allowed. - total_context_tokens: 373570 - long_context_capacity: 1 + # Ornith FP8 reports 1,173,993 GPU KV tokens and 4.48x maximum + # concurrency for full 262144-token requests. Admission remains capped at four. + total_context_tokens: 1048576 + long_context_capacity: 4 runtime: - container_name: iop-vllm-laguna-s21 - docker_image: iop-vllm-laguna-s21:v0.25.1-cu130 - docker_image_id: sha256:4eb2bb71c0ffc6ceb54dd3381030169f5e24ac1477fb1faa1ba90e57e6996331 - vllm_version: 0.25.1 + container_name: iop-vllm-ornith35b-fp8 + docker_image: vllm/vllm-openai:nightly-aarch64 + docker_image_id: sha256:a720df3e84a89d7db47a3b7a0511cb5b312e203fc4956f7493df248299267a6f + vllm_version: 0.23.1rc1.dev223+ga346d589f max_model_len: 262144 max_num_seqs: 4 - gpu_memory_utilization: 0.70 - default_max_tokens: 65536 + gpu_memory_utilization: 0.50 + default_max_tokens: 32768 enable_auto_tool_choice: true - tool_call_parser: poolside_v1 - reasoning_parser: poolside_v1 - chat_template_host_path: /home/toki/iop-gx10-vllm/laguna-s-2.1-thinking.jinja - chat_template_container_path: /run/iop/laguna-s-2.1-thinking.jinja - thinking_generation_prefix: "\n" - default_chat_template_kwargs: - enable_thinking: true - quantization: nvfp4 + tool_call_parser: qwen3_xml + reasoning_parser: qwen3 + trust_remote_code: true + language_model_only: true + quantization: fp8 dtype: bfloat16 override_generation_config: - temperature: 0.7 + temperature: 0.6 top_p: 0.95 - speculative_config: - model: poolside/Laguna-S-2.1-DFlash-NVFP4 - num_speculative_tokens: 7 - method: dflash + top_k: 20 + observed_gpu_kv_cache_tokens: 1173993 + observed_max_concurrency_for_262144_token_requests: 4.48 + stopped_rollback_container: iop-vllm-laguna-s21 direct_providers: - - id: laguna-direct - family: poolside_laguna + - id: ornith-gx10-direct + family: deepreinforce_ornith type: vllm endpoint: http://192.168.0.91:8001/v1 - served_model: laguna-s:2.1 - upstream_model: poolside/Laguna-S-2.1-NVFP4 - upstream_url: https://huggingface.co/poolside/Laguna-S-2.1-NVFP4 - upstream_revision: 07614121b31898586430f189d27a25a0be310843 - status: direct_pi_provider_not_registered_runtime_shares_current_iop_laguna_gx10_provider + served_model: ornith:35b + upstream_model: deepreinforce-ai/Ornith-1.0-35B-FP8 + upstream_url: https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B-FP8 + upstream_revision: 1ab57ce0b44950e498a88756f40ad1ed4d0f30ca + status: active_runtime_shares_current_iop_ornith_gx10_provider pi_model_parameters: context_window: 262144 - max_tokens: 65536 + max_tokens: 32768 reasoning: true thinking_format: chat-template preserve_thinking: true runtime: - container_name: iop-vllm-laguna-s21 - docker_image: iop-vllm-laguna-s21:v0.25.1-cu130 - docker_image_id: sha256:4eb2bb71c0ffc6ceb54dd3381030169f5e24ac1477fb1faa1ba90e57e6996331 - vllm_version: 0.25.1 + container_name: iop-vllm-ornith35b-fp8 + docker_image: vllm/vllm-openai:nightly-aarch64 + docker_image_id: sha256:a720df3e84a89d7db47a3b7a0511cb5b312e203fc4956f7493df248299267a6f + vllm_version: 0.23.1rc1.dev223+ga346d589f port_mapping: 0.0.0.0:8001->8000 hf_home: /models/.cache/huggingface - hf_cache_snapshot_path: /models/.cache/huggingface/hub/models--poolside--Laguna-S-2.1-NVFP4/snapshots/07614121b31898586430f189d27a25a0be310843 - draft_cache_snapshot_path: /models/.cache/huggingface/hub/models--poolside--Laguna-S-2.1-DFlash-NVFP4/snapshots/723794750422b3efbf3a7b3af76dffb4ba035943 - quantization: nvfp4 + hf_cache_snapshot_path: /models/.cache/huggingface/hub/models--deepreinforce-ai--Ornith-1.0-35B-FP8/snapshots/1ab57ce0b44950e498a88756f40ad1ed4d0f30ca + quantization: fp8 vllm_quantization_backend: compressed-tensors dtype: bfloat16 max_model_len: 262144 max_num_seqs: 4 - max_num_scheduled_tokens: 2024 - gpu_memory_utilization: 0.70 + gpu_memory_utilization: 0.50 enable_prefix_caching: true enable_chunked_prefill: true tensor_parallel_size: 1 @@ -1115,22 +1154,15 @@ nodes: kv_cache_dtype: float8_e4m3fn enforce_eager: false enable_auto_tool_choice: true - tool_call_parser: poolside_v1 - reasoning_parser: poolside_v1 + tool_call_parser: qwen3_xml + reasoning_parser: qwen3 reasoning_parser_enable_in_reasoning: false - chat_template_host_path: /home/toki/iop-gx10-vllm/laguna-s-2.1-thinking.jinja - chat_template_container_path: /run/iop/laguna-s-2.1-thinking.jinja - thinking_generation_prefix: "\n" - default_chat_template_kwargs: - enable_thinking: true - attention_backend: FLASHINFER + trust_remote_code: true + language_model_only: true generation_config_override: - temperature: 0.7 + temperature: 0.6 top_p: 0.95 - speculative_config: - model: poolside/Laguna-S-2.1-DFlash-NVFP4 - num_speculative_tokens: 7 - method: dflash + top_k: 20 docker_runtime: network_mode: bridge ipc_mode: host @@ -1139,12 +1171,11 @@ nodes: gpu_device_request: all bind_mounts: /home/toki/Data/models: /models - /home/toki/Data/vllm/cache/laguna-s21: /root/.cache - /home/toki/iop-gx10-vllm/laguna-s-2.1-thinking.jinja: /run/iop/laguna-s-2.1-thinking.jinja - available_kv_cache_memory_gib: 13.08 - gpu_kv_cache_tokens: 373711 - max_concurrency_for_262144_token_requests: 1.43 - thinking_note: The poolside_v1 parser emits reasoning in the provider-native reasoning field. The local chat template uses a trailing newline after the generation-prefix think tag; Pi high emits thinking deltas and Pi off emits no thinking event. + /home/toki/Data/vllm/templates: /templates + available_kv_cache_memory_gib: 23.01 + gpu_kv_cache_tokens: 1173993 + max_concurrency_for_262144_token_requests: 4.48 + thinking_note: The qwen3 reasoning parser and qwen3_xml tool parser are active; chat_template_kwargs.enable_thinking controls reasoning per request. - id: onexplayer-lemonade-node alias: onexplayer-lemonade role: lemonade-provider @@ -1155,11 +1186,13 @@ nodes: provider: id: onexplayer-lemonade type: lemonade - endpoint: http://192.168.0.59:13305/v1 - served_model: Ornith-1.0-35B-GGUF-llamacpp-tp1-Q5_K_M + endpoint: http://127.0.0.1:8001/v1 + lifecycle_endpoint: http://192.168.0.59:13305/v1 + served_model: ornith:35b + backend_model: Ornith-1.0-35B-GGUF-llamacpp-tp1-Q5_K_M served_model_response: ornith-1.0-35b-Q5_K_M.gguf previous_served_model: Qwen3.6-35B-A3B-MTP-GGUF - status: active_ornith_q5_current_iop_ornith_provider_pool_member + status: active_ornith_q5_current_iop_ornith_35b_pool_member qwen_reenable_profile: status: not_loaded_ornith_q5_active model_name: Qwen3.6-35B-A3B-MTP-GGUF @@ -1178,14 +1211,13 @@ nodes: long_context_capacity: 2 validation_required_after_load: true validation_basis: This inactive Qwen profile requires separate validation before re-enable; do not derive its parallelism from the current single-slot Ornith profile. - capacity: 1 + capacity: 3 priority: 2 # Long-context admission policy (maps to edge.yaml nodes[].providers[]). - # IOP admits one Ornith request at a time, so llama.cpp also uses one slot. - # This keeps agent-turn prefix cache on the same slot instead of scattering - # it across auto slots. Do NOT raise ctx_size above 262144. - total_context_tokens: 262144 - long_context_capacity: 1 + # llama.cpp auto-provisions four 262144-token slots from ctx_size 524288; + # IOP advertises a conservative normal capacity of three and two long slots. + total_context_tokens: 524288 + long_context_capacity: 2 load: endpoint: http://192.168.0.59:13305/v1/load model_name: Ornith-1.0-35B-GGUF-llamacpp-tp1-Q5_K_M @@ -1197,13 +1229,13 @@ nodes: gguf_file_size_bytes: 24729130848 cache_snapshot_path: C:/Users/r0bin/.cache/huggingface/hub/models--LordNeel--Ornith-1.0-35B-GGUF-llamacpp-tp1/snapshots/c50d5d4407f70e43208dee836c66bb8a05c1be91 backend: vulkan - ctx_size: 262144 - llamacpp_args: "--spec-type none -np 1 -cb -fa on -b 4096 -ub 1024 --kv-unified --temp 0.6 --top-p 0.95 --top-k 20" - observed_process_args: "--ctx-size 262144 --port 8001 --jinja --context-shift --keep 16 --reasoning-format auto --no-webui --no-mmap -ngl 99 --kv-unified --spec-type none --temp 0.6 --top-k 20 --top-p 0.95 -b 4096 -cb -fa on -np 1 -ub 1024" - observed_total_slots: 1 - observed_total_ctx_size: 262144 + ctx_size: 524288 + llamacpp_args: "--spec-type none --alias ornith:35b -cb -fa on -b 4096 -ub 1024 --kv-unified --temp 0.6 --top-p 0.95 --top-k 20" + observed_process_args: "--ctx-size 524288 --port 8001 --jinja --context-shift --keep 16 --reasoning-format auto --no-webui --no-mmap -ngl 99 --kv-unified --spec-type none --alias ornith:35b --temp 0.6 --top-k 20 --top-p 0.95 -b 4096 -cb -fa on -ub 1024" + observed_total_slots: 4 + observed_total_ctx_size: 524288 observed_slot_n_ctx: 262144 - context_per_slot_note: -np 1 matches the IOP provider capacity and preserves one reusable prefix-cache lineage across agent turns. + context_per_slot_note: Fixed -np is removed; the runtime reports four unified-KV slots while IOP caps normal admission at three. residency: observed_at: "2026-07-12" process_resident_policy: keep llama-server loaded until explicit lemonade unload or service/process restart @@ -1252,7 +1284,7 @@ nodes: low_output_budget_note: max_tokens 128 and one Korean 512-token probe ended in reasoning_content with empty final content; keep Pi/Edge maxTokens at 32768 or use at least 1024 for short smoke prompts. windows_process_start: ownership: user_managed_manual_toggle - observed_at: "2026-07-26" + observed_at: "2026-08-15" boot_autostart: status: disabled startup_shortcuts_matching_iop_or_lemonade: absent @@ -1263,15 +1295,16 @@ nodes: recoverable_backup: C:/Users/r0bin/iop-field/IOP-OnexNode.pre-manual-lemonade-20260725T234232Z.xml manual_remote_llm_toggle: script: C:/Users/r0bin/iop-field/remote-llm-toggle.ps1 - script_sha256: d1729c7928978d1d547f8ed20f86ef25a75cc4e0228857e332c0530541abd9c7 + script_sha256: 0bbda344b0a3991009a57197acf6b416d09b39587607986b105ef46e48e8132c + runtime_script_sha256: f2630af285d20c737bb9d2af45f579226ca1da6fcd382c84486b7bc6be39a36f log: C:/Users/r0bin/iop-field/onex-remote-llm-toggle.log default_action: toggle_by_complete_stack_readiness process_start: Win32_Process.Create readiness_requires: - LemonadeServer.exe running and health status ok - - exact Ornith model with Vulkan ctx_size 262144 and -np 1 profile + - exact Ornith model with Vulkan ctx_size 524288, API alias ornith:35b, and no fixed -np - public listener 0.0.0.0:13305 - - one llama slot with n_ctx 262144 + - four llama slots with n_ctx 262144 each - iop-node.exe running and established Edge TCP connection to port 18084 up_sequence: - start LemonadeServer.exe @@ -1320,7 +1353,35 @@ nodes: endpoint: http://192.168.0.111:13305/v1 served_model: Ornith-1.0-35B-GGUF-llamacpp-tp1-Q5_K_M served_model_response: ornith-1.0-35b-Q5_K_M.gguf - status: active_ornith_q5_current_iop_ornith_provider_pool_member + status: active_ornith_q5_ornith_fast_only_not_ornith_35b_candidate + qwen38_standby_resource: + status: downloaded_not_loaded_not_projected_to_dev_route + model_name: Qwen3.8-27B-GGUF-Q4_K_M + checkpoint: ggml-org/Qwen3.8-27B-GGUF:Q4_K_M + hf_revision: 0669b98607d47046c7c2b3f801011d54a08cfccf + gguf_file: Qwen3.8-27B-Q4_K_M.gguf + gguf_file_size_bytes: 18973870432 + cache_snapshot_path: D:/Models/models--ggml-org--Qwen3.8-27B-GGUF/snapshots/0669b98607d47046c7c2b3f801011d54a08cfccf + saved_recipe_options: + ctx_size: 262144 + llamacpp_backend: cuda + llamacpp_args: "--chat-template-kwargs '{\"preserve_thinking\":true}' --kv-unified --min-p 0.00 --repeat-penalty 1.0 --spec-type none --temp 0.6 --top-k 20 --top-p 0.95 -b 512 -cb -ctk q8_0 -ctv q8_0 -fa on -np 1 -ub 256" + kv_cache_key_type: q8_0 + kv_cache_value_type: q8_0 + slots: 1 + activation_policy: + current_loaded_model: Ornith-1.0-35B-GGUF-llamacpp-tp1-Q5_K_M + explicit_manual_load_required: true + dev_route_and_capacity_membership: absent + note: Keep this resource out of dev provider eligibility and capacity until its own route projection and qualification are explicitly approved. + direct_benchmark: + observed_at: "2026-08-15" + thinking_enabled: false + prompt_tokens: 54 + completion_tokens_per_run: 1024 + sequential_runs: 3 + end_to_end_tok_s: 67.12 + server_decode_tok_s: 68.24 capacity: 1 priority: 0 total_context_tokens: 262144 diff --git a/docs/edge-local-dev-guide.md b/docs/edge-local-dev-guide.md index d0408a8c..a6c18244 100644 --- a/docs/edge-local-dev-guide.md +++ b/docs/edge-local-dev-guide.md @@ -150,15 +150,15 @@ http://:18081/v1 `/v1/chat/completions` 요청은 OpenAI-compatible field를 기본으로 사용한다. Provider-pool pure `passthrough`는 selected provider가 지원하는 OpenAI-compatible 표준 field와 provider extension field를 보존해야 하며, `chat_template_kwargs` 같은 provider-native option을 IOP allowlist로 막지 않는다. `think`, `reasoning_effort`, `thinking_token_budget`, `include_reasoning`은 normalized backend 또는 provider별 차이를 보완하기 위한 IOP 확장 field다. -- dev GX10 `laguna-s:2.1` 공식 sampling baseline은 Laguna S 2.1 모델 카드의 vLLM recipe를 따른다: `temperature=0.7`, `top_p=0.95`, model `generation_config`의 `top_k=20`. -- GX10 공식 근거: https://huggingface.co/poolside/Laguna-S-2.1 -- dev OneXPlayer/RTX5090 `ornith:35b` sampling baseline은 `temperature=0.6`, `top_p=0.95`, `top_k=20`을 유지한다. 공식 예제에 없는 non-neutral `repeat_penalty`는 임의로 추가하지 않는다. RTX5090 saved recipe의 `--repeat-penalty 1.0`과 `--min-p 0.00`은 출력을 바꾸지 않는 neutral runtime serialization로만 유지한다. +- dev GX10/OneXPlayer/RTX5090 Ornith sampling baseline은 `temperature=0.6`, `top_p=0.95`, `top_k=20`을 유지한다. +- dev OneXPlayer `ornith:35b`와 RTX5090 `ornith-fast` sampling baseline은 `temperature=0.6`, `top_p=0.95`, `top_k=20`을 유지한다. 공식 예제에 없는 non-neutral `repeat_penalty`는 임의로 추가하지 않는다. RTX5090 saved recipe의 `--repeat-penalty 1.0`과 `--min-p 0.00`은 출력을 바꾸지 않는 neutral runtime serialization로만 유지한다. - Ornith 공식 근거: https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B - 이 값은 caller가 sampling field를 생략했을 때 쓰는 provider 기본값이다. caller가 명시한 sampling field가 있으면 요청값이 우선한다. -- dev GX10 vLLM은 `poolside/Laguna-S-2.1-NVFP4`와 quantization-matched `Laguna-S-2.1-DFlash-NVFP4`를 사용하고 `--tool-call-parser poolside_v1`, `--reasoning-parser poolside_v1`, `--chat-template /run/iop/laguna-s-2.1-thinking.jinja`, `--default-chat-template-kwargs '{"enable_thinking":true}'`, `--override-generation-config '{"temperature":0.7,"top_p":0.95}'`로 설정한다. -- GX10 host의 `/home/toki/iop-gx10-vllm/laguna-s-2.1-thinking.jinja`는 container에 read-only bind한다. stock generation prefix ``가 첫 생성 토큰 ``를 유도해 reasoning이 비는 현상이 재현됐으므로, dev template은 `\n`을 사용한다. +- dev GX10 vLLM은 `deepreinforce-ai/Ornith-1.0-35B-FP8`을 `ornith:35b`, context `262144`, slots `4`, `--gpu-memory-utilization 0.50`, `--tool-call-parser qwen3_xml`, `--reasoning-parser qwen3`, `--language-model-only`로 제공한다. `total_context_tokens=1048576`, `long_context_capacity=4`를 보수적 admission 기준으로 사용한다. +- GX10의 `iop-vllm-laguna-s21` 컨테이너와 Laguna/DFlash 다운로드는 stopped rollback 자산으로만 보존하며 active dev provider나 capacity로 계산하지 않는다. - dev OneXPlayer/RTX5090 Lemonade는 해당 Ornith recipe의 `llamacpp_args`에 `--temp 0.6 --top-p 0.95 --top-k 20`을 저장한다. RTX5090 수동 toggle은 이 sampling 값과 함께 `preserve_thinking=true`, unified KV, Q8 KV, neutral min-p/repeat-penalty를 exact profile로 검증한다. -- Laguna reasoning은 provider-native `reasoning` field로 나가며 Pi `openai-completions`가 이를 `thinking_delta`로 소비한다. Pi Laguna profile은 thinking level을 `chat_template_kwargs.enable_thinking`으로 전달하고 `preserve_thinking=true`로 이전 assistant reasoning을 유지한다. +- RTX5090에는 `ggml-org/Qwen3.8-27B-GGUF:Q4_K_M` 체크포인트와 CUDA/262144 context/Q8 KV/단일 슬롯 recipe가 standby 자원으로 저장되어 있다. 현재 로드 모델과 dev route/capacity의 활성 대상은 Ornith이며, Qwen3.8은 명시적으로 로드하고 별도 route 정합성 검증을 마치기 전까지 dev provider 자원으로 계산하지 않는다. +- GX10 Ornith reasoning은 `chat_template_kwargs.enable_thinking`으로 제어하고 vLLM `qwen3` reasoning parser 결과를 사용한다. tool-call은 `qwen3_xml` parser로 정규화한다. - 출력 smoke는 같은 요청을 Pi `high`와 `off`로 대조한다. `high`에서는 `thinking_start`/`thinking_delta`/`thinking_end`와 최종 text가, `off`에서는 thinking event 0개와 최종 text가 나와야 한다. agentic multi-turn에서는 tool-call 전후 reasoning, tool result, 최종 text까지 확인한다. - 현재 dev-corp provider-pool device mapping은 `gemma4:26b` -> Mac Studio provider capacity `5`, `ornith:35b` -> DGX Spark 01/02 provider 합산 capacity `8`이다. 세부 endpoint와 runtime args는 `agent-test/inventory-dev-corp.yaml`을 기준으로 한다. - dev-corp provider-pool 안정 smoke: 기본 smoke에서는 `think`, `reasoning_effort`, `thinking_token_budget`을 생략하고 현재 provider 기본값을 유지한다. Provider-native passthrough를 검증할 때는 selected provider가 직접 지원하는 field를 그대로 보낸다. @@ -170,12 +170,14 @@ http://:18081/v1 ### Managed route-qualified capacity smoke -Managed mode의 capacity는 전역 model catalog나 같은 upstream model을 제공하는 모든 provider의 합이 아닙니다. 먼저 OpenAI ingress와 같은 principal token으로 active route alias를 조회하고, route의 `resource_selector`, profile, upstream model과 일치하는 현재 healthy provider snapshot만 eligible capacity로 계산합니다. 현재 dev projection에서는 `ornith:35b`가 `onexplayer-lemonade`, `ornith-fast`가 `rtx5090-lemonade`에 각각 고정되므로 두 capacity를 합치거나 alias를 한 batch에 섞지 않습니다. +Managed mode의 capacity는 전역 model catalog나 같은 upstream model을 제공하는 모든 provider의 무조건적인 합이 아닙니다. 먼저 OpenAI ingress와 같은 principal token으로 active route alias를 조회합니다. explicit `resource_selector`는 그 provider 하나만 선택하고, `default` selector는 route의 profile/upstream model과 model group이 모두 일치하는 healthy provider를 풀링합니다. 현재 `ornith:35b`는 `default` selector로 GX10 `4`와 OneXPlayer `3`, 합계 `7`을 사용하며 RTX5090은 후보가 아니라 `ornith-fast` 전용입니다. -정상 capacity 검증은 `scripts/e2e-openai-managed-capacity-smoke.sh`로 Chat과 Responses를 한 endpoint·route씩 실행합니다. 스크립트는 실제 emitted JSON의 Unicode rune 수와 `runes/4 + runes/16` estimate를 계산해 `normal`임을 확인하고, selected provider의 `capacity + 1`만 전송합니다. 별도 long-context/repeat smoke는 `long_context_capacity`를 사용하며 normal-capacity 완료 근거를 대체하지 않습니다. +정상 capacity 검증은 Chat과 Responses를 한 endpoint·route씩 실행합니다. 실제 emitted JSON의 Unicode rune 수와 `runes/4 + runes/16` estimate를 계산해 `normal`임을 확인합니다. explicit selector는 selected provider의 `capacity + 1`, `default` pool은 eligible provider capacity 합계에 1을 더한 요청을 전송하고 provider별 peak가 각 capacity에 도달했는지 확인합니다. 별도 long-context/repeat smoke는 `long_context_capacity`를 사용하며 normal-capacity 완료 근거를 대체하지 않습니다. 성공 조건은 모든 요청 HTTP 200, Chat의 finish terminal과 Responses의 `response.completed` 각각 정확히 1개, stream별 `[DONE]` 정확히 1개, selected provider peak가 eligible capacity와 같고 queue가 1 이상인 상태, 최종 `in_flight=0`/`queued=0`입니다. route mismatch, context class mismatch, non-selected capacity 포함, terminal 누락/중복, status 관측 누락은 fail-closed입니다. +2026-08-15 `ornith:35b` default pool Chat capacity smoke는 8/8 HTTP 200과 정상 terminal, GX10 peak `4`, OneXPlayer peak `3`, 합산 peak `7`, queue peak `2`, RTX peak `0`, final `0/0`으로 통과했습니다. 별도 Responses smoke는 5/5 HTTP 200과 각 `response.completed`까지 수신했지만 stream별 `[DONE]`이 없어 미통과이므로 Responses stream terminal은 후속 확인 대상으로 유지합니다. + 각 invocation은 ignored `agent-test/runs/**` 아래 mode `0700` unique directory를 생성하고 current-run manifest가 소유한 request/result/status만 판정합니다. Raw route DTO, token/header, route/slot id, prompt, request/response body와 모델 출력은 해당 디렉터리 밖으로 복사하지 않습니다. tracked evidence와 code review에는 allowlist된 sanitized summary의 run id, script hash, route alias, selected provider, endpoint, computed request shape/class, terminal count, peak/queue/final counter와 outcome만 남깁니다. 예시 (dev-corp `gemma4:26b` provider-pool non-stream 측정, think 생략):