From 9cff91f8bfdc92610f8d3b17804c74d19a255752 Mon Sep 17 00:00:00 2001 From: leedongmyun Date: Tue, 14 Jul 2026 06:14:00 +0900 Subject: [PATCH] =?UTF-8?q?docs(benchmark):=20Ornith=20think=20=EC=9D=91?= =?UTF-8?q?=EB=8B=B5=20=EB=B9=84=EA=B5=90=ED=91=9C=EB=A5=BC=20=EC=B6=94?= =?UTF-8?q?=EA=B0=80=ED=95=9C=EB=8B=A4?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- ...NK_INITIAL_RESPONSE_COMPARISON_20260713.md | 49 +++++++++++++++++++ 1 file changed, 49 insertions(+) create mode 100644 ORNITH_THINK_INITIAL_RESPONSE_COMPARISON_20260713.md diff --git a/ORNITH_THINK_INITIAL_RESPONSE_COMPARISON_20260713.md b/ORNITH_THINK_INITIAL_RESPONSE_COMPARISON_20260713.md new file mode 100644 index 0000000..b48db9e --- /dev/null +++ b/ORNITH_THINK_INITIAL_RESPONSE_COMPARISON_20260713.md @@ -0,0 +1,49 @@ +# Mac Studio Ornith Think True/Off 초기 응답 속도 비교 + +측정 목적: Mac Studio Ornith `ornith:35b` 런타임에서 thinking 활성 상태와 thinking 비활성 상태의 사용자 가시 첫 content 응답 속도를 비교한다. + +## 측정 기준 + +- 측정일: 2026-07-13 KST +- Host: `dc-devui-MacStudio.local` +- Endpoint: `http://127.0.0.1:8007/v1` +- Model alias: `ornith:35b` +- API: OpenAI-compatible streaming `/v1/chat/completions` +- Prompt: `Benchmark generation task. Produce exactly 8 short numbered Korean lines about reliable agent runtime operations.` +- 요청 공통값: `max_tokens=1024`, `temperature=0.2`, `top_p=0.95` +- 원본 측정 문서: `agent-test/dev-corp/mac-studio-ornith-think-on-baseline-20260713.md` + +## 모드 정의 + +- Think true: upstream `mlx_lm.server` 기본값 `enable_thinking=true`; 요청에는 별도 think-off field를 넣지 않았다. +- Think off: request-level `chat_template_kwargs: {"enable_thinking": false}`를 사용한 실효 think-off 측정값이다. +- 주의: alias proxy의 plain `think:false`는 당시 `chat_template_args`로만 매핑되어 실효 think-off가 아니었다. 실제 비교에는 `chat_template_kwargs.enable_thinking=false` 결과를 사용한다. + +## 초기 content 응답 속도 비교 + +`첫 content`는 reasoning delta가 아니라 사용자가 실제 답변으로 보는 첫 non-empty `content` delta가 도착한 시점이다. + +| 동시 호출 수 | Think true 첫 content (s) | Think off 첫 content (s) | 단축 시간 (s) | 초기 content 응답 배율 | +|---:|---:|---:|---:|---:| +| 1 | 17.830 | 0.208 | 17.622 | 85.9x | +| 2 | 22.193 | 0.333 | 21.860 | 66.7x | +| 3 | 26.260 | 0.457 | 25.803 | 57.5x | +| 4 | 30.638 | 0.580 | 30.058 | 52.9x | +| 5 | 38.788 | 0.703 | 38.085 | 55.2x | + +## Content 처리량 비교 + +| 동시 호출 수 | Think true content tok/s per call | Think off content tok/s per call | Think true total content tok/s | Think off total content tok/s | +|---:|---:|---:|---:|---:| +| 1 | 53.00 | 52.77 | 5.25 | 47.74 | +| 2 | 44.44 | 44.81 | 3.88 | 79.15 | +| 3 | 38.26 | 37.79 | 4.92 | 98.27 | +| 4 | 32.34 | 32.47 | 5.62 | 111.20 | +| 5 | 25.64 | 25.82 | 5.55 | 111.03 | + +## 요약 + +- Think true에서는 reasoning이 먼저 길게 생성되어 첫 content가 `17.830s`에서 `38.788s` 사이에 도착했다. +- Think off에서는 reasoning token이 `0`으로 측정되었고 첫 content가 `0.208s`에서 `0.703s` 사이에 도착했다. +- 단일 호출 decode 속도는 큰 차이가 아니며, 핵심 개선은 reasoning-before-content 지연 제거다. +- 동시 호출 4~5개에서 think-off total content throughput은 약 `111 tok/s` 수준으로 측정되었다.