Remove ORNITH_THINK comparison file
This commit is contained in:
parent
f14d3a201e
commit
91c10c1c36
1 changed files with 0 additions and 49 deletions
|
|
@ -1,49 +0,0 @@
|
|||
# Mac Studio Ornith Think True/Off 초기 응답 속도 비교
|
||||
|
||||
측정 목적: Mac Studio Ornith `ornith:35b` 런타임에서 thinking 활성 상태와 thinking 비활성 상태의 사용자 가시 첫 content 응답 속도를 비교한다.
|
||||
|
||||
## 측정 기준
|
||||
|
||||
- 측정일: 2026-07-13 KST
|
||||
- Host: `dc-devui-MacStudio.local`
|
||||
- Endpoint: `http://127.0.0.1:8007/v1`
|
||||
- Model alias: `ornith:35b`
|
||||
- API: OpenAI-compatible streaming `/v1/chat/completions`
|
||||
- Prompt: `Benchmark generation task. Produce exactly 8 short numbered Korean lines about reliable agent runtime operations.`
|
||||
- 요청 공통값: `max_tokens=1024`, `temperature=0.2`, `top_p=0.95`
|
||||
- 원본 측정 문서: `agent-test/dev-corp/mac-studio-ornith-think-on-baseline-20260713.md`
|
||||
|
||||
## 모드 정의
|
||||
|
||||
- Think true: upstream `mlx_lm.server` 기본값 `enable_thinking=true`; 요청에는 별도 think-off field를 넣지 않았다.
|
||||
- Think off: request-level `chat_template_kwargs: {"enable_thinking": false}`를 사용한 실효 think-off 측정값이다.
|
||||
- 주의: alias proxy의 plain `think:false`는 당시 `chat_template_args`로만 매핑되어 실효 think-off가 아니었다. 실제 비교에는 `chat_template_kwargs.enable_thinking=false` 결과를 사용한다.
|
||||
|
||||
## 초기 content 응답 속도 비교
|
||||
|
||||
`첫 content`는 reasoning delta가 아니라 사용자가 실제 답변으로 보는 첫 non-empty `content` delta가 도착한 시점이다.
|
||||
|
||||
| 동시 호출 수 | Think true 첫 content (s) | Think off 첫 content (s) | 단축 시간 (s) | 초기 content 응답 배율 |
|
||||
|---:|---:|---:|---:|---:|
|
||||
| 1 | 17.830 | 0.208 | 17.622 | 85.9x |
|
||||
| 2 | 22.193 | 0.333 | 21.860 | 66.7x |
|
||||
| 3 | 26.260 | 0.457 | 25.803 | 57.5x |
|
||||
| 4 | 30.638 | 0.580 | 30.058 | 52.9x |
|
||||
| 5 | 38.788 | 0.703 | 38.085 | 55.2x |
|
||||
|
||||
## Content 처리량 비교
|
||||
|
||||
| 동시 호출 수 | Think true content tok/s per call | Think off content tok/s per call | Think true total content tok/s | Think off total content tok/s |
|
||||
|---:|---:|---:|---:|---:|
|
||||
| 1 | 53.00 | 52.77 | 5.25 | 47.74 |
|
||||
| 2 | 44.44 | 44.81 | 3.88 | 79.15 |
|
||||
| 3 | 38.26 | 37.79 | 4.92 | 98.27 |
|
||||
| 4 | 32.34 | 32.47 | 5.62 | 111.20 |
|
||||
| 5 | 25.64 | 25.82 | 5.55 | 111.03 |
|
||||
|
||||
## 요약
|
||||
|
||||
- Think true에서는 reasoning이 먼저 길게 생성되어 첫 content가 `17.830s`에서 `38.788s` 사이에 도착했다.
|
||||
- Think off에서는 reasoning token이 `0`으로 측정되었고 첫 content가 `0.208s`에서 `0.703s` 사이에 도착했다.
|
||||
- 단일 호출 decode 속도는 큰 차이가 아니며, 핵심 개선은 reasoning-before-content 지연 제거다.
|
||||
- 동시 호출 4~5개에서 think-off total content throughput은 약 `111 tok/s` 수준으로 측정되었다.
|
||||
Loading…
Reference in a new issue