12 KiB
Mac Studio Ornith Think-On Baseline
측정 목적: Mac Studio Ornith runtime의 think off 실험 전 기준 성능을 보존한다.
Runtime Identity
- 측정 시각: 2026-07-13 17:47:59 KST
- Host:
dc-devui-MacStudio.local - Endpoint:
http://127.0.0.1:8007/v1 - Model alias:
ornith:35b - Current runtime path:
ornith35b_alias_proxy.pyon0.0.0.0:8007->mlx_lm.serveron127.0.0.1:8008 - Upstream model path:
/Users/dc_dev/.cache/huggingface/manual/mlx-community/Ornith-1.0-35B-8bit - Runtime note: the active
vllm_mlx.cli serveprocess on this host is currently Gemma4 on8004; this Ornith baseline uses the current Mac Studio Ornith alias path above. - Thinking mode: current default. Upstream
mlx_lm.serverwas started with--chat-template-args '{"enable_thinking":true}'; benchmark requests did not passthink=false. - API shape: OpenAI-compatible streaming
/v1/chat/completions - Token counting:
Qwen2Tokenizerover concatenated streamingreasoning,reasoning_content, andcontentdeltas.
Request Parameters
max_tokens:256temperature:0.2top_p:0.95- Warmup: one short warmup request, excluded from the table.
- Prompt:
Benchmark generation task. Do not use tools. Produce a compact Korean report with exactly 24 numbered lines about reliable agent runtime operations. Keep each line short, but continue until the list is complete.
Metric Definitions
avg_ttft_s: average seconds from request start to first non-empty streaming delta. This includesreasoningdeltas and is not the user-visible content latency.avg_per_call_tok_s: average per-call decode throughput, computed as output tokens divided by time from first non-empty delta to completion.total_tok_s: aggregate group throughput, computed as total output tokens divided by wall-clock time from earliest request start to latest completion.group_wall_s: wall-clock duration for the concurrency group.
Content-First Baseline
This corrected measurement records when the first non-empty content delta appears. It is the user-visible answer latency requested for the think-off comparison.
- 측정 시각: 2026-07-13 17:55:18 KST
- Endpoint:
http://127.0.0.1:8007/v1 - Model alias:
ornith:35b - Thinking mode: current default. Upstream
mlx_lm.serverwas started with--chat-template-args '{"enable_thinking":true}'; benchmark requests did not passthink=false. max_tokens:1024temperature:0.2top_p:0.95- Prompt:
Benchmark generation task. Produce exactly 8 short numbered Korean lines about reliable agent runtime operations.
| Concurrent calls | Success | Avg first reasoning/any delta (s) | Avg first content delta (s) | Avg reasoning tokens/call | Avg content tokens/call | Avg all tok/s per call | Avg content tok/s per call | Total all tok/s | Total content tok/s | Group wall (s) | Finish reason |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 1/1 | 0.227 | 17.830 | 919.0 | 104.0 | 52.29 | 53.00 | 51.69 | 5.25 | 19.792 | length |
| 2 | 2/2 | 0.339 | 22.193 | 978.0 | 45.0 | 44.74 | 44.44 | 88.17 | 3.88 | 23.206 | length |
| 3 | 3/3 | 0.462 | 26.260 | 978.0 | 45.0 | 37.93 | 38.26 | 111.85 | 4.92 | 27.439 | length |
| 4 | 4/4 | 0.598 | 30.638 | 978.0 | 45.0 | 32.55 | 32.34 | 127.74 | 5.62 | 32.035 | length |
| 5 | 5/5 | 0.728 | 38.788 | 978.0 | 45.0 | 25.69 | 25.64 | 126.13 | 5.55 | 40.553 | length |
Content-First Per-Call Detail
| Concurrent calls | Call | First content delta (s) | Reasoning tokens | Content tokens | All tok/s | Content tok/s | Duration (s) |
|---|---|---|---|---|---|---|---|
| 1 | 1 | 17.830 | 919 | 104 | 52.29 | 53.00 | 19.792 |
| 2 | 1 | 22.195 | 978 | 45 | 44.74 | 44.49 | 23.206 |
| 2 | 2 | 22.191 | 978 | 45 | 44.73 | 44.38 | 23.205 |
| 3 | 1 | 26.258 | 978 | 45 | 37.93 | 38.33 | 27.432 |
| 3 | 2 | 26.263 | 978 | 45 | 37.93 | 38.29 | 27.439 |
| 3 | 3 | 26.258 | 978 | 45 | 37.92 | 38.15 | 27.437 |
| 4 | 1 | 30.637 | 978 | 45 | 32.55 | 32.47 | 32.023 |
| 4 | 2 | 30.637 | 978 | 45 | 32.54 | 32.22 | 32.034 |
| 4 | 3 | 30.642 | 978 | 45 | 32.55 | 32.31 | 32.035 |
| 4 | 4 | 30.637 | 978 | 45 | 32.55 | 32.34 | 32.028 |
| 5 | 1 | 38.787 | 978 | 45 | 25.69 | 25.50 | 40.551 |
| 5 | 2 | 38.787 | 978 | 45 | 25.70 | 25.70 | 40.538 |
| 5 | 3 | 38.787 | 978 | 45 | 25.69 | 25.61 | 40.544 |
| 5 | 4 | 38.793 | 978 | 45 | 25.69 | 25.58 | 40.552 |
| 5 | 5 | 38.788 | 978 | 45 | 25.70 | 25.81 | 40.531 |
Content-First Observations
- The user-visible first content delta was much later than the first reasoning delta.
- At concurrency 1, first reasoning/any delta arrived at
0.227s, but first content arrived at17.830s. - At concurrency 5, first reasoning/any delta arrived at
0.728s, but first content arrived at38.788s. - The run consistently spent most of the
1024token budget on reasoning before content. Most calls produced about978reasoning tokens and45content tokens beforefinish_reason=length. - For think-off comparison,
Avg first content delta (s)is the latency column to compare first.
Think-Off Result
Compatibility Note
The alias proxy currently maps request-level think:false to upstream chat_template_args.enable_thinking=false. The installed mlx_lm.server request parser reads chat_template_kwargs, not request-level chat_template_args (mlx_lm/server.py request parse path stores self.chat_template_kwargs = self.body.get("chat_template_kwargs")). Because of that mismatch, plain think:false did not produce a real think-off result in this runtime path.
Observed think:false compatibility probe:
| Mode | First content delta (s) | Reasoning chars | Content chars | Finish reason | Interpretation |
|---|---|---|---|---|---|
think:false via alias proxy |
17.975 | 3300 | 203 | length | Not effective; still generated reasoning first |
The effective think-off measurement below used request chat_template_kwargs: {"enable_thinking": false}.
Effective Think-Off Content-First Table
- 측정 시각: 2026-07-13 18:05:09 KST
- Endpoint:
http://127.0.0.1:8007/v1 - Model alias:
ornith:35b - Thinking mode: effective request-level think-off via
chat_template_kwargs.enable_thinking=false max_tokens:1024temperature:0.2top_p:0.95- Prompt: same as the Content-First Baseline.
| Concurrent calls | Success | Avg first content delta (s) | Avg reasoning tokens/call | Avg content tokens/call | Avg content tok/s per call | Total content tok/s | Group wall (s) | Finish reason |
|---|---|---|---|---|---|---|---|---|
| 1 | 1/1 | 0.208 | 0.0 | 104.0 | 52.77 | 47.74 | 2.179 | stop |
| 2 | 2/2 | 0.333 | 0.0 | 113.0 | 44.81 | 79.15 | 2.855 | stop |
| 3 | 3/3 | 0.457 | 0.0 | 113.0 | 37.79 | 98.27 | 3.450 | stop |
| 4 | 4/4 | 0.580 | 0.0 | 113.0 | 32.47 | 111.20 | 4.065 | stop |
| 5 | 5/5 | 0.703 | 0.0 | 113.0 | 25.82 | 111.03 | 5.089 | stop |
Think-On vs Effective Think-Off
| Concurrent calls | Think-on first content (s) | Effective think-off first content (s) | Delta (s) | Speedup |
|---|---|---|---|---|
| 1 | 17.830 | 0.208 | 17.622 | 85.9x |
| 2 | 22.193 | 0.333 | 21.860 | 66.7x |
| 3 | 26.260 | 0.457 | 25.803 | 57.5x |
| 4 | 30.638 | 0.580 | 30.058 | 52.9x |
| 5 | 38.788 | 0.703 | 38.085 | 55.2x |
Effective Think-Off Observations
- Effective think-off eliminated reasoning output in this benchmark (
0reasoning tokens for all calls). - First user-visible content became the first non-empty stream delta.
- Total content throughput peaked around concurrency 4-5 at about
111 tok/s. - The largest improvement is first content latency, not single-stream decode tok/s.
- To make
think:falsework through the alias endpoint, update the proxy to mapthinktochat_template_kwargs.enable_thinking, not onlychat_template_args.enable_thinking.
First-Content Optimization Candidates
| Priority | Candidate | Expected effect | Risk / note |
|---|---|---|---|
| 1 | Fix alias proxy think mapping to upstream chat_template_kwargs.enable_thinking |
Makes simple think:false clients get the measured 0.2-0.7s first-content path |
Very small code change, but verify /v1/chat/completions streaming and existing chat_template_kwargs passthrough |
| 2 | Start a separate always-non-thinking Ornith endpoint with mlx_lm.server --chat-template-args '{"enable_thinking":false}' |
Makes non-thinking the default without per-request kwargs | Keep on a separate port first; quality/tool-call behavior may differ |
| 3 | Use vllm_mlx Ornith path if stable, with --default-chat-template-kwargs '{"enable_thinking": false}' or request enable_thinking=false |
Better OpenAI-compatible semantics, continuous batching, prefix cache, warm prompts, and request-level thinking controls | This changes runtime stack from current mlx_lm.server; test as a parallel endpoint before replacing 8007 |
| 4 | If reasoning must stay on, enforce a small thinking budget where supported | Bounds reasoning delay before content instead of fully disabling reasoning | Current mlx_lm.server path has no observed thinking budget control; vllm_mlx/vLLM-style serving has this concept |
| 5 | Warm repeated prompt prefixes / prompt cache | Reduces prefill/TTFT for repeated system prompts and agent templates | Helps prefill, not the long reasoning-before-content delay. Think-off gives the bigger win here |
| 6 | Tune decode-concurrency, prompt-concurrency, and prefill-step-size on mlx_lm.server |
May improve batch scheduling and prefill behavior | Current content-first bottleneck was reasoning generation. Tune only after think-off route is correct |
References checked:
- Local
mlx_lm.server --help: supports--chat-template-args,--decode-concurrency,--prompt-concurrency,--prefill-step-size,--prompt-cache-size,--prompt-cache-bytes, and draft-model options. - Local installed
mlx_lm/server.py: request body useschat_template_kwargs; startup CLI useschat_template_args. - Qwen docs: previous Qwen3-style hybrid thinking can be controlled by
enable_thinking=False; newer Qwen3 Instruct/Thinking variants may be mode-specific. - vLLM docs: request-level
chat_template_kwargscan override server defaults, and thinking budgets can bound reasoning tokens. vllm-mlxdocs: continuous batching, prefix cache, SSD-tiered cache, and warm prompts are intended Apple Silicon serving optimizations.
Reasoning-Inclusive First-Delta Table
| Concurrent calls | Success | Avg output tokens/call | Total output tokens | Avg TTFT (s) | Avg duration (s) | Avg decode duration (s) | Avg tok/s per call | Total tok/s | Group wall (s) | Finish reason |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 1/1 | 256.0 | 256 | 0.221 | 5.085 | 4.864 | 52.63 | 50.34 | 5.085 | length |
| 2 | 2/2 | 256.0 | 512 | 0.340 | 6.072 | 5.732 | 44.66 | 84.31 | 6.073 | length |
| 3 | 3/3 | 256.0 | 768 | 0.462 | 7.213 | 6.751 | 37.92 | 106.42 | 7.217 | length |
| 4 | 4/4 | 256.0 | 1024 | 0.570 | 8.445 | 7.875 | 32.51 | 121.17 | 8.451 | length |
| 5 | 5/5 | 256.0 | 1280 | 0.698 | 10.639 | 9.940 | 25.75 | 120.20 | 10.649 | length |
Per-Call Detail
| Concurrent calls | Call | Output tokens | TTFT (s) | Duration (s) | Decode duration (s) | Tok/s |
|---|---|---|---|---|---|---|
| 1 | 1 | 256 | 0.221 | 5.085 | 4.864 | 52.63 |
| 2 | 1 | 256 | 0.339 | 6.072 | 5.734 | 44.65 |
| 2 | 2 | 256 | 0.342 | 6.072 | 5.730 | 44.68 |
| 3 | 1 | 256 | 0.465 | 7.216 | 6.751 | 37.92 |
| 3 | 2 | 256 | 0.461 | 7.216 | 6.755 | 37.90 |
| 3 | 3 | 256 | 0.460 | 7.208 | 6.748 | 37.94 |
| 4 | 1 | 256 | 0.569 | 8.443 | 7.874 | 32.51 |
| 4 | 2 | 256 | 0.569 | 8.438 | 7.869 | 32.53 |
| 4 | 3 | 256 | 0.569 | 8.450 | 7.881 | 32.48 |
| 4 | 4 | 256 | 0.574 | 8.451 | 7.877 | 32.50 |
| 5 | 1 | 256 | 0.697 | 10.647 | 9.950 | 25.73 |
| 5 | 2 | 256 | 0.697 | 10.632 | 9.935 | 25.77 |
| 5 | 3 | 256 | 0.696 | 10.639 | 9.943 | 25.75 |
| 5 | 4 | 256 | 0.704 | 10.648 | 9.944 | 25.74 |
| 5 | 5 | 256 | 0.698 | 10.626 | 9.929 | 25.78 |
Observations
- Every measured request completed successfully and stopped at
finish_reason=length. - Aggregate throughput increased from concurrency 1 to 4, then flattened at concurrency 5.
- In this baseline, total throughput was
121.17 tok/sat concurrency 4 and120.20 tok/sat concurrency 5. - Average TTFT rose with concurrency, from
0.221sat concurrency 1 to0.698sat concurrency 5.