Skip to content

Commit 5574335

Browse files
committed
Add direct native llama.cpp benchmark baseline with timing provenance
1 parent 819b1b3 commit 5574335

13 files changed

Lines changed: 1037 additions & 122 deletions

File tree

‎.github/workflows/verify.yml‎

Lines changed: 23 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -27,12 +27,13 @@ jobs:
2727
os: ubuntu-24.04
2828
runtime_identifier: linux-x64
2929
runs-on: ${{ matrix.os }}
30-
timeout-minutes: 20
30+
timeout-minutes: 35
3131
env:
3232
DOTNET_NOLOGO: true
3333
DOTNET_SKIP_FIRST_TIME_EXPERIENCE: true
3434
SYNAPSE_REQUIRED_RUNTIME_IDENTIFIER: ${{ matrix.runtime_identifier }}
3535
SYNAPSE_DOTLLM_VERSION: d88040451d7db56e5dfef9d5754ad0955b0f7fe5
36+
SYNAPSE_LLAMACPP_VERSION: b29c606e28a01b1bc8c1351026a0fa6e616bf6c4
3637
SYNAPSE_MODEL_ROOT: ${{ github.workspace }}/artifacts/models
3738
steps:
3839
- name: Checkout
@@ -46,6 +47,14 @@ jobs:
4647
path: _external/dotLLM
4748
persist-credentials: false
4849

50+
- name: Checkout pinned native llama.cpp benchmark subject
51+
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
52+
with:
53+
repository: ggml-org/llama.cpp
54+
ref: b29c606e28a01b1bc8c1351026a0fa6e616bf6c4
55+
path: _external/llama.cpp
56+
persist-credentials: false
57+
4958
- name: Install pinned .NET SDK
5059
uses: actions/setup-dotnet@a98b56852c35b8e3190ac28c8c2271da59106c68 # v6.0.0
5160
with:
@@ -77,10 +86,23 @@ jobs:
7786
- name: Build pinned dotLLM benchmark subject
7887
run: dotnet build _external/dotLLM/src/DotLLM.Cli/DotLLM.Cli.csproj --configuration Release
7988

89+
- name: Build pinned CPU llama.cpp benchmark subjects
90+
shell: bash
91+
run: |
92+
cmake -S _external/llama.cpp -B _external/llama.cpp/build \
93+
-DCMAKE_BUILD_TYPE=Release -DGGML_METAL=OFF \
94+
-DLLAMA_BUILD_TESTS=OFF -DLLAMA_BUILD_SERVER=OFF
95+
cmake --build _external/llama.cpp/build --config Release --parallel 4 \
96+
--target llama-completion llama-bench
97+
8098
- name: Register dotLLM smoke-test executable
8199
shell: bash
82100
run: echo "SYNAPSE_DOTLLM_EXECUTABLE=${GITHUB_WORKSPACE}/_external/dotLLM/src/DotLLM.Cli/bin/Release/net10.0/DotLLM.Cli" >> "${GITHUB_ENV}"
83101

102+
- name: Register native llama.cpp smoke-test executable
103+
shell: bash
104+
run: echo "SYNAPSE_LLAMACPP_EXECUTABLE=${GITHUB_WORKSPACE}/_external/llama.cpp/build/bin/llama-completion" >> "${GITHUB_ENV}"
105+
84106
- name: Test .NET
85107
run: dotnet test Synapse.slnx --configuration Release --no-build
86108

‎README.md‎

Lines changed: 49 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -121,12 +121,57 @@ llama.cpp remains far ahead in TTFT and decode. Synapse TTFT varied from 304.8
121121
to 574.2 ms, so scheduling stability is an explicit optimization target. The
122122
formal benchmark gate still requires 30 paired measurements, a thread-scaling
123123
sweep, the locked 10-turn growing-context dialogue, cache hit/miss workloads,
124-
embeddings, and direct llama.cpp. Energy is
124+
and embeddings. Energy is
125125
`not_run_missing_privilege`: `/usr/bin/powermetrics` requires superuser access,
126126
and missing energy data is never reported as zero.
127127

128+
Direct native `llama.cpp` 0.4.1 (`b29c606e2`) now passes the same Qwen prompt
129+
token IDs and eight-token continuation. It was measured separately, not
130+
interleaved with the three-subject run above, so it is **not a fourth paired
131+
row** in that table. After three warm-ups, five fresh-process native runs had
132+
median process wall 566.2 ms and max observed RSS 1,208.3 MiB. The native CLI
133+
reported median prompt-eval 14.9 ms and eval 57.9 ms (120.84 eval tok/s).
134+
These are native internal phases, **not** comparable load, TTFT, or end-to-end
135+
generation measurements; those fields remain `null`. For a 128-token output,
136+
five measured native completion runs ranged from 47.28 to 99.40 internal eval
137+
tok/s (median 80.73), with identical generated text. This large spread is a
138+
reason to avoid a winner claim while the Mac also runs other development work.
139+
One later repeat exceeded 90 seconds for the same 128-token limit and was
140+
terminated; it is not folded into the five successful samples.
141+
Separately, native `llama-bench` reported 115.36 ± 5.94 tok/s for five
142+
128-token repetitions; [its own methodology](https://github.com/ggml-org/llama.cpp/blob/master/tools/llama-bench/README.md)
143+
excludes tokenization and sampling, so it is a kernel diagnostic, not a
144+
completion-process result.
145+
146+
| Catalog model / architecture | Synapse | dotLLM | LLamaSharp | native llama.cpp | MLX | ONNX Runtime |
147+
|---|---|---|---|---|---|---|
148+
| Qwen2.5 0.5B Q8_0 · Qwen2 | D8 | D8 | D8 | D8 + D128 | NR | NR |
149+
| SmolLM2 135M BF16 · Llama | NR | NR | NR | NR | NR | NR |
150+
| Qwen3 0.6B Q8_0 · Qwen3 | NR | NR | NR | NR | NR | NR |
151+
| Mamba 130M F32 · SSM | NR | NR | NR | NR | NR | NR |
152+
| Phi-3 Mini 3.8B Q4 · Phi3 | NR | NR | NR | NR | NR | NR |
153+
| DeepSeek-R1 Distill 1.5B BF16 · Qwen2 | NR | NR | NR | NR | NR | NR |
154+
| Ministral 3 3B Q4 · Mistral3 | NR | NR | NR | NR | NR | NR |
155+
| all-MiniLM-L6-v2 F32 · BERT embedding | NR | NR | NR | NR | NR | NR |
156+
| BGE-small-en-v1.5 F32 · BERT embedding | NR | NR | NR | NR | NR | NR |
157+
158+
`D8`/`D128` mean measured diagnostic output lengths, not formal benchmark
159+
victories; `NR` means no compatible measurement yet, not zero performance.
160+
Each newly qualified model will receive the same load/TTFT/generation/decode/
161+
CPU/memory comparison table in its own hardware and precision cohort; the
162+
coverage matrix does not substitute made-up numbers for those future tables.
163+
Only Qwen GGUF currently runs through Synapse's inference path. Other catalog
164+
packages have different architectures and/or SafeTensors/embedding formats;
165+
the existence of a pinned download is not evidence of executable inference.
166+
MLX is a candidate Python-free native/Swift Apple Silicon baseline and ONNX
167+
Runtime GenAI is a candidate C# baseline using a separately pinned ONNX model.
168+
Their GPU/format/precision cohorts will not be silently mixed with CPU GGUF.
169+
128170
Raw measured samples and binary/model fingerprints are stored in
129171
[`benchmarks/results/2026-09-28-m2-pro-qwen2.5-0.5b-q8_0-smoke.json`](benchmarks/results/2026-09-28-m2-pro-qwen2.5-0.5b-q8_0-smoke.json).
172+
The [native llama.cpp diagnostic samples](benchmarks/results/2026-09-28-m2-pro-native-llamacpp-qwen2.5-0.5b-q8_0-diagnostic.json)
173+
include the separate 8/128-token completion repeats, kernel microbenchmark,
174+
model digest, and binary fingerprints.
130175
The locked dialogue and embedding workloads are in `benchmarks/scenarios/`.
131176

132177
## Build, test, and run
@@ -189,7 +234,9 @@ docs/ architecture, ADRs, features, commands, task regist
189234
|---|---|---|
190235
| [dotLLM](https://github.com/kkokosa/dotLLM) | Pure-.NET correctness/performance competitor and architecture study | GPL-3.0; pinned checkout and separate process only |
191236
| [LLamaSharp](https://github.com/SciSharp/LLamaSharp) | Required llama.cpp-backed CPU benchmark | MIT; benchmark-project package 0.27.0 |
192-
| [llama.cpp](https://github.com/ggml-org/llama.cpp) | GGUF/quantization reference and required direct native baseline | MIT; direct runner is still pending |
237+
| [llama.cpp](https://github.com/ggml-org/llama.cpp) | GGUF/quantization reference and direct native baseline | MIT; external CPU process pinned at `b29c606e2` |
238+
| [MLX Swift LM](https://github.com/ml-explore/mlx-swift-lm) | Candidate Python-free Apple Silicon/Metal baseline | External subject planned; no measurement yet |
239+
| [ONNX Runtime GenAI](https://onnxruntime.ai/docs/genai/api/csharp.html) | Candidate C# ONNX-format baseline | Preview API; verified ONNX package and measurement pending |
193240
| [ZoneTree](https://github.com/ZoneTree/ZoneTree) | Durable cache metadata, prefix indexes, journals, evidence indexes | MIT; runtime package 1.9.8 |
194241
| [Microsoft Orleans](https://github.com/dotnet/orleans) | Request/control plane, leases, epochs, placement, recovery | Planned D3 dependency; never tensor/KV transport |
195242
| [Aspire](https://github.com/dotnet/aspire) | Multi-process topology, health, telemetry, test orchestration | Added only with the first real distributed topology |

0 commit comments

Comments
 (0)