Reproducible measurements

Compare test conditions before results

This page preserves the context and missing fields of public results so teams can establish their own capacity baseline.

Records
3
Verified
2026-07-26
Not directly comparable

RTFx changes with hardware, audio distribution, VAD, concurrency, warmup, and timing boundaries. A result missing any condition is not a procurement or capacity promise.

1

Fix inputs

Record total duration, file distribution, language, channels, and sample rate.

2

Fix system

Record CPU/GPU, driver, model, commit, runtime, and every parameter.

3

Define timing

State whether loading, warmup, I/O, queueing, and post-processing are included.

Public records

Read every number with its qualifications

vLLM GPU throughput

Fun-ASR-Nano-2512

RTFx 340; CER 8.20%
Runtime
vLLM batch
Hardware
GPU model not recorded in the cited public table
Workload
Offline batch
Audio
184 long-form files; 11,541 seconds
Settings
Dynamic VAD; the public table does not record batch size or the full software and hardware stack
Timing scope
Offline throughput; the cited table does not state whether warmup, file I/O, and decoding are excluded
Verified
2026-07-26

Qualification: Incomplete hardware and timing record. Use only as public reference evidence, not a capacity promise.

Open primary source

llama.cpp / GGUF standalone

SenseVoiceSmall Q8 GGUF

Approximately 20x realtime; Q8 CER 8.17%
Runtime
llama.cpp/GGUF with built-in FSMN-VAD
Hardware
CPU, 8 threads; CPU model not recorded in the cited report
Workload
Offline Mandarin CPU
Audio
184 Mandarin clips, about 44-60 seconds each; total duration not recorded in the cited report
Settings
Q8 runtime; built-in FSMN-VAD; micro-CER normalize_zh; model load excluded
Timing scope
Sum of compute time divided by sum of audio duration; model load excluded
Verified
2026-07-26

Qualification: CPU model and full software stack are not recorded. Reproduce on target hardware before comparison.

Open primary source

llama.cpp / GGUF standalone

Paraformer Q8 GGUF

Approximately 21x realtime; Q8 CER 9.89%
Runtime
llama.cpp/GGUF with built-in FSMN-VAD
Hardware
CPU, 8 threads; CPU model not recorded in the cited report
Workload
Offline Mandarin CPU
Audio
184 Mandarin clips, about 44-60 seconds each; total duration not recorded in the cited report
Settings
Q8 runtime; built-in FSMN-VAD; micro-CER normalize_zh; model load excluded
Timing scope
Sum of compute time divided by sum of audio duration; model load excluded
Verified
2026-07-26

Qualification: CPU model and full software stack are not recorded. Reproduce on target hardware before comparison.

Open primary source

Realtime services

Streaming capacity needs a separate test

File throughput does not describe draft latency, final latency, long-lived connection concurrency, or backpressure. Use the realtime WebSocket method to record them.

Realtime test method Offline RTFx method