MOSS 统一转写与说话人分离
OpenMOSS-Team/MOSS-Transcribe-Diarize@e8681d68
Direct vLLM HTTP and the FunASR adapter each returned two non-empty monotonic segments with speakers S01 and S02.
- 运行时
- vLLM 0.27.1 / Torch 2.13.0+cu129
- 硬件
- NVIDIA H100 80GB HBM3
- 负载
- OpenAI-compatible diarized_json API smoke plus FunASR AutoModel structured-response mapping
- 音频
- 15.1685 s two-speaker probe; SHA256 43dccc068506439cb633b382b6b98185baa837363d08cc5f7152ca89b0fdc3c8
- 设置
- Pinned model revision; vLLM 0.27.1; vllm[audio]; response_format=diarized_json; temperature=0; max_completion_tokens=768
- 计时口径
- Functional response validation after /health became ready; no latency or throughput claim
- 核对日期
- 2026-08-30
限定:A structured response contract smoke, not a diarization accuracy study, overlap test, throughput benchmark, or production capacity promise.
查看原始来源
MOSS 统一转写与说话人分离
OpenMOSS-Team/MOSS-Transcribe-Diarize@e8681d68
1088 / 1088 completed with 0 failures; raw WER 2.74% and WER excluding >50% outliers 2.45%.
- 运行时
- SGLang Omni merge 8458f76a
- 硬件
- NVIDIA H100 80GB
- 负载
- Seed-TTS EN full set, non-streaming, max concurrency 16, one repeat
- 音频
- 1088 single-speaker English clips; mean duration 4.736 s
- 设置
- Single H100; 1088 / 1088 requests completed; same upstream benchmark harness used for model comparison
- 计时口径
- Upstream end-to-end benchmark report; throughput 55.487 req/s, mean latency 0.287 s, mean RTF 0.0625
- 核对日期
- 2026-07-04
限定:Single-speaker English ASR benchmark after stripping timestamp and speaker markup; diarization/timestamp correctness is not evaluated, and the result is not a production capacity promise.
查看原始来源
vLLM 原生 FunASR 服务
allendou/Fun-ASR-Nano-2512-vllm@e718b36e
Chinese 200 in 0.968 s; repeated hotword corrected 开饭时间 to 开放时间 in 0.214 s; English and Japanese two-request batch completed in 1.123 s wall time
- 运行时
- vLLM 0.27.1+cu129 / Torch 2.13.0+cu129
- 硬件
- NVIDIA H100 80GB
- 负载
- OpenAI-compatible transcription; Chinese baseline and hotword probes plus two concurrent requests (English and Japanese)
- 音频
- Pinned model-repository examples: 6 s Chinese, 8 s English, and 8 s Japanese
- 设置
- FP32; eager mode; gpu-memory-utilization 0.40; max model length 40,960; explicit language; warmed server
- 计时口径
- Client wall time after /health became ready; local HTTP and decoding included; cold model download and 20.3 s engine warmup excluded
- 核对日期
- 2026-08-13
限定:Single-H100 correctness and concurrency probe with a community-converted checkpoint, not an accuracy study or production capacity promise. Revalidate the checkpoint, target GPU, languages, traffic, and hotword policy before rollout.
查看原始来源
SenseVoice TensorRT / Triton
SenseVoiceSmall
527,504,916 bytes; 113.9 s engine build; 100% CTC top-1 agreement; exact bundled-audio transcript
- 运行时
- TensorRT 10.0.1 with Triton 24.05 model contract
- 硬件
- NVIDIA H100; device memory capacity and host CPU are not part of the recorded result
- 负载
- Native FP16 engine build, profile-bound execution, and PyTorch parity validation
- 音频
- Bundled Chinese example plus synthetic feature tensors at 30 and 64 frames
- 设置
- FP16; batch profile 1/8/16; post-LFR frame profile 1/512/4096; 8 GiB builder workspace
- 计时口径
- Engine build wall time only; inference timings are correctness probes and not a throughput benchmark
- 核对日期
- 2026-08-04
限定:Single-H100 compatibility and parity evidence. Build on the target GPU and TensorRT version, then load-test with production audio before capacity planning.
查看原始来源
llama.cpp / GGUF 独立运行
SenseVoiceSmall F16 GGUF
100/100 processes exited successfully; 0 empty transcripts; one output SHA-256; transcript 我想问我在滨海新区有房。
- 运行时
- runtime-llamacpp-v0.2.4 Linux x64 AVX2 release asset
- 硬件
- Intel Xeon Platinum 8480+, 8 runtime threads
- 负载
- 100 sequential cold-process SenseVoice transcriptions
- 音频
- One 6.00 s Mandarin PCM sample; mono, 16 kHz, 16-bit
- 设置
- Published AVX2 archive; 470,197,600-byte F16 GGUF; direct SenseVoice CLI without VAD
- 计时口径
- Functional stability validation only; every process included model loading and decoding, and timings were not retained
- 核对日期
- 2026-08-29
限定:Single-host, single-sample regression evidence for the former intermittent empty-output defect, not an accuracy study or production capacity promise.
查看原始来源
llama.cpp / GGUF 独立运行
SenseVoiceSmall Q8 GGUF
Approximately 20x realtime; Q8 CER 8.17%
- 运行时
- llama.cpp/GGUF with built-in FSMN-VAD
- 硬件
- CPU, 8 threads; CPU model not recorded in the cited report
- 负载
- Offline Mandarin CPU
- 音频
- 184 Mandarin clips, about 44-60 seconds each; total duration not recorded in the cited report
- 设置
- Q8 runtime; built-in FSMN-VAD; micro-CER normalize_zh; model load excluded
- 计时口径
- Sum of compute time divided by sum of audio duration; model load excluded
- 核对日期
- 2026-07-26
限定:CPU model and full software stack are not recorded. Reproduce on target hardware before comparison.
查看原始来源
llama.cpp / GGUF 独立运行
Paraformer Q8 GGUF
Approximately 21x realtime; Q8 CER 9.89%
- 运行时
- llama.cpp/GGUF with built-in FSMN-VAD
- 硬件
- CPU, 8 threads; CPU model not recorded in the cited report
- 负载
- Offline Mandarin CPU
- 音频
- 184 Mandarin clips, about 44-60 seconds each; total duration not recorded in the cited report
- 设置
- Q8 runtime; built-in FSMN-VAD; micro-CER normalize_zh; model load excluded
- 计时口径
- Sum of compute time divided by sum of audio duration; model load excluded
- 核对日期
- 2026-07-26
限定:CPU model and full software stack are not recorded. Reproduce on target hardware before comparison.
查看原始来源
audio.cpp 原生 Fun-ASR-Nano 与 SenseVoice
Fun-ASR-Nano-2512 Q8_0 GGUF
Approximately 0.0725 RTF; transcript matched the official safetensors path
- 运行时
- audio.cpp native Fun-ASR-Nano runtime
- 硬件
- NVIDIA H100 CUDA; exact driver and toolchain versions are not recorded in the cited guide
- 负载
- Single offline transcription through the OpenAI-compatible server
- 音频
- One 14.07-second reference WAV
- 设置
- Standalone Q8_0 GGUF; CUDA backend; concurrency one; other runtime settings are not recorded
- 计时口径
- End-to-end server validation; the cited guide does not state warmup, file I/O, or process-start boundaries
- 核对日期
- 2026-07-29
限定:Single-sample parity smoke, not a capacity benchmark. Reproduce on target hardware with production audio and concurrency.
查看原始来源
audio.cpp 原生 Fun-ASR-Nano 与 SenseVoice
SenseVoice-Small Q8_0 GGUF
919 tensors loaded without sidecar overrides; transcript 开饭时间早上9点至下午5点。; GGUF SHA-256 4dedf169f625437fb336f2959674f399819729a765e184128c0e25a6e16ff0ec
- 运行时
- audio.cpp native sense_asr main@979e070f
- 硬件
- CPU; validation host details are not a cross-hardware benchmark contract
- 负载
- Offline Chinese transcription plus buffered streaming partials
- 音频
- Known 16 kHz reference WAV and PCM stream
- 设置
- Standalone schema-v1 Q8_0 GGUF; CPU backend; two-second streaming windows during validation
- 计时口径
- Functional parity and loader validation; no capacity claim
- 核对日期
- 2026-08-13
限定:Merged into audio.cpp main with all six contribution CI jobs passing, but not yet included in a tagged release; pin the exact main commit.
查看原始来源