Reproducible measurements

Compare test conditions before results

This page preserves the context and missing fields of public results so teams can establish their own capacity baseline.

Records
6
Verified
2026-08-13
Not directly comparable

RTFx changes with hardware, audio distribution, VAD, concurrency, warmup, and timing boundaries. A result missing any condition is not a procurement or capacity promise.

1

Fix inputs

Record total duration, file distribution, language, channels, and sample rate.

2

Fix system

Record CPU/GPU, driver, model, commit, runtime, and every parameter.

3

Define timing

State whether loading, warmup, I/O, queueing, and post-processing are included.

Public records

Read every number with its qualifications

Native FunASR on vLLM

allendou/Fun-ASR-Nano-2512-vllm@e718b36e

Chinese 200 in 0.968 s; repeated hotword corrected 开饭时间 to 开放时间 in 0.214 s; English and Japanese two-request batch completed in 1.123 s wall time
Runtime
vLLM 0.27.1+cu129 / Torch 2.13.0+cu129
Hardware
NVIDIA H100 80GB
Workload
OpenAI-compatible transcription; Chinese baseline and hotword probes plus two concurrent requests (English and Japanese)
Audio
Pinned model-repository examples: 6 s Chinese, 8 s English, and 8 s Japanese
Settings
FP32; eager mode; gpu-memory-utilization 0.40; max model length 40,960; explicit language; warmed server
Timing scope
Client wall time after /health became ready; local HTTP and decoding included; cold model download and 20.3 s engine warmup excluded
Verified
2026-08-13

Qualification: Single-H100 correctness and concurrency probe with a community-converted checkpoint, not an accuracy study or production capacity promise. Revalidate the checkpoint, target GPU, languages, traffic, and hotword policy before rollout.

Open primary source

SenseVoice TensorRT / Triton

SenseVoiceSmall

527,504,916 bytes; 113.9 s engine build; 100% CTC top-1 agreement; exact bundled-audio transcript
Runtime
TensorRT 10.0.1 with Triton 24.05 model contract
Hardware
NVIDIA H100; device memory capacity and host CPU are not part of the recorded result
Workload
Native FP16 engine build, profile-bound execution, and PyTorch parity validation
Audio
Bundled Chinese example plus synthetic feature tensors at 30 and 64 frames
Settings
FP16; batch profile 1/8/16; post-LFR frame profile 1/512/4096; 8 GiB builder workspace
Timing scope
Engine build wall time only; inference timings are correctness probes and not a throughput benchmark
Verified
2026-08-04

Qualification: Single-H100 compatibility and parity evidence. Build on the target GPU and TensorRT version, then load-test with production audio before capacity planning.

Open primary source

llama.cpp / GGUF standalone

SenseVoiceSmall Q8 GGUF

Approximately 20x realtime; Q8 CER 8.17%
Runtime
llama.cpp/GGUF with built-in FSMN-VAD
Hardware
CPU, 8 threads; CPU model not recorded in the cited report
Workload
Offline Mandarin CPU
Audio
184 Mandarin clips, about 44-60 seconds each; total duration not recorded in the cited report
Settings
Q8 runtime; built-in FSMN-VAD; micro-CER normalize_zh; model load excluded
Timing scope
Sum of compute time divided by sum of audio duration; model load excluded
Verified
2026-07-26

Qualification: CPU model and full software stack are not recorded. Reproduce on target hardware before comparison.

Open primary source

llama.cpp / GGUF standalone

Paraformer Q8 GGUF

Approximately 21x realtime; Q8 CER 9.89%
Runtime
llama.cpp/GGUF with built-in FSMN-VAD
Hardware
CPU, 8 threads; CPU model not recorded in the cited report
Workload
Offline Mandarin CPU
Audio
184 Mandarin clips, about 44-60 seconds each; total duration not recorded in the cited report
Settings
Q8 runtime; built-in FSMN-VAD; micro-CER normalize_zh; model load excluded
Timing scope
Sum of compute time divided by sum of audio duration; model load excluded
Verified
2026-07-26

Qualification: CPU model and full software stack are not recorded. Reproduce on target hardware before comparison.

Open primary source

audio.cpp native Fun-ASR-Nano and SenseVoice

Fun-ASR-Nano-2512 Q8_0 GGUF

Approximately 0.0725 RTF; transcript matched the official safetensors path
Runtime
audio.cpp native Fun-ASR-Nano runtime
Hardware
NVIDIA H100 CUDA; exact driver and toolchain versions are not recorded in the cited guide
Workload
Single offline transcription through the OpenAI-compatible server
Audio
One 14.07-second reference WAV
Settings
Standalone Q8_0 GGUF; CUDA backend; concurrency one; other runtime settings are not recorded
Timing scope
End-to-end server validation; the cited guide does not state warmup, file I/O, or process-start boundaries
Verified
2026-07-29

Qualification: Single-sample parity smoke, not a capacity benchmark. Reproduce on target hardware with production audio and concurrency.

Open primary source

audio.cpp native Fun-ASR-Nano and SenseVoice

SenseVoice-Small Q8_0 GGUF

919 tensors loaded without sidecar overrides; transcript 开饭时间早上9点至下午5点。; GGUF SHA-256 4dedf169f625437fb336f2959674f399819729a765e184128c0e25a6e16ff0ec
Runtime
audio.cpp native sense_asr main@979e070f
Hardware
CPU; validation host details are not a cross-hardware benchmark contract
Workload
Offline Chinese transcription plus buffered streaming partials
Audio
Known 16 kHz reference WAV and PCM stream
Settings
Standalone schema-v1 Q8_0 GGUF; CPU backend; two-second streaming windows during validation
Timing scope
Functional parity and loader validation; no capacity claim
Verified
2026-08-13

Qualification: Merged into audio.cpp main with all six contribution CI jobs passing, but not yet included in a tagged release; pin the exact main commit.

Open primary source

Consumer GPU community validation

Fun-ASR Nano on one RTX 4090

The SGLang-Omni community completed full SeedTTS English and Chinese sweeps plus a 30-minute stability run. These results describe the tested feature branch, not a released SGLang-Omni integration.

SGLang-Omni 0.1.0

Fun-ASR Nano

c=16: 114.26 EN / 127.25 ZH samples/s
Hardware
RTX 4090 24,564 MiB; Ryzen 9 7950X; 125 GiB RAM
Software
CUDA 13.0; PyTorch 2.11.0; Transformers 5.6.0; Fun-ASR 854d88f
SeedTTS EN
1,088/1,088; c=1: 20.16 samples/s, 0.050 s; c=16: 114.26 samples/s, 0.139 s; WER 0.0175
SeedTTS ZH
2,020/2,020; c=1: 14.36 samples/s, 0.071 s; c=16: 127.25 samples/s, 0.125 s; CER 0.0164
Stability
30-minute mixed c=1/4/8/16 soak; 105,067 successful requests; 0 unexpected errors
Memory
15.98 GiB after static allocation; 16.11 GiB after graph capture; GPU memory released on exit
Timing scope
Three measured repeats after warmup; complete SeedTTS sweeps; zero skipped clips
Request boundary
Maximum 30 seconds per request; only terminal transcript events were observed

Qualification: Results came from a dirty feature worktree while upstream integration remained in progress; they do not establish released support or low-latency streaming deltas.

Realtime services

Streaming capacity needs a separate test

File throughput does not describe draft latency, final latency, long-lived connection concurrency, or backpressure. Use the realtime WebSocket method to record them.

Realtime test method Offline RTFx method