可复现实测

先比较实验条件,再比较结果

这里保留公开结果的上下文和缺失项,帮助团队建立自己的容量基线。

记录
6
验证日期
2026-08-13
不可直接横向比较

RTFx 会随硬件、音频分布、VAD、并发、预热和计时边界变化。缺少任一条件的结果不能作为采购或容量承诺。

1

固定输入

记录音频总时长、文件分布、语言、声道和采样率。

2

固定系统

记录 CPU/GPU、驱动、模型、提交、运行时和所有参数。

3

定义计时

说明是否包含加载、预热、I/O、排队和后处理。

公开记录

带完整限定条件阅读每一个数字

vLLM 原生 FunASR 服务

allendou/Fun-ASR-Nano-2512-vllm@e718b36e

Chinese 200 in 0.968 s; repeated hotword corrected 开饭时间 to 开放时间 in 0.214 s; English and Japanese two-request batch completed in 1.123 s wall time
运行时
vLLM 0.27.1+cu129 / Torch 2.13.0+cu129
硬件
NVIDIA H100 80GB
负载
OpenAI-compatible transcription; Chinese baseline and hotword probes plus two concurrent requests (English and Japanese)
音频
Pinned model-repository examples: 6 s Chinese, 8 s English, and 8 s Japanese
设置
FP32; eager mode; gpu-memory-utilization 0.40; max model length 40,960; explicit language; warmed server
计时口径
Client wall time after /health became ready; local HTTP and decoding included; cold model download and 20.3 s engine warmup excluded
核对日期
2026-08-13

限定:Single-H100 correctness and concurrency probe with a community-converted checkpoint, not an accuracy study or production capacity promise. Revalidate the checkpoint, target GPU, languages, traffic, and hotword policy before rollout.

查看原始来源

SenseVoice TensorRT / Triton

SenseVoiceSmall

527,504,916 bytes; 113.9 s engine build; 100% CTC top-1 agreement; exact bundled-audio transcript
运行时
TensorRT 10.0.1 with Triton 24.05 model contract
硬件
NVIDIA H100; device memory capacity and host CPU are not part of the recorded result
负载
Native FP16 engine build, profile-bound execution, and PyTorch parity validation
音频
Bundled Chinese example plus synthetic feature tensors at 30 and 64 frames
设置
FP16; batch profile 1/8/16; post-LFR frame profile 1/512/4096; 8 GiB builder workspace
计时口径
Engine build wall time only; inference timings are correctness probes and not a throughput benchmark
核对日期
2026-08-04

限定:Single-H100 compatibility and parity evidence. Build on the target GPU and TensorRT version, then load-test with production audio before capacity planning.

查看原始来源

llama.cpp / GGUF 独立运行

SenseVoiceSmall Q8 GGUF

Approximately 20x realtime; Q8 CER 8.17%
运行时
llama.cpp/GGUF with built-in FSMN-VAD
硬件
CPU, 8 threads; CPU model not recorded in the cited report
负载
Offline Mandarin CPU
音频
184 Mandarin clips, about 44-60 seconds each; total duration not recorded in the cited report
设置
Q8 runtime; built-in FSMN-VAD; micro-CER normalize_zh; model load excluded
计时口径
Sum of compute time divided by sum of audio duration; model load excluded
核对日期
2026-07-26

限定:CPU model and full software stack are not recorded. Reproduce on target hardware before comparison.

查看原始来源

llama.cpp / GGUF 独立运行

Paraformer Q8 GGUF

Approximately 21x realtime; Q8 CER 9.89%
运行时
llama.cpp/GGUF with built-in FSMN-VAD
硬件
CPU, 8 threads; CPU model not recorded in the cited report
负载
Offline Mandarin CPU
音频
184 Mandarin clips, about 44-60 seconds each; total duration not recorded in the cited report
设置
Q8 runtime; built-in FSMN-VAD; micro-CER normalize_zh; model load excluded
计时口径
Sum of compute time divided by sum of audio duration; model load excluded
核对日期
2026-07-26

限定:CPU model and full software stack are not recorded. Reproduce on target hardware before comparison.

查看原始来源

audio.cpp 原生 Fun-ASR-Nano 与 SenseVoice

Fun-ASR-Nano-2512 Q8_0 GGUF

Approximately 0.0725 RTF; transcript matched the official safetensors path
运行时
audio.cpp native Fun-ASR-Nano runtime
硬件
NVIDIA H100 CUDA; exact driver and toolchain versions are not recorded in the cited guide
负载
Single offline transcription through the OpenAI-compatible server
音频
One 14.07-second reference WAV
设置
Standalone Q8_0 GGUF; CUDA backend; concurrency one; other runtime settings are not recorded
计时口径
End-to-end server validation; the cited guide does not state warmup, file I/O, or process-start boundaries
核对日期
2026-07-29

限定:Single-sample parity smoke, not a capacity benchmark. Reproduce on target hardware with production audio and concurrency.

查看原始来源

audio.cpp 原生 Fun-ASR-Nano 与 SenseVoice

SenseVoice-Small Q8_0 GGUF

919 tensors loaded without sidecar overrides; transcript 开饭时间早上9点至下午5点。; GGUF SHA-256 4dedf169f625437fb336f2959674f399819729a765e184128c0e25a6e16ff0ec
运行时
audio.cpp native sense_asr main@979e070f
硬件
CPU; validation host details are not a cross-hardware benchmark contract
负载
Offline Chinese transcription plus buffered streaming partials
音频
Known 16 kHz reference WAV and PCM stream
设置
Standalone schema-v1 Q8_0 GGUF; CPU backend; two-second streaming windows during validation
计时口径
Functional parity and loader validation; no capacity claim
核对日期
2026-08-13

限定:Merged into audio.cpp main with all six contribution CI jobs passing, but not yet included in a tagged release; pin the exact main commit.

查看原始来源

消费级 GPU 社区验证

单张 RTX 4090 上的 Fun-ASR Nano

SGLang-Omni 社区完成了 SeedTTS 英文、中文全量测试和 30 分钟稳定性测试。以下结果对应被测 feature branch,不代表 SGLang-Omni 已发布该集成。

SGLang-Omni 0.1.0

Fun-ASR Nano

c=16: 114.26 EN / 127.25 ZH samples/s
硬件
RTX 4090 24,564 MiB; Ryzen 9 7950X; 125 GiB RAM
软件
CUDA 13.0; PyTorch 2.11.0; Transformers 5.6.0; Fun-ASR 854d88f
SeedTTS EN
1,088/1,088; c=1: 20.16 samples/s, 0.050 s; c=16: 114.26 samples/s, 0.139 s; WER 0.0175
SeedTTS ZH
2,020/2,020; c=1: 14.36 samples/s, 0.071 s; c=16: 127.25 samples/s, 0.125 s; CER 0.0164
稳定性
30 分钟混合 c=1/4/8/16;105,067 次成功请求;0 个意外错误
显存
静态分配后 15.98 GiB;CUDA Graph 捕获后 16.11 GiB;退出后显存释放
测量口径
预热后三次测量;完整 SeedTTS sweep;无跳过样本
请求边界
每次请求最长 30 秒;只观察到最终转写事件

限定:结果来自有未提交改动的 feature worktree,上游集成仍在推进;不能据此声称已发布支持或具备低延迟增量输出。

实时服务

流式容量需要单独测量

文件吞吐不能说明首稿延迟、最终结果延迟、长连接并发和背压。使用实时 WebSocket 方法记录这些指标。

实时测试方法 离线 RTFx 方法