vLLM 原生 FunASR 服务
allendou/Fun-ASR-Nano-2512-vllm@e718b36e
- 运行时
- vLLM 0.27.1+cu129 / Torch 2.13.0+cu129
- 硬件
- NVIDIA H100 80GB
- 负载
- OpenAI-compatible transcription; Chinese baseline and hotword probes plus two concurrent requests (English and Japanese)
- 音频
- Pinned model-repository examples: 6 s Chinese, 8 s English, and 8 s Japanese
- 设置
- FP32; eager mode; gpu-memory-utilization 0.40; max model length 40,960; explicit language; warmed server
- 计时口径
- Client wall time after /health became ready; local HTTP and decoding included; cold model download and 20.3 s engine warmup excluded
- 核对日期
- 2026-08-13
限定:Single-H100 correctness and concurrency probe with a community-converted checkpoint, not an accuracy study or production capacity promise. Revalidate the checkpoint, target GPU, languages, traffic, and hotword policy before rollout.
查看原始来源