Native FunASR on vLLM
allendou/Fun-ASR-Nano-2512-vllm@e718b36e
- Runtime
- vLLM 0.27.1+cu129 / Torch 2.13.0+cu129
- Hardware
- NVIDIA H100 80GB
- Workload
- OpenAI-compatible transcription; Chinese baseline and hotword probes plus two concurrent requests (English and Japanese)
- Audio
- Pinned model-repository examples: 6 s Chinese, 8 s English, and 8 s Japanese
- Settings
- FP32; eager mode; gpu-memory-utilization 0.40; max model length 40,960; explicit language; warmed server
- Timing scope
- Client wall time after /health became ready; local HTTP and decoding included; cold model download and 20.3 s engine warmup excluded
- Verified
- 2026-08-13
Qualification: Single-H100 correctness and concurrency probe with a community-converted checkpoint, not an accuracy study or production capacity promise. Revalidate the checkpoint, target GPU, languages, traffic, and hotword policy before rollout.
Open primary source