FunASR Ecosystem
Open-source projects and integrations powered by FunASR, SenseVoice, and Paraformer.
Video & Media Tools
6K+ stars
ASR-powered video clipping. Automatic subtitles, keyword and speaker-based clip extraction. Download v2.1.1
VideoOfficial
10.5K+ stars
Real-time streaming transcription. The merged SenseVoiceSmall backend runs through LocalAgreement and the VAC/VAD pipeline, preserves word timestamps, and supports auto detection plus Mandarin, Cantonese, English, Japanese, and Korean.
Live TranscriptionSenseVoiceSmall
13.7K+ stars
v5.2.0-beta2 ships local Fun-ASR Nano and SenseVoice video subtitle engines for Windows, macOS, and Linux, with Q4, Q8, and F16 models. PR #13063 updates the CrispASR runtime and verifies the SenseVoice Metal path; read the setup and model guide.
Video SubtitlesFun-ASR NanoSenseVoice
17.6K stars
Video translation tool. Uses FunASR for Chinese speech recognition with subtitle generation.
VideoTranslation
10.4K stars
Gradio WebUI for TTS, voice cloning, and audio processing with ASR capabilities.
AudioTTS
3.2K stars
AI video dubbing toolkit. Automatic speech recognition, translation, and voice cloning for multilingual dubbing.
DubbingVideo
Voice Input & Desktop Apps
13.1K+ stars
An open-source AI wearable and mobile app. Merged #10447 adds a default-on “Send raw audio to Omi” control. When disabled,
raw audio goes only to the selected STT provider; transcripts and non-audio data may still reach Omi. Unsupported Custom STT codecs fail closed instead of falling back to Omi audio upload. The provider can be a compatible FunASR or SenseVoice service.WearableCustom STTPrivacy
5.5K stars
PC voice input tool with offline recognition. Hold CapsLock to speak, release to paste. Powered by FunASR Paraformer.
DesktopVoice Input
1.8K stars
Voice input for macOS & Windows. Hold a key, speak, release — text appears at cursor. Uses SenseVoice via Sherpa ONNX.
DesktopmacOS
30K+ stars
Cross-platform local AI desktop app. The SenseVoice integration was merged in #7436 and uses sherpa-onnx for offline automatic Chinese, English, Japanese, Korean, and Cantonese recognition on macOS, Linux, and Windows. It is available in the v1.4.159-rc.1 prerelease; v1.4.158 stable predates the integration.
DesktopOfflineSenseVoice
2.2K stars
Open-source Wispr Flow alternative. Desktop voice workflow integrating FunASR local models with configurable LLMs.
DesktopVoice Input
1.2K stars
Multi-function desktop app with audio/video processing, image editing, and AI-enhanced speech transcription.
DesktopToolkit
704 stars
Privacy-first local voice input tool. Converts speech to text via hotkey and auto-types into any app. Supports MCP integration.
Voice InputPrivacy
325 stars
Real-time audio translation. Captures system audio + mic, uses SenseVoice for ASR, then LLM streaming translation.
TranslationReal-time
139 stars
High-performance Linux offline Chinese voice input. Based on FunASR, 0.1s instant display, IBus/Fcitx5 support.
LinuxInput Method
103 stars
Free offline voice-to-text for macOS. Push-to-talk, works in any app. Fully local processing with SenseVoice.
macOSVoice Input
74 stars
Voice-driven writing, input, and cross-app work for your desktop. Speech-to-text with AI refinement.
DesktopWriting
73 stars
Open-source offline voice dictation — a free Typeless alternative. 100% local, SenseVoice + DirectML, ideal for air-gapped environments.
OfflineSecurity
Voice Assistants & Agents
12.8K stars
Digital human agent framework connecting 2.5D/3D avatars with LLMs. Uses FunASR for real-time speech recognition.
Digital HumanAgent
13.4K stars
Open-source AI digital human toolkit. Offline video generation and real-time interaction. Uses FunASR for speech recognition.
Digital HumanVideo
7.1K stars
Chinese voice assistant / smart speaker on Raspberry Pi. Supports ChatGPT, brain-computer interaction. FunASR as ASR engine.
IoTAssistant
6.9K stars
Coding agent from your phone, desktop, and CLI. Uses Paraformer and SenseVoice for speech recognition via Sherpa ONNX.
CodingAgent
3.3K stars
Digital avatar conversational system. Combines ASR, LLM, and TTS for natural dialogue with virtual characters. Uses FunASR.
Digital AvatarDialogue
2.1K stars
Extract audio/video content into structured markdown notes. Uses FunASR for accurate transcription.
NotesProductivity
1.7K stars
GPT-4o-style voice chatbot. Full ASR + LLM + TTS pipeline for natural voice conversations. Powered by FunASR.
Voice ChatGPT-4o
AI Platforms & Frameworks
75K+ stars
A visual workflow and agent platform with a dedicated FunASR speech-to-text component merged in #14225. It defaults to local
/v1 and sensevoice, with optional gateway authentication and language selection; add a self-hosted service host to Langflow's SSRF allow-list when that protection is enabled.WorkflowLocal FunASROpenAI-compat
88.6K+ stars
An open-source RAG and agent platform. Merged #17388 lets local FunASR use the default
http://localhost:8000/v1 endpoint without an API key, preserves Bearer authentication when a key is configured, and validates the model name before building the multipart transcription request instead of panicking on missing configuration.RAGLocal FunASROptional auth
313 stars
An audio-model evaluation framework. Merged #47 registers
fun-asr-nano-2512, pins the Hugging Face model revision and runtime dependencies, and evaluates through an isolated subprocess. A fixed-sample H100 smoke test passed before merge; it is not evidence for every dataset, language, or deployment target.EvaluationFun-ASR-NanoPinned revision
58K stars
Dataset annotation and WebUI transcription support Fun-ASR-Nano, SenseVoice, and classic FunASR models. Merged #2824 corrects the project requirement to
Transformers >=4.51,<5: 4.51 is the first version in its validation matrix that constructs Qwen3, preventing KeyError: qwen3 before Fun-ASR-Nano transcription starts. See also the runtime fallback and backend guide.TTSDataset transcriptionQwen3
33K stars
Self-hosted OpenAI alternative. FunASR integration as speech-to-text backend (PR in review).
LLMSelf-hosted
54.3K stars
OpenAI-compatible AI Gateway. FunASR / SenseVoice can be routed through
custom_openai as a self-hosted transcription endpoint.AI GatewayOpenAI-compatible
12.5K stars
Voice and multimodal conversational AI framework. FunASR as community STT integration.
Conversational AI
9.8K stars
Open-source audio, music, and speech generation toolkit from OpenMMLab. Uses FunASR for TTS evaluation and data processing.
Audio ToolkitOpenMMLab
150K stars · 352 installs
The official FunASR plugin 0.1.1 is live on Dify Marketplace with SenseVoice, Fun-ASR-Nano, and Paraformer presets, supporting 25 MB uploads.
Official PluginSpeech-to-Text
9.3K stars
Distributed inference framework. Built-in FunASR speech recognition backend with one-click ASR model deployment.
InferenceDistributed
5K stars
All-in-one AI digital human system with video synthesis, voice cloning. Integrates FunASR for speech recognition.
Digital HumanAIGC
1.6K stars
Lightweight multimodal model combining vision, audio, and language understanding. Uses FunASR for speech recognition module.
MultimodalLLM
95 stars
ComfyUI custom nodes for SenseVoice and CosyVoice. Visual workflow builder for speech recognition and synthesis.
ComfyUIWorkflow
SenseVoice Community Extensions
894 stars
Enhanced SenseVoice with high-accuracy word-level timestamps. Same speed as original model.
TimestampsSenseVoice
541 stars
API and WebSocket server for SenseVoice. Supports VAD detection, real-time streaming, and speaker verification.
APIWebSocket
451 stars
Pseudo-streaming SenseVoice with hotword boosting. Low-latency near-realtime speech recognition.
StreamingHotwords
111 stars
Enterprise-grade SenseVoice inference with ONNX Runtime. No PyTorch dependency, production-ready deployment.
ONNXDeployment
109 stars
FastAPI wrapper for SenseVoice with ONNX inference. Smaller footprint, quantized models, GPU acceleration.
FastAPIQuantized
93 stars
SenseVoice API service compatible with OneAPI. Unified interface for managing multiple speech recognition models.
OneAPIAPI
Cross-Platform Inference
5K+ stars
Cross-platform speech processing with ONNX. Runs SenseVoice and Paraformer on iOS, Android, Raspberry Pi, and browsers.
MobileEdge
1.1K+ stars
The merged native Fun-ASR-Nano implementation runs with pure C++ and GGML on CPU or CUDA, exposing a CLI and local OpenAI-compatible HTTP transcription service. Read the pinned usage and verification guide.
Fun-ASR-NanoC++ / GGML
7.7K+ stars
An Apple Silicon-native MLX audio runtime. Main now documents Fun-ASR-Nano conversion, Python/CLI, and OpenAI-compatible transcription, with uniform hotwords/context support merged in #885. This path is transcription-only with no timestamps; VAD and diarization remain external. See the model scope in the Fun-ASR repository.
Apple SiliconFun-ASR-NanoOpenAI-compat
20.7K+ stars
A multi-agent interactive classroom with a native local FunASR provider merged in #1044. Follow the setup guide and set
ASR_FUNASR_BASE_URL to send WAV audio to the OpenAI-compatible transcription endpoint and select SenseVoiceSmall, Paraformer, or Fun-ASR-Nano without an API key. See the FunASR repository for server capabilities and deployment boundaries.Local ASRMulti-agentOpenAI-compat
608 stars
Cross-platform ASR inference library based on ONNX Runtime and FunASR. Ready to use, supports Chinese-English mixed recognition.
ONNXCross-Platform
550 stars
C/C++ implementation of SenseVoice model. Pure C++ inference with no Python dependency.
C++Embedded
142K stars
Hugging Face Transformers library. Fun-ASR-Nano integration (PR in review) — use FunASR models with the familiar HF API.
ML Framework
211 stars
OpenAI-compatible speech server supporting FunASR, Whisper, Bark, and CosyVoice backends.
API ServerOpenAI-compat
136 stars
Out-of-the-box local speech service. Microservice architecture, compatible with Alibaba Cloud Speech API and OpenAI TTS API.
API ServerSelf-hosted
113 stars
C++ inference engine based on GGML. CPU/CUDA support, real-time mic streaming, single GGUF file deployment.
GGMLC++
79 stars
Multi-model ASR inference solution supporting Paraformer, SenseVoice, Whisper, and more. ONNX-based, multi-scenario ready.
Multi-modelONNX
Get Started
The fastest way to try FunASR:
pip install funasr # Python API from funasr import AutoModel model = AutoModel(model="iic/SenseVoiceSmall") result = model.generate(input="audio.wav") # Or start an OpenAI-compatible server pip install vllm fastapi uvicorn python-multipart funasr-server --device cuda