AI Audio
-
Deepgram API AIStartMap
Unified speech AI API platform, providing real-time speech recognition, synthesis, intelligent analysis and conversational AI Agent infrastructure
-
Ecrett Music ecrett
AI royalty-free soundtrack generation tool for creators, quickly generate commercially available music by scene, mood and genre
-
3D-Speaker
ModelScope is an open source multi-modal voiceprint recognition and speaker separation toolkit.
-
Eleven Music ElevenLabs
ElevenLabs’ AI music generation platform combines a creator workbench, commercial licensing marketplace, and embeddable Music API
-
ACE-Step ACE-Step
ACE Studio and StepFun jointly open source the basic model of music generation to support efficient generation and lyrics editing.
-
Fish Audio Fish
A text-to-speech and speech cloning platform focusing on emotional expressiveness, with the open source Fish-Speech model behind it
-
Aero-1-Audio Aero-1-Audio
A lightweight audio model with only 150 million parameters, supporting 15 minutes of continuous audio processing
-
Fun Asr Fun-ASR
Tongyi Lab’s open source end-to-end speech recognition model supports Chinese dialects and 31 languages.
-
AInterview AInterview
You are the guest and AI is the host, generating a podcast interview in a few minutes
-
Fun-AudioGen-VD Alibaba Tongyi Lab
The timbre design model launched by Alibaba Tongyi Lab supports natural language description to generate timbre, emotion and scene-based audio.
-
Gemini 3.1 Flash TTS Google
Next-generation text-to-speech model launched by Google supports 70+ languages and audio tag director-level control
-
AssemblyAI API AssemblyAI
A voice AI API platform for developers, providing a full-link infrastructure from transcription, understanding to voice agent











