AI Audio
-
GPT Realtime Whisper OpenAI
OpenAI’s low-latency streaming transcription model for real-time subtitles, meeting recordings, and speech workflows
-
AudioCraft
Meta open source AI audio generation framework supports music, sound effects and compression encoding
-
Boomy Boomy
AI platform that lets anyone create original music in seconds
-
Cleanvoice AI Cleanvoice
AI audio post-processing tools for podcasting and interview scenarios, the real value is to clean up idioms, pauses and noise in batches instead of just transcribing
-
Deepgram API AIStartMap
Unified speech AI API platform, providing real-time speech recognition, synthesis, intelligent analysis and conversational AI Agent infrastructure
-
Ecrett Music ecrett
AI royalty-free soundtrack generation tool for creators, quickly generate commercially available music by scene, mood and genre
-
3D-Speaker
ModelScope is an open source multi-modal voiceprint recognition and speaker separation toolkit.
-
Eleven Music ElevenLabs
ElevenLabs’ AI music generation platform combines a creator workbench, commercial licensing marketplace, and embeddable Music API
-
ACE-Step ACE-Step
ACE Studio and StepFun jointly open source the basic model of music generation to support efficient generation and lyrics editing.
-
Fish Audio Fish
A text-to-speech and speech cloning platform focusing on emotional expressiveness, with the open source Fish-Speech model behind it
-
Aero-1-Audio Aero-1-Audio
A lightweight audio model with only 150 million parameters, supporting 15 minutes of continuous audio processing
-
Fun Asr Fun-ASR
Tongyi Lab’s open source end-to-end speech recognition model supports Chinese dialects and 31 languages.











