Speechmatics
Speechmatics is the most accurate and inclusive speech-to-text API ever released.
Speechmatics is the world’s leading expert in Speech Technology, combining the latest breakthroughs in AI and ML to unlock the business value in human speech.
Businesses use Speechmatics worldwide to accurately understand and transcribe human-level speech into text regardless of demographic, age, gender, accent, dialect, or location in real-time and on recorded media. Combining these transcripts with the latest AI-driven speech capabilities, businesses build products that utilize summarization, topic detection, sentiment analysis, translation, and more.
How is Speechmatics different?
* The most accurate speech recognition on the market
* 55 languages with vast accent and dialect coverage
* Cloud-based or on-premises deployment options for data security
* Real-time transcription with low latency and high accuracy
* Real-time translation with 69 language pairs
* Speech Understanding features such as Summaries, Sentiment, Topic Detection, Chapters, Audio Events
* Fast and secure transcriptions for pre-recorded audio
* Automatic translation and language identification
* A culture of R&D in deep learning and speech recognition
Learn more
Twilio Voice
Create a scalable voice experience with the API that connects millions globally. With Twilio Voice, you can build unique phone call experiences with one API, to create, receive, control and monitor calls with just a few lines of code. Customize your experience the way you want by using a wide range of customization resources, such as our Voice SDK, speech recognition, Interactive Voice Response (IVR), and recording transcriptions.
Whether you're looking to set up global conferencing or alerts & notifications, Twilio has the support you need for building with Voice, such as our Twilio Runtime and Studio developer tools. Find docs, code samples, and helper libraries to start building today.
Learn more
Google Cloud Speech-to-Text
An API powered by Google's AI technology allows you to accurately convert speech into text. You can accurately caption your content, provide a better user experience with products using voice commands, and gain insight from customer interactions to improve your service. Google's deep learning neural network algorithms are the most advanced in automatic speech recognition (ASR). Speech-to-Text allows for experimentation, creation, management, and customization of custom resources. You can deploy speech recognition wherever you need it, whether it's in the cloud using the API or on-premises using Speech-to-Text O-Prem. You can customize speech recognition to translate domain-specific terms or rare words. Automated conversion of spoken numbers into addresses, years and currencies. Our user interface makes it easy to experiment with your speech audio.
Learn more
Rubidium
Rubidium allows leading companies to embed voice commands or text to speech into their products. Voice Trigger is an engine that is "always on" and listens to your voice. It wakes up when the correct "magic word" is said. Voice Trigger uses a miniature footprint Automatic Speech Recognition engine to distinguish between the trigger phrase, the rest of speech, sounds, and noise. Automated Speech Recognition is able to control any function through voice commands. You can use it to accept and reject calls, set up and install devices (pairing, calibration and interconnection), and so on. Voice dialing, music streaming control, and music selection. Today, Rubidium technology can be found in more than 50 million consumer products. Customers and partners include leading global brands like RIM (Blackberry), GN Netcom(Jabra), Panasonic and Uniden.
Learn more