An API powered by Google's AI technology allows you to accurately convert speech into text. You can accurately caption your content, provide a better user experience with products using voice commands, and gain insight from customer interactions to improve your service. Google's deep learning neural network algorithms are the most advanced in automatic speech recognition (ASR). Speech-to-Text allows for experimentation, creation, management, and customization of custom resources. You can deploy speech recognition wherever you need it, whether it's in the cloud using the API or on-premises using Speech-to-Text O-Prem. You can customize speech recognition to translate domain-specific terms or rare words. Automated conversion of spoken numbers into addresses, years and currencies. Our user interface makes it easy to experiment with your speech audio.
Learn more

Most AI video tools hand you a black box: closed weights, a subscription, and no way to see what is happening under the hood. LTX takes the opposite approach. Built by Lightricks, LTX is an open foundation model that generates and simulates across video, audio, and the physical world, and it puts the weights, the code, and the control in your hands.
At the center of the model is LTX-2.5, a 22B-parameter dual-stream diffusion transformer that produces native 4K video at up to 50 frames per second, with audio and video generated together in a single pass rather than stitched together afterward. Artificial Analysis, an independent benchmarking group, currently ranks LTX among the top three AI video models in the world.
You choose how you want to use it. Download the open weights and run LTX-2.5 on your own hardware. License the model for on-premise deployment backed by enterprise support. Or build directly on LTX Studio, the production suite that turns the model into a full creative workflow. Companies like ElevenLabs, Asteria Film Co., Magnopus, and NVIDIA already rely on LTX for their own work.
LTX is not built for one-off social clips. It is infrastructure for teams that generate motion, audio, and physical environments as part of their own products and pipelines.
Learn more
Cartesia Sonic-3.6
Sonic is an advanced text-to-speech model designed specifically for real-time voice agents, featuring a natural delivery system with a response time of less than 90 milliseconds and supporting over 40 languages seamlessly. Its primary aim is to facilitate effortless voice interactions, characterized by a tone that adapts to various contexts, a steady pacing, and speech that aligns with the natural flow of conversation. Sonic automatically interprets the emotional nuances within transcripts, adjusting its delivery accordingly, and allows for the direct insertion of non-verbal cues like laughter into the spoken text. Faithful to the original transcripts, the model generates clear audio across different languages and voice options while effortlessly managing alphanumeric data, including order and phone numbers, email addresses, and IDs, without requiring any prior processing. Its context-aware pronunciation ensures that heteronyms are articulated correctly based on surrounding terms, and customizable pronunciation dictionaries empower teams to dictate how specific proper nouns and industry-related terminology should be pronounced. This comprehensive approach not only enhances the quality of interactions but also tailors the user experience to meet diverse communication needs.
Learn more
ElevenLabs
The most versatile and realistic AI speech software ever. Eleven delivers the most convincing, rich and authentic voices to creators and publishers looking for the ultimate tools for storytelling. The most versatile and versatile AI speech tool available allows you to produce high-quality spoken audio in any style and voice. Our deep learning model can detect human intonation and inflections and adjust delivery based upon context. Our AI model is designed to understand the logic and emotions behind words. Instead of generating sentences one-by-1, the AI model is always aware of how each utterance links to preceding or succeeding text. This zoomed-out perspective allows it a more convincing and purposeful way to intone longer fragments. Finally, you can do it with any voice you like.
Learn more