Average Ratings 0 Ratings
Average Ratings 0 Ratings
Description
Bland Speech v3 is an innovative text-to-speech model that aims to generate voice audio that closely resembles that of a human, particularly in contexts like phone calls where overly polished speech may come off as inauthentic. This model captures essential human elements such as breathing, stutters, pauses, laughter, and throat-clearing by utilizing performance tags that are acted out rather than simply read. Users have the option to input their own scripts or utilize the Director feature to outline a conversation, allowing Bland to craft the dialogue, timing, and delivery prior to speech generation. Additionally, it offers voice cloning capabilities: a quick clone can be made from approximately 10 seconds of audio, whereas professional-grade cloning requires 30 minutes or more of verified audio, with users affirming that each voice belongs to them. Bland Speech can be accessed via a web studio and a single /v1/speak API endpoint, which employs bearer-key authentication for security. Audio is streamed through HTTP chunked transfer or WebSocket, returning PCM16 WAV files at a sample rate of 44.1 kHz, ensuring high-quality output for diverse applications. This versatility makes Bland Speech an essential tool for developers looking to enhance their audio experiences.
Description
Zyphra is thrilled to unveil the beta release of Zonos-v0.1, which boasts two sophisticated and real-time text-to-speech models that include high-fidelity voice cloning capabilities. Our release features both a 1.6B transformer and a 1.6B hybrid model, all under the Apache 2.0 license. Given the challenges in quantitatively assessing audio quality, we believe that the generation quality produced by Zonos is on par with or even surpasses that of top proprietary TTS models currently available. Additionally, we are confident that making models of this quality publicly accessible will greatly propel advancements in TTS research. You can find the Zonos model weights on Huggingface, with sample inference code available on our GitHub repository. Furthermore, Zonos can be utilized via our model playground and API, which offers straightforward and competitive flat-rate pricing options. To illustrate the performance of Zonos, we have prepared a variety of sample comparisons between Zonos and existing proprietary models, highlighting its capabilities. This initiative emphasizes our commitment to fostering innovation in the field of text-to-speech technology.
API Access
Has API
API Access
Has API
Integrations
Amazon Connect
Bland AI
Cal.com
Calendly
Five9
Genesys Cloud CX
GitHub
HubSpot CRM
Hugging Face
Make
Integrations
Amazon Connect
Bland AI
Cal.com
Calendly
Five9
Genesys Cloud CX
GitHub
HubSpot CRM
Hugging Face
Make
Pricing Details
$0.11 per minute
Free Trial
Free Version
Pricing Details
$0.02 per minute
Free Trial
Free Version
Deployment
Web-Based
On-Premises
iPhone App
iPad App
Android App
Windows
Mac
Linux
Chromebook
Deployment
Web-Based
On-Premises
iPhone App
iPad App
Android App
Windows
Mac
Linux
Chromebook
Customer Support
Business Hours
Live Rep (24/7)
Online Support
Customer Support
Business Hours
Live Rep (24/7)
Online Support
Types of Training
Training Docs
Webinars
Live Training (Online)
In Person
Types of Training
Training Docs
Webinars
Live Training (Online)
In Person
Vendor Details
Company Name
Bland AI
Founded
2023
Country
United States
Website
www.bland.ai/speech
Vendor Details
Company Name
Zyphra
Country
United States
Website
www.zyphra.com/post/beta-release-of-zonos-v0-1
Product Features
Product Features
Text to Speech
API
Adjust Speaking Rate / Pitch
Audio Optimization
Custom Lexicons
Different Voice Choices
Multi-Language Support
Synchronize Speech