Average Ratings 0 Ratings
Average Ratings 0 Ratings
Description
Celeris-1 stands out as a swift and versatile language model platform, complemented by a diffusion model that achieves cutting-edge intelligence at unprecedented speeds. Unlike conventional autoregressive models that generate tokens sequentially, Celeris employs a diffusion-based inference architecture that allows for simultaneous generation, resulting in response times that can be measured in mere milliseconds. On the MMLU-Pro benchmark, Celeris-1 boasts an impressive accuracy of 75.9% while achieving a median response time of 158 milliseconds and producing an astonishing 1,664 output tokens per second, positioning it closely to leading models but operating over ten times faster. This powerful model is accessible through an API that is compatible with OpenAI, enabling developers to seamlessly integrate it into existing SDKs and applications with minimal modifications. Additionally, it supports streaming capabilities for real-time applications, allowing for response times as low as 24 milliseconds without any buffering or delays, making it an ideal choice for interactive use cases. Overall, Celeris-1 represents a significant advancement in the efficiency and performance of language models.
Description
DiffusionGemma is an innovative open model that investigates text diffusion, representing a remarkably rapid method for generating text. Released under the Apache 2.0 license, this 26 billion parameter Mixture of Experts (MoE) model advances beyond the usual sequential token generation typical of autoregressive models. Instead, it produces entire blocks of text at once, achieving text generation speeds that are up to four times faster on GPUs. Drawing from the parameter efficiency of the Gemma 4 family and Gemini Diffusion research, DiffusionGemma incorporates a unique diffusion head that enhances generation speed significantly. It is particularly aimed at researchers and developers looking to optimize speed-sensitive, interactive local workflows, including in-line editing, swift iterations, and non-linear narrative forms. By reallocating the decode bottleneck from memory bandwidth to computational power, it can produce over 1,000 tokens per second on a single NVIDIA H100 and more than 700 tokens per second on an NVIDIA GeForce RTX 5090. This breakthrough allows for a new level of efficiency in text generation that could reshape various applications in natural language processing.
API Access
Has API
API Access
Has API
Pricing Details
$0.20 per 1M tokens
Free Trial
Free Version
Pricing Details
Free
Free Trial
Free Version
Deployment
Web-Based
On-Premises
iPhone App
iPad App
Android App
Windows
Mac
Linux
Chromebook
Deployment
Web-Based
On-Premises
iPhone App
iPad App
Android App
Windows
Mac
Linux
Chromebook
Customer Support
Business Hours
Live Rep (24/7)
Online Support
Customer Support
Business Hours
Live Rep (24/7)
Online Support
Types of Training
Training Docs
Webinars
Live Training (Online)
In Person
Types of Training
Training Docs
Webinars
Live Training (Online)
In Person
Vendor Details
Company Name
Celeris-1
Country
United States
Website
celeris.ai/
Vendor Details
Company Name
Founded
1998
Country
United States
Website
blog.google/innovation-and-ai/technology/developers-tools/diffusion-gemma-faster-text-generation/
Product Features
Product Features
Alternatives
No Alternatives