Average Ratings 3 Ratings

Total
ease
features
design
support

Average Ratings 0 Ratings

Total
ease
features
design
support

No User Reviews. Be the first to provide a review:

Write a Review

Description

Experience a robust, self-service machine learning platform that enables you to transform models into scalable APIs with just a few clicks. Create an account with Deep Infra through GitHub or log in using your GitHub credentials. Select from a vast array of popular ML models available at your fingertips. Access your model effortlessly via a straightforward REST API. Our serverless GPUs allow for quicker and more cost-effective production deployments than building your own infrastructure from scratch. We offer various pricing models tailored to the specific model utilized, with some language models available on a per-token basis. Most other models are charged based on the duration of inference execution, ensuring you only pay for what you consume. There are no long-term commitments or upfront fees, allowing for seamless scaling based on your evolving business requirements. All models leverage cutting-edge A100 GPUs, specifically optimized for high inference performance and minimal latency. Our system dynamically adjusts the model's capacity to meet your demands, ensuring optimal resource utilization at all times. This flexibility supports businesses in navigating their growth trajectories with ease.

Description

NVIDIA DGX Cloud Serverless Inference provides a cutting-edge, serverless AI inference framework designed to expedite AI advancements through automatic scaling, efficient GPU resource management, multi-cloud adaptability, and effortless scalability. This solution enables users to reduce instances to zero during idle times, thereby optimizing resource use and lowering expenses. Importantly, there are no additional charges incurred for cold-boot startup durations, as the system is engineered to keep these times to a minimum. The service is driven by NVIDIA Cloud Functions (NVCF), which includes extensive observability capabilities, allowing users to integrate their choice of monitoring tools, such as Splunk, for detailed visibility into their AI operations. Furthermore, NVCF supports versatile deployment methods for NIM microservices, granting the ability to utilize custom containers, models, and Helm charts, thus catering to diverse deployment preferences and enhancing user flexibility. This combination of features positions NVIDIA DGX Cloud Serverless Inference as a powerful tool for organizations seeking to optimize their AI inference processes.

API Access

Has API Yes 

API Access

Has API Yes 

Screenshots View All

Screenshots View All

Integrations

Llama Yes 
AI SpendOps Yes 
Code Llama Yes 
Codestral Mamba Yes 
CoreWeave No 
Google Cloud Platform No 
Helm No 
Higgs Audio / Avatar Yes 
Llama 2 Yes 
Llama 3 Yes 
Llama 3.1 Yes 
Mathstral Yes 
Microsoft Azure No 
Ministral 8B Yes 
Mistral 7B Yes 
Mistral AI Yes 
Mistral NeMo Yes 
NVIDIA AI Foundations No 
Oracle Cloud Infrastructure No 
Pixtral Large Yes 

Integrations

Llama Yes 
AI SpendOps No 
Code Llama No 
Codestral Mamba No 
CoreWeave Yes 
Google Cloud Platform Yes 
Helm Yes 
Higgs Audio / Avatar No 
Llama 2 No 
Llama 3 No 
Llama 3.1 No 
Mathstral No 
Microsoft Azure Yes 
Ministral 8B No 
Mistral 7B No 
Mistral AI No 
Mistral NeMo No 
NVIDIA AI Foundations Yes 
Oracle Cloud Infrastructure Yes 
Pixtral Large No 

Pricing Details

$0.70 per 1M input tokens
Free Trial Yes 
Free Version No 

Pricing Details

No price information available.
Free Trial No 
Free Version No 

Deployment

Web-Based Yes 
On-Premises No 
iPhone App No 
iPad App No 
Android App No 
Windows No 
Mac No 
Linux No 
Chromebook No 

Deployment

Web-Based Yes 
On-Premises No 
iPhone App No 
iPad App No 
Android App No 
Windows No 
Mac No 
Linux No 
Chromebook No 

Customer Support

Business Hours No 
Live Rep (24/7) No 
Online Support Yes 

Customer Support

Business Hours Yes 
Live Rep (24/7) No 
Online Support Yes 

Types of Training

Training Docs Yes 
Webinars No 
Live Training (Online) No 
In Person No 

Types of Training

Training Docs Yes 
Webinars Yes 
Live Training (Online) Yes 
In Person Yes 

Vendor Details

Company Name

Deep Infra

Website

deepinfra.com

Vendor Details

Company Name

NVIDIA

Founded

1993

Country

United States

Website

developer.nvidia.com/dgx-cloud/serverless-inference

Product Features

Machine Learning

Deep Learning No 
ML Algorithm Library No 
Model Training No 
Natural Language Processing (NLP) No 
Predictive Modeling No 
Statistical / Mathematical Tools No 
Templates No 
Visualization No 

Product Features

Alternatives

Alternatives

SambaNova Reviews

SambaNova

SambaNova Systems