Average Ratings 1 Rating

Total
ease
features
design
support

Average Ratings 0 Ratings

Total
ease
features
design
support

No User Reviews. Be the first to provide a review:

Write a Review

Description

Muse Spark 1.1 is Meta’s upgraded multimodal reasoning model designed to support advanced agentic workflows, coding tasks, computer use, and complex tool orchestration. Developed by Meta Superintelligence Labs, it builds on Muse Spark with major gains in planning, tool use, long-context reasoning, multimodal perception, and real-world task execution. The model can work across external apps and services, native tools, MCP servers, custom skills, browsers, scripts, images, video, PDFs, and audio inputs. Muse Spark 1.1 can act as a main agent by gathering context, creating a plan, and delegating work to parallel subagents, or operate as a subagent that follows instructions and escalates when needed. Its 1 million token context window allows it to retain earlier actions, retrieve information from long workflows, and compact context while preserving critical details. The model is also trained for computer-use tasks, deciding when to automate with scripts and when to interact directly with an interface. In coding workflows, Muse Spark 1.1 can diagnose bugs, implement features, migrate large codebases, generate web applications, take screenshots, identify UI issues, and validate fixes. Its multimodal strengths include visual-to-code generation, detailed image and video captioning, grounded perception, and workflows where seeing, reasoning, and acting happen together. Available through the Meta Model API public preview and in Thinking mode inside Meta AI, Muse Spark 1.1 gives developers and users a more capable foundation for building agents, automations, coding assistants, and multimodal productivity tools.

Description

Muse Voice Transcribe represents Meta’s inaugural venture into real-time audio perception, providing instantaneous automatic speech recognition (ASR), speaker diarization, and endpointing capabilities. This autoregressive multimodal model, part of the Muse Spark series, analyzes audio segments of 80 milliseconds and makes real-time decisions on whether to keep listening or to convert the spoken words into text. The adaptive delay mechanism allows it to adjust the audio context utilized for each word according to the complexity of the speech, thus optimizing the balance between transcription precision and response time. With training encompassing over 70 languages, 25 of which were rigorously validated at the time of its release, the model also seamlessly accommodates arbitrary code-switching, allowing transitions within and across sentences. Furthermore, language, keyword, and contextual biasing features enhance the recognition capabilities for specific names, locations, contacts, or specialized terms. The streaming diarization functionality enables the model to recognize shifts in speakers and can differentiate between more than 20 individual voices. Additionally, the endpointing feature is adept at identifying the commencement of speech and knowing when a user has completed their statement, ensuring a fluid interaction experience. Overall, Muse Voice Transcribe stands out as a cutting-edge tool in the realm of speech recognition technology, merging advanced features with user-friendly application.

API Access

Has API Yes 

API Access

Has API Yes 

Screenshots View All

Screenshots View All

Integrations

.NET Yes 
C Yes 
C++ Yes 
CSS Yes 
Claude Agent SDK Yes 
Continue Yes 
Dart Yes 
Facebook Yes 
Instagram Yes 
JavaScript Yes 
Kotlin Yes 
Kubernetes Yes 
Lua Yes 
Model Context Protocol (MCP) Yes 
OpenAI Agents SDK Yes 
OpenAI Codex Yes 
OpenClaw Yes 
R Yes 
Ruby Yes 
Swift Yes 

Integrations

.NET No 
C No 
C++ No 
CSS No 
Claude Agent SDK No 
Continue No 
Dart No 
Facebook No 
Instagram No 
JavaScript No 
Kotlin No 
Kubernetes No 
Lua No 
Model Context Protocol (MCP) No 
OpenAI Agents SDK No 
OpenAI Codex No 
OpenClaw No 
R No 
Ruby No 
Swift No 

Pricing Details

$1.25 per 1M tokens (input)
$1.25 per million tokens in input, and $4.25 per million tokens of output
Free Trial Yes 
Free Version No 

Pricing Details

No price information available.
Free Trial Yes 
Free Version No 

Deployment

Web-Based Yes 
On-Premises No 
iPhone App No 
iPad App No 
Android App No 
Windows No 
Mac No 
Linux No 
Chromebook No 

Deployment

Web-Based Yes 
On-Premises No 
iPhone App No 
iPad App No 
Android App No 
Windows No 
Mac No 
Linux No 
Chromebook No 

Customer Support

Business Hours No 
Live Rep (24/7) No 
Online Support Yes 

Customer Support

Business Hours No 
Live Rep (24/7) No 
Online Support Yes 

Types of Training

Training Docs Yes 
Webinars No 
Live Training (Online) No 
In Person No 

Types of Training

Training Docs Yes 
Webinars No 
Live Training (Online) Yes 
In Person No 

Vendor Details

Company Name

Meta

Founded

2004

Country

United States

Website

meta.ai

Vendor Details

Company Name

Meta

Founded

2004

Country

United States

Website

research.meta.ai/blog/introducing-muse-voice-transcribe

Product Features

Alternatives

Grok 4.6 Reviews

Grok 4.6

SpaceXAI

Alternatives

MAI-Transcribe-2 Reviews

MAI-Transcribe-2

Microsoft AI
GPT-5.6 Sol Reviews

GPT-5.6 Sol

OpenAI
MAI-Transcribe-1.5 Reviews

MAI-Transcribe-1.5

Microsoft AI
Claude Opus 5 Reviews

Claude Opus 5

Anthropic