Meta logo

Muse Glimmer 30B

Open weights · Apache 2.0

Meta's first open release since Llama 4: a 29.6B dense model (including a 1.8B perception encoder) distilled from Muse Spark, tuned for agentic tasks on a single consumer GPU, under Apache 2.0.

Price via Together
Input
$0.35 / 1M tokens
Output
$1.50 / 1M tokens

About $6.50 for 10M input and 2M output tokens. Estimate yours

Together Meta 2 providers

Key Specifications

Context window
131K tokens
Knowledge cutoff
Inputs
Text, Image Outputs: Text

Muse Glimmer 30B pricing by provider

Provider Input / 1M tokens Output / 1M tokens Cost at 10M in + 2M out
Together logo Together $0.35 $1.50 $6.50 View
Meta logo Meta Creator Open weights, no hosted price View

Heads up: We do our best to keep these specs & prices accurate. However, cloud costs may fluctuate based on region, usage, and other factors not listed here. These are estimates based on common setups and are for informational purposes only. Always verify current rates & exact specs with the provider before provisioning. LLM rates are base-tier, on-demand prices per 1M tokens; cached-input, batch and long-context tiers are not included.

Compare every model at this volume in the LLM cost calculator.

Capabilities

Function calling

Function calling

Connect to external tools, APIs, and systems.

Estimated cost to self-host Muse Glimmer 30B

Muse Glimmer 30B has 29.6B parameters. At 4-bit it needs about 18 GB of GPU memory, at 8-bit 36 GB and at BF16 71 GB, counting 20% on top of the weights for KV cache and runtime overhead.

Precision Memory needed Cheapest rentals that fit Per month, 24/7
4-bit 18 GB Nvidia logo RTX 3090 ($0.13/hr) Nvidia logo RTX 3090 Ti ($0.19/hr) $94
8-bit 36 GB Nvidia logo RTX A6000 ($0.56/hr) Nvidia logo A40 ($0.74/hr) $403
BF16 71 GB Nvidia logo A100 ($1.76/hr, tight) Nvidia logo RTX PRO 6000 ($2.19/hr) $1,267

$0.13 an hour (RTX 3090) buys about 240K tokens an hour at the cheapest hosted rate we track ($0.35 in / $1.50 out per 1M tokens on Together, 5:1 input to output); below that volume the API is cheaper, before idle time. Compare hosted costs.

Memory is parameters × bytes per weight at each precision, plus 20% for KV cache and runtime overhead. Rentals are the cheapest cards that hold it, at provider-weighted median on-demand prices for the week of August 17, 2026, consumer cards included, in nodes of up to 8 GPUs; a fit with under 15% headroom is marked tight. A month is 720 hours. See the best-value GPUs guide for the same table across models, and cloud GPU pricing for every card.

More from Meta

Model Context Input / 1M Output / 1M
Meta logo Llama 4 Maverick Open weights 1M $0.25 $0.95
Meta logo Llama 3 70B Open weights 8K $0.65 $2.75
Meta logo Llama 4 Scout Open weights 10M $0.17 $0.65
Meta logo Llama 3.3 70B Open weights 128K $0.14 $0.40
Meta logo Muse Spark 1.1 1M $1.25 $4.25
Meta logo Muse Spark 1.2 1M $1.25 $4.25
Meta logo Llama 3 8B Open weights 8K $0.05 $0.25
Meta logo Llama 3.1 8B Open weights 131K $0.02 $0.05

Models near this price

The models nearest this one by input rate, five each way, at each one's cheapest listed provider.

Frequently Asked Questions

How much does Muse Glimmer 30B cost?

Muse Glimmer 30B costs $0.35 per 1M input tokens and $1.50 per 1M output tokens via Together. Meta publishes the weights, so it can also be run on your own hardware.

What does Muse Glimmer 30B cost for 10M input and 2M output tokens?

At Together's rates, 10M input tokens and 2M output tokens cost about $6.50: $3.50 for input and $3.00 for output. Input is prompts and context, output is what the model writes back; a workload that generates more than it reads shifts the cost toward the output rate of $1.50 per 1M tokens.

Which providers offer Muse Glimmer 30B?

2 providers list Muse Glimmer 30B: Together ($0.35 in / $1.50 out) and Meta (open weights, no hosted price).

What is Muse Glimmer 30B's context window?

Muse Glimmer 30B accepts up to 131K tokens of input per request. The context window is the prompt plus any documents, conversation history and tool results sent with it; every token in it is billed at the input rate.

What is Muse Glimmer 30B's knowledge cutoff?

Muse Glimmer 30B's knowledge cutoff is January 2026: its training data runs up to that month and it has no built-in knowledge of later events.

What inputs and outputs does Muse Glimmer 30B support?

Muse Glimmer 30B accepts text and images as input and produces text. Its listed capabilities are function calling.

Can I self-host Muse Glimmer 30B?

Yes. Meta publishes Muse Glimmer 30B's weights under the Apache 2.0 license. At 4-bit it needs about 18 GB of GPU memory; the cheapest rental that holds it is RTX 3090 at $0.13 per hour, about $94 a month. $0.13 an hour (RTX 3090) buys about 240K tokens an hour at the cheapest hosted rate we track ($0.35 in / $1.50 out per 1M tokens on Together, 5:1 input to output); below that volume the API is cheaper, before idle time. The estimate above prices 8-bit and BF16 too.

How does Muse Glimmer 30B compare with Llama 4 Maverick?

Muse Glimmer 30B costs $0.35 per 1M input tokens against Llama 4 Maverick's $0.25, 1.4x more expensive ($1.50 vs $0.95 per 1M output tokens). The context window is 131K tokens against 1M.

What are cheaper alternatives to Muse Glimmer 30B?

Models from other creators with a lower input rate and at least Muse Glimmer 30B's 131K-token context window: Gemini 3.5 Flash-Lite at $0.30 per 1M input tokens (1M context), Gemini 2.5 Flash at $0.30 per 1M input tokens (1M context) and MiniMax-M3 at $0.30 per 1M input tokens (1M context). Rates are the cheapest listed provider for each; whether the quality holds for a given task is a separate question.

Cheaper alternatives to Muse Glimmer 30B