Forgot your password?
typodupeerror

Comment Re:Is this supposed to be new? (Score 1) 49

1) 30B, dense. Optimized to fit Q4 quantized in a 24GB card with a speculative decoding model as well.

2) Because reporters don't know what weights are and assume you don't know either.

3) Way better than Gemma 4 on text tasks, slightly better on multimodal. Numbers below are all: Benchmark: Muse Glimmer score Gemma 4 31B score difference

Artificial Analysis Intelligence Index: Muse Glimmer: 35 30 +5
MCP Atlas (Public): 75.5 54.2 +21.3
DeepSearch QA: 74.6 61.7 +12.9
SWE-Bench Pro: 51.2 36.9 +14.3
SWE-Bench Verified: 76.0 66.6 +9.4
OSWorld-Verified: 65.9 58.5 +7.4
GAIA2: 43.3 36.4 +6.9
WildClawBench: 47.6 37.6 +10.0
TerminalBench 2.1: 51.7 43.4 +8.3
Tau3-Banking: 23.5 15.1 +8.4
MMMU Pro: 74.0 73.0 +1.0
Charxiv Reasoning: 78.8 77.7 +1.1
OmniDocBench v1.5: 75.8 72.5 +3.3
ScreenSpot Pro: 75.4 75.9 -0.4

4) 128k tokens

5) No, sadly.

Comment Re:'24 GB or 32 GB envelope' (Score 1) 49

What you want is a DGX Spark.

It costs $4000.

And yes, there are fundamental advantages to cloud services, such as large-scale batching, little idle downtime, high speed, and hardware optimized to the specific models / serving needs. That said, one can weigh that off against sovereign control over your server...

Comment Re:What card? (Score 1) 49

Define "tolerable".

The model in question - Muse Glimmer - with DFlash/speculative decoding - will probably get you ~60 to 124 tok/s on a 3090 (a quite dated GPU). For a 5090, it's said to clock in at 233,4 tok/s. On CPU you're looking at maybe 3-5 tok/s.

If you call that "tolerable", I guess you're more patient than me? And as mentioned, you're not just wasting time, but also wasting a lot of power too - CPU is a very power-inefficient way to run ML models.

If you insist on CPU, this isn't the right kind of model anyway. You want to take advantage of the fact that you probably have lots of (comparably cheap) RAM, and compensate for the fact that you have (comparably) terrible memory bandwidth, and for that, you want a MoE with a high total parameters but a low active parameters. Not a dense model like this.

Then I was basically aiming at Mac Minis

That's very much a special case which you didn't mention in your post that I responded to, but still the answer is "meh". You couldn't run it at all on a 16GB Mac Mini, and I think you'd struggle to run it at all on a 24GB (because you have to share the ram with the OS, the inference server, etc). For the base Mini you might get 10-12 tok/s, and for the M2 Pro / M4 Pro, maybe 25-35 tok/s. Still pretty far from a GPU, though.

This model is designed for >= 24GB GPUs.

Comment Re:What card? (Score 2) 49

MoEs don't save RAM (for a given quality), they increase it. You have to store all of the parameters in memory, not just the active params. But inference only uses a subset of the total params for each token, so it reduces the memory bandwidth requirements and improves token generation rate. But this comes at the cost of a higher total param count for a given quality.

Comment Re:Zuck Wakes Up (Score 1) 49

Meta has always been releasing open models (Llama was famously the first powerful open model). The change has actually been in the opposite direction, with Muse Spark *not* being open.

Zuck descrbed his motivation way back when, about how they got burned with Facebook on app stores, in that Apple and Google could basically bully them however they wanted, on whatever extractive terms they wanted, and there was nothing Meta could do about it. He's now paranoid about "others controlling the platform", and wanted to make sure that doesn't happen with AI, that they have their own AI base to work with.

Comment Re: What card? (Score 2) 49

This really isn't a good model for CPU. For CPU, you want a MoE with a large number of total params but a tiny number of active params. Something like DeepSeek V4 Flash 0731 if you have at least 128GB of RAM - you might get 2-3 tok/s or so on that. The goal is to minimize the memory bandwidth requirements per token, at the cost of a greater total RAM footprint.

For GPU, you're highly VRAM limited but not bandwidth limited, so your best option is generally a dense model (non-MoE) with speculative decoding to make up for the performance limitations.

Comment Re:How can you watermark a song ... (Score 1) 33

Actually, no, what they usually do is much more insidious: fingerprinting rather than watermarking. The fingerprint isn't actually included in the audio, it's included in a database. If they want to tell if the track was generated, they just try to match the fingerprint in the database. Can't filter it out of the track like you can with a watermark.

Comment Re:Open source solution (Score 4, Informative) 33

The two mainstream options (both about half a year old now, we're due for something better) are Ace Step-1.5 and Stable Audio 3 Medium.

Ace Step has full vocals with its music. See here for examples. The downside in my opinion (judge for yourself) is that it has Suno's flaw of sounding too "clean", "mainstream" and "uncreative", except even moreso (Udio was always much better than Suno at this, albeit "less well behaved" - but Udio is out of the game now).

Stable Audio 3 doesn't do vocals, though it has other useful features, like inpainting (great for fixing glitches in recorded tracks for example). In my view, it sounds a lot better. It can also be used to create sound effects. License is a bit more restrictive, though, if you actually care about that, but generally won't affect the average user.

Both are trained on fully licensed data and both are open-weights. So even if the devs released a version that included some sort of watermarking in the future (or someone developed a fingerprinting algo for the existing versions), you could just finetune them to break that.

TL/DR, if you want open source and want to just churn out a full track: choose Ace Step 1.5.
If you want open source and want to do instrumental tracks or to supplement manual work (such as your own vocals) or do audio editing: chose Stable Audio 3.

Comment Re:Won't affect me (Score 1) 207

I often wonder how all these people who believed that COVID vaccines were a plot to sterilize or murder humanity have justified to themselves the continued-unchanged existence of babies-in-specific and humans-in-general, respectively.

Comment Re:Business Interests (Score 1) 183

Hawaiian eruption last a long time and the lava flows persist for years.

Your effusive eruptions last on average no longer than our effusive eruptions. The fact that Iceland also has phreatic eruptions doesn't change the fact that most of ours are effusive basalt, just like yours, with corresponding eruptive dynamics. Also, FYI, for a given total effusion volume, slow eruptions are easier to divert. Diversion is much more challenging when an equivalent volume of lava is effused in a short period of time. And it's not like you have vastly larger average lava flows than we do. The largest lava flow on Earth in the Holocene (modern era) is in Iceland (the THjórsárhraun - 900km2, 25km3, 70km underground travel followed by 130-140km on the surface). For comparison, Hawaii's largest holocene flow is the Ailaau Flow at 430km2, 4-5km3, and 40km.

If we want to limit to more recent flows (since Ailaau was fairly recent, in case you think Iceland stopped having big eruptions), Iceland's biggest *after* Alia'au was Laki (1783–1784 AD), at 600km2, 15km3, and 80km. It lasted for 8 months to 2 years, depending on whether you count the concurrent Grímsvötn eruption as part of it. So much gas was released from that huge volume of lava that despite us being nearly arctic (polar volcanoes have much lower near-term climate impacts than equatorial volcanoes), it caused a winter so severe that the Mississippi froze at New Orleans and ice was seen floating in the Gulf of Mexico.

We simply get bigger effusive eruptions than you, full stop. Not all eruptions, of course, both of us have a huge variance in eruption scales and dynamics. But trying to claim ours are somehow "easier" is just not in line with the facts. Also, Hawaiian eruptions tend to occur on terrain that slopes significantly more consistently toward the sea, aka where you want to divert lava toward; on flatter or less-consistently-angled terrain, lava "piles up" more and makes it easier to overtop berms.

There is a practical side that it is easier to move people and build coffee shacks

I'm not talking about "coffee shacks", I'm talking about whole communities and beloved places. In particular, I'm thinking about the 2018 Lower Puna eruption. Of course the homes near the vents in Leilani Estates were doomed, but there is absolutely NO reason why you should have lost anywhere close to as much as you did so far from the vents. Losing Kapoho Bay for example was simply cultural fatalism. I watched a person crying on the news saying she didn't understand why Pele chose to take such a beautiful place. It's not Pele who took it, it's your insane resistance against doing anything to protect yourselves. That was one of the easiest possible flows to divert. Most of the lava fronts moved at a crawl, you had warning long in advance of the eruption and a very slow eruptive rampup (didn't really take off until June), and you had a steep slope to the south into unpopulated areas you could have easily diverted it down. Losing all that was a choice.

"Oh, we can't!" is just an excuse for your fatalism that gives you an excuse to not put forth the effort to actually do it. It's not the 1940s anymore, we're not chucking surplus WWII bombs at lava tubes anymore. Lava diversion is a science and is routinely effectively done. There are limits, but you were nowhere near said limits for sites far from the vents. There are costs, but they are nowhere near the costs that you paid by refusing to act.

Slashdot Top Deals

The disks are getting full; purge a file today.

Working...