Define "tolerable".
The model in question - Muse Glimmer - with DFlash/speculative decoding - will probably get you ~60 to 124 tok/s on a 3090 (a quite dated GPU). For a 5090, it's said to clock in at 233,4 tok/s. On CPU you're looking at maybe 3-5 tok/s.
If you call that "tolerable", I guess you're more patient than me? And as mentioned, you're not just wasting time, but also wasting a lot of power too - CPU is a very power-inefficient way to run ML models.
If you insist on CPU, this isn't the right kind of model anyway. You want to take advantage of the fact that you probably have lots of (comparably cheap) RAM, and compensate for the fact that you have (comparably) terrible memory bandwidth, and for that, you want a MoE with a high total parameters but a low active parameters. Not a dense model like this.
Then I was basically aiming at Mac Minis
That's very much a special case which you didn't mention in your post that I responded to, but still the answer is "meh". You couldn't run it at all on a 16GB Mac Mini, and I think you'd struggle to run it at all on a 24GB (because you have to share the ram with the OS, the inference server, etc). For the base Mini you might get 10-12 tok/s, and for the M2 Pro / M4 Pro, maybe 25-35 tok/s. Still pretty far from a GPU, though.
This model is designed for >= 24GB GPUs.