Perplexity Will Open Source Its Faster Lily AI Engine For Apple Silicon 17
BrianFagioli writes: Perplexity has built a local artificial intelligence engine designed specifically for Apple silicon and the Qwen3.6-35B-A3B model. Called Lily, the engine uses a Rust runtime and custom Metal kernels, with neither PyTorch nor MLX in its execution path. Perplexity says Lily averaged 23 percent faster prompt processing and 35 percent faster token generation than MLX-LM on an M5 Max MacBook Pro with 128GB of unified memory. Lily is more specialized than MLX-LM, which supports a much wider range of models and architectures. Perplexity says it plans to release Lily as open source, but the code is not available yet, leaving its performance claims dependent on internal testing for now.
M5 Max MacBook Pro with 128GB (Score:2)
Now that is promising for many people.
Re: M5 Max MacBook Pro with 128GB (Score:2)
I mean, if youâ(TM)re running AIs, youâ(TM)re gonna want 128GB at least, and a fast GPU. At that point, it being a MacBook Pro isnâ(TM)t going to make much difference. Itâ(TM)s gonna cost $7000, which I grant you is a lot of money, however⦠now find me another laptop with a GPU around the performance of the laptop 5080, and 128GB of RAM the GPU can access.
Re: (Score:3)
"...now find me another laptop with a GPU around the performance of the laptop 5080, and 128GB of RAM the GPU can access."
For the market that must run local AI as fast as possible on a portable device. Wonder what application needs that?
Re: (Score:2)
I mean, if you like, we can loosen the requirements and you can look for a desktop. Then the mac costs $5099. I'm not sure you're going to get a GPU with 128GB for that much, let alone the whole system.
Re: M5 Max MacBook Pro with 128GB (Score:3)
You can buy old GPUs with 32GB and put four of them in the box and do it cheaper. Since the bandwidth is more important than the processing this is doable.
Re: M5 Max MacBook Pro with 128GB (Score:2)
Yeh, except then youâ(TM)re paying a bandwidth penalty for accessing ram over PCIe or SLI, and a bandwidth penalty for having old GPUs, and a performance penalty for having GPUs that arenâ(TM)t as fast, or as flexible. So the result overall is not capable of doing this. When you use multiple GPUs to run AI, you load the whole model onto each GPU, you donâ(TM)t split it across GPUs, because the entire game is about having all the bandwidth available.
Re: (Score:3)
When you use multiple GPUs to run AI, you load the whole model onto each GPU, you donÃ(TM)t split it across GPUs
You absolutely can.
because the entire game is about having all the bandwidth available.
Between the GPU and the memory, yes. From your comments it looks like you have no idea how any of this works.
Re: (Score:3)
Well, since you seem to know everything about it, lets assume you do, and look at how much your proposed system would cost...
Lets see what GPUs have ever shipped with 32GB of RAM...
the 5090...
oh... that's it. So... absolutely no "old GPUs with 32GB of RAM" are available.
Well, lets assume you meant 24GB for the sake of it, what GPUs shipped with that?
The 4090, the 3090, the A5000, the 4500, and the 7900. We're going to need 6 of them to get past 128GB. Lets be kind and call it 5, because the mac is going
Re: (Score:2)
GPU based systems with that amount of memory cost about $30k
And: in case of Mac Minis, you can chain them via thunderbolt, I think up to 16 computers.
Re: (Score:2)
Application for a digital nomad, traveling around with his laptop and doing coding or AI related stuff.
Wow ... that was easy.
Re: (Score:2)
A laptop or desktop with 128GB of RAM will likely be around $7000, with RAM costs being what they are. Even a 64GB la
Re: (Score:3)
So, the trick is to use memory, search, fetch, vector database.
The purpose of a model is to reason. It needs enough training to perform further research. Context is very-short term memory, vector d
Re: (Score:2)
Okay, maybe, but... my MacBook Pro (M5 Max) can run Qwen 3.8 27b with 162k context window at 65 tokens per second. So your $2500 system is aproximately half the speed of a MacBook Pro. A MacStudio with 128GB of RAM and an M5 ultra costs $5000 - twice as much as your system, but would run it at somewhere around 120-130 tokens per second.
but promising what? (Score:2)
What is it promising?
Optimise all the things! (Score:2)
They built a custom engine for a specific model and it went 30% faster. That's great, but it's inflexible and of limited use.
Of course these days, coding agents make designing and writing new code very cheap. With the right tools and toolchains it shouldn't be hard to automate the process of building a custom engine for each new model that comes along. The cost of doing this would be small compared to the cost of training the model in the first place, so really we should just be doing this as a matter of co
Re: Optimise all the things! (Score:3)
I somewhat disagree. They developed an engine using somewhat different building blocks. The first iteration of their design works on only one model, but guess what - a lot of the blocks will transfer to other AIs. Iâ(TM)d bet that this can be expanded to run at least other versions of Qwen, and plausibly other LLMs.
Ultimately, all software is about trade offs, and here theyâ(TM)ve made a somewhat different trade from others, but given how popular high end Apple silicon devices are becoming for
Re: (Score:2)
"That's great, but it's inflexible and of limited use."
It won't be able to kill people nearly fast enough.
"...coding agents make designing and writing new code very cheap. "
Writing new code has always been "very cheap" as long as you don't care if it works.
"With the right tools and toolchains it shouldn't be hard to automate the process of building a custom engine for each new model that comes along."
Sure Sam, it's all easy. That's why no one has to work any more.
"The cost of doing this would be small comp