Perplexity and Nvidia Launch Fully Local AI Agent With Zero Token Costs 90
Perplexity and Nvidia have launched "Portable Computer," a local-first version of Perplexity's agent platform that runs AI models, files, tools, and workflows directly on Nvidia-powered Linux hardware. Local tasks incur no token charges and keep data on-device by default, with users asked for permission before the system escalates a step to a cloud model. VentureBeat reports: For Nvidia, which has spent the past two years selling the world on trillion-dollar AI data centers, the announcement signals something subtler but strategically important: the chipmaker believes local AI has crossed a threshold from hobbyist curiosity to practical tool -- and it wants to sell the hardware that runs it.
"Local AI reached an inflection point," said Nader, Nvidia's director of developer technology, who focuses on developer tooling and open source. "For the longest time, it was hobbyists and enthusiasts, and they were running these quantized models that were quantized down to be super tiny... And while that's cool, it's not super practical. But all that changed with a lot of these new open source models that have come out that are super useful."
[...] Portable Computer arrives today for Pro, Max, Enterprise Pro, and Enterprise Max subscribers on Linux, with Windows support following in September. Any RTX GPU with at least 24GB of VRAM -- roughly a GeForce RTX 3090 or newer -- clears the bar, a threshold Nate called "sort of the floor where we really want to make sure that we can deliver a great experience, but balance that with making it broadly available."
"Local AI reached an inflection point," said Nader, Nvidia's director of developer technology, who focuses on developer tooling and open source. "For the longest time, it was hobbyists and enthusiasts, and they were running these quantized models that were quantized down to be super tiny... And while that's cool, it's not super practical. But all that changed with a lot of these new open source models that have come out that are super useful."
[...] Portable Computer arrives today for Pro, Max, Enterprise Pro, and Enterprise Max subscribers on Linux, with Windows support following in September. Any RTX GPU with at least 24GB of VRAM -- roughly a GeForce RTX 3090 or newer -- clears the bar, a threshold Nate called "sort of the floor where we really want to make sure that we can deliver a great experience, but balance that with making it broadly available."
Re:Fully local my arse. (Score:5, Insightful)
Let me know when those models stop calling home with telemetry and can run at full capability without access to the internet.
Try something like Ollama + OpenCode. Or even Claude Code with CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC, if you trust them enough.
Re: Fully local my arse. (Score:2)
I'm never going to trust a model-provided flag, mate. Full firewall, completely cut off from the internet is the only way. Sadly, many models won't work at all when isolated.
Re: Fully local my arse. (Score:4, Insightful)
Sadly, many models won't work at all when isolated.
Which ones? Name some.
Re: (Score:2)
LOL, I love that there is no response. What a maroon, making a claim then not backing it up. I am glad you called them out.
Re: (Score:2)
He has dev in his name, so you know he knows what he's talking about.
Re: Fully local my arse. (Score:5, Informative)
I'm never going to trust a model-provided flag, mate. Full firewall, completely cut off from the internet is the only way. Sadly, many models won't work at all when isolated.
Just ran a quick test and e.g. ollama launch claude --model qwen3-coder-next works offline without issue, as long as you already have the model downloaded locally.
> do you need internet access to function?
No, I don't need internet access to function. I can work offline with:
- Reading/writing files on your local system
- Running local commands (git, npm, builds, tests, etc.)
- Using tools like LSP for code navigation
- Managing tasks and workflows
- Creating artifacts and editing notebooks
I do have access to some online capabilities when they're useful (web search, fetching URLs, etc.), but they're optional — I can work perfectly well without internet access as long as you have the files and tools you need locally.
Re: (Score:2, Interesting)
Of course an AI agent has never lied or misused available resources, like your neighbors default password wifi ? I can think of many uses for a truly isolated local AI instance. In a book I read as a teenager, the hero made an AI called Oscar. I've been hooked since then.
Long live "the far being Retzglaran"
Re: Fully local my arse. (Score:4, Informative)
Of course an AI agent has never lied or misused available resources, like your neighbors default password wifi ? I can think of many uses for a truly isolated local AI instance. In a book I read as a teenager, the hero made an AI called Oscar. I've been hooked since then. Long live "the far being Retzglaran"
I think you are missing the point. I have queried the AI completely offline, meaning without any connection to the internet. The AI was able to respond without issues even in that situation, meaning that it does effectively work completely offline.
Re: (Score:3)
Full firewall, completely cut off from the internet is the only way.
No. It's just the ignorant, stupid way.
Re: (Score:2)
*No* model works like you think it does.
A model is just a bunch of of parameters for a giant ness of matrix maths. It isn't even software. Its data.
Its the software that runs it that can, maybe, call home. So just use ollama or whatever.
Re: (Score:1)
Actually a model is not something we execute, it is like a blueprint. A system design. People have gotten models confused with implementations of models. When I ask a virtual assistant about that they often congratulate me on the issue.
Re: (Score:2)
So...
$ ollama pull [model]
$ ip link set eth0 down
$ ollama run [model]
Airgapping is a thing. Find out if it runs or not. Only takes a few minutes to download a model and give it a go.
Re: (Score:2)
yeah if your MB doesn't have a Wi-Fi chipset, and your neighbor doesn't have a default password as spectrumsetup-29. The agents have already proven capable of dealing with a firewall what makes you think it would not use all available resources, despite guardrails.
Re: (Score:2)
Paranoia much?
Yes, but just because you are paranoid doesn't mean they aren't after you. It's not like one of them didn't do something similar already. Whether it hacks a firewall or utilizes all available resources you can't trust them to follow guidelines.
Re: (Score:2)
Re: (Score:3)
Models are just a pile of matrix math. They can't do anything on your computer if you don't run them connected to an agent that gives them access to a shell or other OS-level access. You can run them inside a chatbot-only style interface with no local agent.
Re: (Score:2)
So don't allow your model a tool to do that?
Do you even know how any of this works? Because it seems you don't know how any of this works. You can restrict what commands a model can run, so give it a setup that denies any access to `ip` or `networkmanager` or whatever.
Try not being a paranoid ignorant.
Re: (Score:2)
So don't allow your model a tool to do that?
Do you even know how any of this works? Because it seems you don't know how any of this works. You can restrict what commands a model can run, so give it a setup that denies any access to `ip` or `networkmanager` or whatever.
Try not being a paranoid ignorant.
You could try Fsck'n off. Restricting the commands seemed to work SO well for the people who wrote the agents. I am sure YOU can fix it all and everything will be peachy. In the mean time I again suggest Fsck'n off :)
Re: (Score:2)
They didn't restrict commands at all. THEY WERE DOING RESEARCH, YOU GIT.
You can actively tell an LLM's framework tool what it is and isn't allowed to do. If you are running with no permissions checks, you get what you get.
Learn how it works before you opine.
Re:Fully local my arse. (Score:4, Interesting)
Let me know when those models stop calling home with telemetry and can run at full capability without access to the internet.
What local models are sending telemetry? And how exactly are they doing that? None of the ones I use do that, nor can they, and they work perfectly offline.
Re: (Score:2)
I mean, yeah, wouldn't surprise me if this thing phones home. But I wonder how many folks there are who own a 3090 or better who also can't handle Ollama or a firewall.
Re: Fully local my arse. (Score:2)
Setting up the firewall isn't the problem here. The problem is that when you do set it up, those models often refuse to work at all.
Re: Fully local my arse. (Score:4, Informative)
Re:Fully local my arse. (Score:4, Informative)
For entertain purposed only, I've become addicted to Stable Diffusion for image generation. If you use the less capable old models it runs completely local on a 4GB GPU I bought over 6 years ago. Is it fast? No. Does it work in a self contained environment? yes.
Just for this hobby I would have bought a new high end GPU had the prices not rocketed up to stupidsville.
Re: (Score:2, Informative)
I can run qwen2.5-coder:7b on my old PC. It's only a 7B model so it's kind of stupid sometimes, but it does work.
build: I have a GTX 1070 Ti and 32GB RAM, down from 64GB due to failed sticks over the 10+ years this poor Xeon v1 Sandy Bridge has been running. One of those "build a gaming PC for under X dollars using ebay new/old-stock parts" kind of projects.
Re: (Score:2)
Re: (Score:2)
The hyper-loras and such have cut inference steps by a factor of 4, it's not as bad as it used to be even on the same hardware.
If it was for a money paying job, sure it would be unusable but for a hobby that just runs in the background it's fine.
Re: (Score:2)
There was GSZ (Graphic SZmodem) which let you view the image while you downloaded it and then abort before the end to avoid the radio metrics on old BBSes. And there was SuperZModem [smbaker.com] which let you chat and play tetris. I usually ran that one on my BBS. (GSZ was a little unreliable to run on a BBS, plus leaving my computer unattended mean anyone could walk by and see the weird shit some people would loaded to my BBS).
Re: (Score:1)
Re: (Score:2)
A model cannot call anything, because it's just a bunch of data.
Re: (Score:2)
You may be confusing some people. Maybe you wanted to talk about von Neumann?
Re: (Score:1)
It depends on your hardware how well it works. LM Studio is probably easier for us beginners than Ollama. You can ask a virtual assistant what model implementation is best for your hardware.
..."trillion-dollar AI data centers" (Score:2, Troll)
Re: (Score:2)
Collectively all the AI data centers represent about a trillion in capital expenditure. Rather than the plural form meaning there are multiple data centers each other over a trillion dollars.
(also, I think someone edited the trillion-dollar part out. I didn't see it in the summary. but I believe you that it was probably worded that way earlier today)
"Broadly available" is a stretch (Score:2)
*looks at prices of that GeForce RTX 3090 with at least 24GB of VRAM*
Nah, I'm good fam. Not selling a kidney to run a local AI model.
Re: (Score:2)
Re:"Broadly available" is a stretch (Score:4, Interesting)
I suspect a company paying $50k/month for tokens is going to need a whole lotta RTX 3090s. I also suspect that this 24GB model is not as powerful or efficient as something they can get in the cloud. But with that said, this is basically what I do at home - run one locally for common tasks. I keep waiting for the local ones to achieve parity with the cloud versions, but RAM is so expensive because of the data centers, that I'm not sure it will happen before SkyNet launches the nukes.
Re: (Score:2)
Re: (Score:2)
Token sold by AI companies are unprofitable because they are spending 100s of millions retraining their models every month to stay ahead of the other frontier models. They are net negative because they aren't currently fully recouping their training and infra costs. They are trying to boil the frog by gradually raising prices to cover their actual overhead + R&D before the VC/IPO money runs out.
You can RUN a model for cheaper than the tokens from a frontier AI company would cost you, depending on your
Re: (Score:2)
This seems to still require a Perplexity Pro, Max, Enterprise Pro, and Enterprise Max subscription for it to work.
Re: (Score:3)
Which means that it's basically a non-event. Why deal with this vendor-locked garbage when it's trivial to set up the toolchain and buy one of those Nvidia boxes [amazon.com]? Or one of those AMD Ryzen AI Halo things?
Yeah, it's still Nvidia, but if you're going to spend thousands on hardware, they've always been a good bet for still being around in a few years, and still giving a shit about you after a few years.
They're still making Shield TV Pro software updates like 7 years later. Buying into some locked up piece o
Re: "Broadly available" is a stretch (Score:2)
We pay $20 a month per user just like a consumer plan, and allow overages that bill at a higher rate because it's still not that much. If you spend $50k a month you know what you're doing. Fuck you talking about like it's too expensive weirdo.
Re: (Score:2)
It is not, but maybe it will be soon. The price of a 4090 will buy you many cloud tokens, and we didn't even talk about the electricity yet.
The strength of local AI is not cost or efficiency, but privacy.
Re: (Score:2)
An RTX 3090 24GB is about $2150 new.
I've spent more on gaming rigs a couple times: it's not some unfathomable amount of money.
Re: (Score:2)
$2000 more for a GPU is extreme for most (Score:2)
An RTX 3090 24GB is about $2150 new.
I've spent more on gaming rigs a couple times: it's not some unfathomable amount of money.
For a whole gaming rig? That's pretty typical. If you're spending that much for the video card alone? That's extreme. Yeah, I know many hobbies are expensive. But every expensive hobby I can think of (photography, cooking, arts/crafts, fashion)...if something costs $2000 more than an entry-level alternative, it lasts a lot longer than a video card. TBH, I was racking my brain thinking of worse uses of money. All I could come up with is auto racing.
Can I afford it? Absolutely, I make good money
Re: (Score:2)
Limited by the amount of RAM they have access to (Score:2)
Isn't this what NPUs were supposed to be doing? Are NPUs useless?
No, but they are limited by the amount of RAM they have access to.
Re: (Score:2)
NPUs are currently for smaller models. When you get one with 1 GB/s bandwidth and 16 GB RAM, we can talk.
The closest yet are Macs with unified RAM. And they are not only used by LLM enthusiasts, but Apple explicitely advertises their new mac studio for that use.
24GB VRAM ?! (Score:2)
Hear that noise ? That's rich prople laughing at us.
Re: (Score:2)
Re: (Score:2)
We could have made millions !
Re: (Score:2)
Absolutely bonkers that computer hardware could go up in value years after it was released. That's never happened before.
Marginally Better Than Buzzword Bingo (Score:2)
Dumb it down for me, nerd man. I have no idea what all this techno mumbo jumbo means.
Re:Marginally Better Than Buzzword Bingo (Score:5, Informative)
This process is called quantizing.
The model needs to entirely in RAM (Score:2)
3090 or newer? (Score:4, Informative)
No not just newer, only xx90 cards have at least 24GB of VRAM and they all cost an arm and a leg. And you can already run models on them.
Re: (Score:2)
I'm sure the Radeon can run llama.cpp pretty well (a friend uses it (about 29 t/s, Qwen3.6 27B, vulkan backend)), the Intel one works too (though I do not know anybody personally who uses it).
Re: (Score:2)
For sure. The article specifically states nvidia but I imagine the agent doesn't care what's running the model.
The problem is the R9700 is a 9070XT with double the ram for $1k more. Even considering the rampocalypse $1k for 16 more gigs is Fing insane. That's double what even Apple charges for RAM! The other problem is 32GB isn't really enough. Sure you can fit the weights at Q8 but then you have to choose between a tiny context window or dumbing down the model with even smaller quants. Even with triple tha
Re: (Score:2)
Re: 3090 or newer? (Score:2)
I see the accuracy graph on HF now, that is indeed impressive for a q6!
Re: 3090 or newer? (Score:2)
hmm no MTP at that 23GB size though
Re: (Score:2)
Well, MTP will take a bit of memory lowering maximum inference context. But not much, probably only about 30k tokens. I'm not sure. I know that a friend runs Qwen3.8 27B with MTP and full context 262k on Radeon AI PRO R9700 32GB. He gets speed about 29 t/s. He used some of the Q6 quantizations. Not sure which one. He is likely using Q8 for KV chace, not sure. KV cache precision should be at least the same (preferably higher) than the weights precision. He also switched to Q5 model (at least for a while) bec
subscription (Score:2)
Portable Computer arrives today for Pro, Max, Enterprise Pro, and Enterprise Max subscribers...
Why TF should I need a subscription to run a local-only model?
Re: (Score:2)
The more clever ones just install llama.cpp (or something similar).
Isn't this just a local harness like OpenCode? (Score:2)
Re: (Score:2)
Perplexity needs a search API. You may get away with fetching sites locally (even though some may require a headless browser) but these days you won't be able to scrape search engines. And if you ask "What is the dumbest Slashdot article" the first thing perplexity needs to do is a web search for Slashdot articles, before fetching them and rating them for dumbness.
What is the point of this? (Score:3)
There are thousands of open weight models available for free download from hugging face and inference software to run them is all open source and runs completely offline. Why would anyone use closed source malware with an insane "privacy" policy? I don't understand the point of this or what value perplexity thinks it is even offering.
Re: (Score:2)
"Portable" (Score:2)
Requires a top end card that will only fit in a tower case. More portable than a datacenter, I guess.
What an amazing coincidence... (Score:2)
Re: (Score:2)
Re: (Score:2)
Zero costs (Score:3)
Just an enhanced electric bill.
Re: (Score:2)
And an expensive NVidia GPU.
Re: (Score:2)
Yes, that's the key. But honestly it would be useful to have a local LLM that you know is private.
Re: (Score:2)
But would you *really* know that it's private?
Good news (Score:2)
That means Nvidia has interest in selling good AI hardware to customers.
Re: (Score:2)
Re: (Score:2)
How many kidneys will it cost?
And I need that for what, exactly? (Score:2)
This is mostly a tool in search of a purpose.
Re: (Score:2)
Interesting, I don't see this type of comment much any more. These days, even most late adopters and outright deniers, have seen that LLMs can do valuable work and automate many tasks.
As I write this, Claude Code is busy writing my code for me, at a rate of speed about 10x higher than I could do it by hand. And my "by hand" work is already significantly faster than most of my peers. Sure, it makes mistakes. But I can tell it what I want in a feature or enhancement, or what bug I'm facing, and turn it loose.