Forgot your password?
typodupeerror
AI

Perplexity and Nvidia Launch Fully Local AI Agent With Zero Token Costs 90

Perplexity and Nvidia have launched "Portable Computer," a local-first version of Perplexity's agent platform that runs AI models, files, tools, and workflows directly on Nvidia-powered Linux hardware. Local tasks incur no token charges and keep data on-device by default, with users asked for permission before the system escalates a step to a cloud model. VentureBeat reports: For Nvidia, which has spent the past two years selling the world on trillion-dollar AI data centers, the announcement signals something subtler but strategically important: the chipmaker believes local AI has crossed a threshold from hobbyist curiosity to practical tool -- and it wants to sell the hardware that runs it.

"Local AI reached an inflection point," said Nader, Nvidia's director of developer technology, who focuses on developer tooling and open source. "For the longest time, it was hobbyists and enthusiasts, and they were running these quantized models that were quantized down to be super tiny... And while that's cool, it's not super practical. But all that changed with a lot of these new open source models that have come out that are super useful."

[...] Portable Computer arrives today for Pro, Max, Enterprise Pro, and Enterprise Max subscribers on Linux, with Windows support following in September. Any RTX GPU with at least 24GB of VRAM -- roughly a GeForce RTX 3090 or newer -- clears the bar, a threshold Nate called "sort of the floor where we really want to make sure that we can deliver a great experience, but balance that with making it broadly available."
This discussion has been archived. No new comments can be posted.

Perplexity and Nvidia Launch Fully Local AI Agent With Zero Token Costs

Comments Filter:
  • Citation needed for these trillion-dollar data centers (even just one).
    • Collectively all the AI data centers represent about a trillion in capital expenditure. Rather than the plural form meaning there are multiple data centers each other over a trillion dollars.

      (also, I think someone edited the trillion-dollar part out. I didn't see it in the summary. but I believe you that it was probably worded that way earlier today)

  • *looks at prices of that GeForce RTX 3090 with at least 24GB of VRAM*

    Nah, I'm good fam. Not selling a kidney to run a local AI model.

    • You're not the target audience. The companies getting tired of paying $50K for tokens are the target audience. Believe me, the hardware cost is trivial compared to how much the AI companies are raping businesses for.
      • by MobyDisk ( 75490 ) on Tuesday August 25, 2026 @03:58PM (#66306648) Homepage

        I suspect a company paying $50k/month for tokens is going to need a whole lotta RTX 3090s. I also suspect that this 24GB model is not as powerful or efficient as something they can get in the cloud. But with that said, this is basically what I do at home - run one locally for common tasks. I keep waiting for the local ones to achieve parity with the cloud versions, but RAM is so expensive because of the data centers, that I'm not sure it will happen before SkyNet launches the nukes.

        • Exactly. Especially as it seems that the AI companies are selling their tokens at below cost. Small-scale installations for local models may make sense for individuals who want to keep their data isolated, but they do not make economic sense for companies.
          • by flink ( 18449 )

            Token sold by AI companies are unprofitable because they are spending 100s of millions retraining their models every month to stay ahead of the other frontier models. They are net negative because they aren't currently fully recouping their training and infra costs. They are trying to boil the frog by gradually raising prices to cover their actual overhead + R&D before the VC/IPO money runs out.

            You can RUN a model for cheaper than the tokens from a frontier AI company would cost you, depending on your

      • This seems to still require a Perplexity Pro, Max, Enterprise Pro, and Enterprise Max subscription for it to work.

        • Which means that it's basically a non-event. Why deal with this vendor-locked garbage when it's trivial to set up the toolchain and buy one of those Nvidia boxes [amazon.com]? Or one of those AMD Ryzen AI Halo things?

          Yeah, it's still Nvidia, but if you're going to spend thousands on hardware, they've always been a good bet for still being around in a few years, and still giving a shit about you after a few years.

          They're still making Shield TV Pro software updates like 7 years later. Buying into some locked up piece o

      • We pay $20 a month per user just like a consumer plan, and allow overages that bill at a higher rate because it's still not that much. If you spend $50k a month you know what you're doing. Fuck you talking about like it's too expensive weirdo.

      • by allo ( 1728082 )

        It is not, but maybe it will be soon. The price of a 4090 will buy you many cloud tokens, and we didn't even talk about the electricity yet.
        The strength of local AI is not cost or efficiency, but privacy.

    • by Tailhook ( 98486 )

      An RTX 3090 24GB is about $2150 new.

      I've spent more on gaming rigs a couple times: it's not some unfathomable amount of money.

      • I've spent that much on an ENTIRE gaming rig. Back when the 3090's were new and "only" cost around $1400, I bought a 3070ti as part of my $2000-and-change rig. Now, a used 3090 runs around $17-1800. I'm seeing new ones at $2099, and the idea that 6-year-old hardware somehow costs more now than it did when released is mind-boggling. And it was released when crypto miners were buying up all the stock and had already driven the prices up a couple hundred bucks.
      • An RTX 3090 24GB is about $2150 new.

        I've spent more on gaming rigs a couple times: it's not some unfathomable amount of money.

        For a whole gaming rig? That's pretty typical. If you're spending that much for the video card alone? That's extreme. Yeah, I know many hobbies are expensive. But every expensive hobby I can think of (photography, cooking, arts/crafts, fashion)...if something costs $2000 more than an entry-level alternative, it lasts a lot longer than a video card. TBH, I was racking my brain thinking of worse uses of money. All I could come up with is auto racing.

        Can I afford it? Absolutely, I make good money

    • This a prelude for the RTX Spark due to be released in the fall. It is also how AI should have been done from the beginning if, you know, 'local' hardware was capable at the time.
  • 24GB VRAM ?

    Hear that noise ? That's rich prople laughing at us.
    • A used 3090 costs over $1700 now. USED!
      • Where were all the investment advisers ten years ago, telling us to *invest in* GPUs and RAM ?!

        We could have made millions !
        • If I had blown my credit limit on 5090's when they came out, I'd have one a hefty profit now. If I'd only known!

          Absolutely bonkers that computer hardware could go up in value years after it was released. That's never happened before.

  • And while that's cool, it's not super practical. But all that changed with a lot of these new open source models that have come out that are super useful.

    Dumb it down for me, nerd man. I have no idea what all this techno mumbo jumbo means.

    • by Fly Swatter ( 30498 ) on Tuesday August 25, 2026 @04:58PM (#66306796) Homepage
      AI inference is all math calculations, math numbers can be rounded off to fewer significant digits but you lose precision. The data within LLM models can be rounded down to less precise numbers, sort of like rounding the number 2.4368966734 to 2.44 requires less memory to contain. In this way big models can be made much smaller to fit within less memory but are less precise - the results become poor compared to the original big model.

      This process is called quantizing.
    • Its about RAM. The model needs to entirely in RAM, no swapping. Then GPUs and NPUs can do their magic efficiently.
  • 3090 or newer? (Score:4, Informative)

    by GrahamJ ( 241784 ) on Tuesday August 25, 2026 @05:51PM (#66306888)

    No not just newer, only xx90 cards have at least 24GB of VRAM and they all cost an arm and a leg. And you can already run models on them.

    • by vyvepe ( 809573 )
      You can get Intel or AMD VGA much cheaper (than Nvidia VGA with the same VRAM size) and run a local model. E.g. Radeon AI PRO R9700 32GB is about 1600; Intel Arc Pro B65 Creator 32GB is about 1200.
      I'm sure the Radeon can run llama.cpp pretty well (a friend uses it (about 29 t/s, Qwen3.6 27B, vulkan backend)), the Intel one works too (though I do not know anybody personally who uses it).
      • by GrahamJ ( 241784 )

        For sure. The article specifically states nvidia but I imagine the agent doesn't care what's running the model.

        The problem is the R9700 is a 9070XT with double the ram for $1k more. Even considering the rampocalypse $1k for 16 more gigs is Fing insane. That's double what even Apple charges for RAM! The other problem is 32GB isn't really enough. Sure you can fit the weights at Q8 but then you have to choose between a tiny context window or dumbing down the model with even smaller quants. Even with triple tha

        • by vyvepe ( 809573 )
          Yeah, 32 GiB is on the low side, but it is usable with small models. E.g. Qwen3.8-27B-Q6_K_M has KL divergence to F16 only 0.002 [unsloth.ai] and will allow for full training context. It will do for some applications.
          • I see the accuracy graph on HF now, that is indeed impressive for a q6!

          • hmm no MTP at that 23GB size though

            • by vyvepe ( 809573 )

              Well, MTP will take a bit of memory lowering maximum inference context. But not much, probably only about 30k tokens. I'm not sure. I know that a friend runs Qwen3.8 27B with MTP and full context 262k on Radeon AI PRO R9700 32GB. He gets speed about 29 t/s. He used some of the Q6 quantizations. Not sure which one. He is likely using Q8 for KV chace, not sure. KV cache precision should be at least the same (preferably higher) than the weights precision. He also switched to Q5 model (at least for a while) bec

  • Portable Computer arrives today for Pro, Max, Enterprise Pro, and Enterprise Max subscribers...

    Why TF should I need a subscription to run a local-only model?

    • by vyvepe ( 809573 )
      The subscription is a tax on stupidity :D
      The more clever ones just install llama.cpp (or something similar).
  • I'm curious to know what they're actually building. Isn't this just a local harness? A friend is currently working on one with some interesting differences from the main commercial ones available: https://darwin-finch.github.io... [github.io] He'll be doing a new release soon.
    • by allo ( 1728082 )

      Perplexity needs a search API. You may get away with fetching sites locally (even though some may require a headless browser) but these days you won't be able to scrape search engines. And if you ask "What is the dumbest Slashdot article" the first thing perplexity needs to do is a web search for Slashdot articles, before fetching them and rating them for dumbness.

  • by WaffleMonster ( 969671 ) on Tuesday August 25, 2026 @09:47PM (#66307220)

    There are thousands of open weight models available for free download from hugging face and inference software to run them is all open source and runs completely offline. Why would anyone use closed source malware with an insane "privacy" policy? I don't understand the point of this or what value perplexity thinks it is even offering.

  • Requires a top end card that will only fit in a tower case. More portable than a datacenter, I guess.

  • So amazing, so coincidental that this Nvidia announcement comes out on the day that Apple show a new 512Gb Mac Studio with killer AI support. Purely by chance. Of Course It Is.
  • by wakeboarder ( 2695839 ) on Wednesday August 26, 2026 @06:52AM (#66307552)

    Just an enhanced electric bill.

  • That means Nvidia has interest in selling good AI hardware to customers.

  • This is mostly a tool in search of a purpose.

    • Interesting, I don't see this type of comment much any more. These days, even most late adopters and outright deniers, have seen that LLMs can do valuable work and automate many tasks.

      As I write this, Claude Code is busy writing my code for me, at a rate of speed about 10x higher than I could do it by hand. And my "by hand" work is already significantly faster than most of my peers. Sure, it makes mistakes. But I can tell it what I want in a feature or enhancement, or what bug I'm facing, and turn it loose.

"Everyone is entitled to an *informed* opinion." -- Harlan Ellison

Working...