Forgot your password?
typodupeerror

Comment Re: M5 Max MacBook Pro with 128GB (Score 2) 17

Qwen 3.8 27b runs with mtp and a 72k context window on two RTX 5060 Ti 16GB at about 35tps (in MTP, that compares to 75tps non mtp). And, yesterday, it nearly matched GLM 5.2 running on 8xH100 in every task I tossed it. (My glm rig is rate limited to 30rps). My comparison is "ability to handle long complex tasks".

So, the trick is to use memory, search, fetch, vector database.

The purpose of a model is to reason. It needs enough training to perform further research. Context is very-short term memory, vector databases are their long term memory, rag is the books they read, and search and fetch are their libraries.

I would LOVE to switch back to two 3090 cards for 48GB (they do image and video now), but when I run qwen 3.8 27b fp8 with 256k context, I find the quality is higher in some cases, but drops because the model then favors short term memory iver research.

Memory bandwidth is much more exciting. 500tps when running on an HPC is very very nice.

So, if it were me measuring, a single 64GB HBM3e GPU is really where we should aim as this should be consumer cost friendly by 2030. (now I have to try on an A100 later today)

Comment Re:Integral layer of the Trusted-Computing/DRM sta (Score 1) 34

In the long term, the goal is for ISPs to use NAC/TNC to interrogate your computer for Trusted Computing compliance, and deny you any internet access whatsoever if your machine isn't compliant.

You said home ISPs sought to deploy Trusted Network Connect about 20 years ago. It hasn't happened. What's holding it up? The rise of mobile devices, smart TVs, and other devices incompatible with the remediation means available in quarantine?

Comment Re:Doesn't matter, obsolete designs (Score 1) 159

Inference is neural network limited. On systems with entirely separated data and compute, this is a memory bus constraint. In quantized systems, ALUs and FPUs are illogical. A simple 16 cell LUT is suitable for noise free multiplication. A neural network circuit with a dedicated LUT and addressable buffers more or less eliminates bandwidth constraints. But this requires neural networks more similar to FPGA cells rather than classic compute.

That said, separating inference from data drops network depth considerably. And in transformers depth increases memory access exponentially which is why nearly identical MoE and dense models perform so differently. We of course need data centers for models with massive numbers of active parameters. But massive numbers of active parameters weaken models as that means depending on training rather than looking up facts. Huge active parameter sets will never be a good design. A real world analogy is that huge models are like Jeopardy champions who have crap loads of partial trivia facts in their heads. Smaller models are like librarians who don't have a trillions useless facts memorized but knows how to find the data for the researchers who will use the librarian repetitively to follow through to the answer. The model only needs enough training to use the data.

What really matters is, how fast it can find the data.

Inference suffers greatly when you treat weights as data. And bigger models have more trivia like facts memorized from their training. This means more active parameters and greater network depths with higher memory bandwidth needs. Even now, you see the chatbots are improving drastically because they rely substantially more on data and way less on inference.

I am not sure which part you consider gibberish, but it makes perfect sense to me. But it might sound better in my head.

P.S. I actually have considerable data to backup many of my claims. At work I'm sitting on about 50MW of compute. We build in shipping crates in a repurposed mine. There is a group of us where our goal is to avoid spinning up more compute. We need to deliver inference to 200,000 users eventually. (not customers, employees) and we could throw millions at NVidia and end up hosting some crap model, but we focus instead on cutting that memory bandwidth need. And yes, I'm far from being the smart guy on the team.

Comment How much tax? (Score 0) 165

This seems smart.

The federal government needs
  a) more tax money to reduce their dependence on more bonds
  b) more tax money to redistribute to startups
  c) reshoring to avoid just being too far behind.

So, taxing imports is a great idea. It doesn't matter if the trickle down effect works or not. The government has bills to pay. The government is owned by the people. Therefore, the people need to pay their bills. These taxes are good... we need to reward Trump for finally bringing new socialism to America. And what is best is that this really will steal from the rich and give to the poor.

BTW... Little secret, if TSMC spun down there business right now over the next 12 months, we'd be fine. We might have to rewind our tech a year or two and it might take some time to grow capacity, but it really wouldn't matter. TSMC just gives us a 12-18 month boost over their competitors. If they scale down as Intel, GF, Samsung, SMIC and a few others scale up... it really wouldn't matter.

Comment Doesn't matter, obsolete designs (Score 5, Interesting) 159

The absolute best data centers today should be written into history as a dark dark stain on the progression of humanity. It is a sign that no matter how stupid we have been up until now, we managed to find that hole has no bottom.

Consider this... The reason we need data centers that massive today is because the tech isn't ready.

NVidia B300 filled racks using 100kw each are barely capable of running modern frontier models. Even if we designed more powerful chips that could do the job, there isn't enough silicon wafers produced in the world to do the job. Even the most amazing semiconductor technologies on earth are not good enough. Consider that even if we could build 1TB HBM4 NPUs, the time it would take to generate a token on a transformer that large would be slow.

What we know

Transformers will leave the data center between 2030-2032 as it will never make sense to run transformer inference in racks. It's at least 1000x less green than local inference.

Intel, AMD, NVidia, Apple and every RISC-V/ARM chip supplier on earth is 100% focused on making local inference the thing.

We will soon stop training massive models that start becoming obsolete the moment the last parameter is fed. Even now, Qwen 3.8 27b is good enough for a year at least (first model to do this).

For LLMs, we'll just stop wasting cycles on huge networks and will focus more on stronger harnesses attached to better data sources.

Cloud AI will be more about acting as data providers. Models will become reasoning engines, the "smarts" will come from things like vector databases.

Vector databases will use kilowatts, not megawatts. Data storage will be much smaller.

Things like Intel's dead Optane product will become the future of datacenters. Google Search, Microsoft Bing, maybe Baidu are the best positioned companies for what comes next and they don't need 100KW per rack to deliver. If I were to enter this segment, I would pair a customized KLV object query engine onto a 1024 bit wide ecc memory bus and connect it with low cost networking (think Intel's Ethernet fabric tech). Then add a super capacitor to flush to flash on power loss... Or just hope Optane comes back. Superfast, massive, token based vector databases are the future of cloud LLM

What's next. In 3-5 years, every single rack in every single AI data center is trash

This is fact.

Even now H100 racks should be heading to scrap.

The chips are slow, they use old heat spreader tech, they get almost no performance per watt compared to B300. They are much harder to cool. They have so little VRAM it takes piles of them to run newer models. PCIe data center GPUs were always irresponsible purchases. SXM GPUs only run on special systems Noone but high end data centers can even run.

Consider that every single B300 or older equipped rack on earth is landfill in 3-5 years.

By then, the world will have overprovisioned RAM, we'll have moved 2-3 generations of cooling forward, we'll have designed denser racks, started employing all the awesome tech China built to work around lack of EUV... No really, Korean companies are already licensing CXMT tech for wafer stacking. We will have moved to lower fresh water depen....

You know what? It is going to just be cheaper to bankrupt the shell corporations who own the current data centers than to upgrade them to new tech. Besides, the environmental issues will be such a mess. Imagine trying to dispose of that much fr4 epoxy without causing environmental disasters? It's probably at least as much epoxy as the total mass of the twin towers.

I can keep ranting but I promise this, there isn't a single piece of tech in the most modern AI data center on earth that is worth using in 5 years... Not even the racks. It's all garbage.

Slashdot Top Deals

"Though a program be but three lines long, someday it will have to be maintained." -- The Tao of Programming

Working...