Forgot your password?
typodupeerror

Comment Re:What a load of feces (Score 1) 187

But it doesn't line up.

There is an old paper (by Bernard Barrs) "In the Theatre of Consciousness" that likens the mind full of ideas to a theater's stage, with consciousness being the spotlight that illuminates some ideas, leaving others dark. This theory of consciousness refers to itself as "Global Workspace Theory".

All Anthropic have done here is choose to call the middle layers of a Transformer a "Global Workspace" (may as well have called it a "Pink Wombat") and say "wow, look, we've got a workspace too, so it may be conscious".

The spotlight of Barr's theory is the bit corresponding to consciousness, so does a Transformer have anything that can be called a "Spotlight" to try to make some analogy here? No.

Necessarily a transformer has many language patterns/predictions internally that aren't in the output it generates, since it is trained to predict all possible continuations of the input and therefore predicts a bunch of alternate next tokens, from which the framework calling the model picks one AT RANDOM.

So Baar's workspace/theatre of ideas has consciousness shining a spotlight on one idea, and a Transformer has a freakin' random number generator picking one.

Do you consider that "lining up"?

It'd be like saying that a banana and a car are similar because you can call a banana skin a "car body", and ignore the fact that the banana has no engine and can't move.

Comment Re:"Reasoning" (Score 1) 187

Nah - completely wrong.

Find a model that does this, and ask it to SPELL strawberry instead, and I'll guarantee it will get it right.

The problems some models have (esp. if you don't allow them to think out loud) is COUNTING, not "mapping tokens to letters".

The reason they don't have a problem doing this mapping is because to them it's not mapping - it's prediction, which is the one thing they have been trained to do. They are predicting that the input "how do you spell strawberry?" will be followed by 'S', 'T' ,etc. It makes no difference to them that "strawberry" may be 6 tokens - it's just one sequence of tokens predicting another sequence of tokens. It is what they are trained to do.

Comment Re:AI Company says their AI is the bestest boy (Score 1) 187

> neural networks are definitely doing something

No shit, Sherlock !

And what language models have been trained to do is find language patterns that help predict next token.

Anthropic : Holy fuck! This thing is finding language patterns that help predict next token! It's alive! It's worth a bajillion dollars!

Comment Re:Well it isn't (Score 1) 187

Yes, although Baars' "Global Workspace Theory" (pertaining to the human brain and consciousness) is really irrelevant here. Anthropic only refer to it because they want, as usual, to make grandiose claims and want you to believe that LLMs may be conscious (while of course not being willing to commit to an actual definition of consciousness).

Of course Baar's theater (not ocean) metapor is central to his theory of consciousness which he compares to a spotlight in a theater drawing attention to some players on stage while ignoring others.

OK, so if you strip away Anthropic's poorly applied comparison to Baars' GWT, and appeal to you to believe that LLMs may be conscious, then what are we left with? What is the actual technical content of their blog post (which the Slashdot story doesn't even bother to link)?

https://transformer-circuits.p...

Their title "Verbalizable Representations Form a Global Workspace in Language Models" cuts to the heart of it.

The "Workspace" that Anthropic are referring to (fig 2) is just the the activations of the middle layers of a Transformer which is where most of the high level pattern recognition and prediction occurs (since the input/output layers need to convert word sequences to/from these richer middle layer embeddings).

OK, so why call it a GLOBAL workspace (which is central to their grandiose brain analogy)? A: Because the diagram of a Transformer you are likely familiar with puts all the focus on the stacked Transformer blocks with their attention and feed-forward components, and presents the residual connections as more of a technicality (maybe you think of these as just to help gradiant
propagation as in a ResNet). In fact the residual connections are really central to a Transformer, which may be better thought of as an embedding highway (the "residual" connections) passing thought the layers, with the job of each layer being to incrementally augment these embeddings by adding data to them (derived from attention and feed-forward blocks).

So, what makes these middle layer "workspace" activations (i.e. embeddings) global is that the residual highway interconnects all layers and anything added in lower layers will therefore be globally accessible to all layers above it.

So, there you have it: Surprise surprise (NOT) Transformers share embedding values across layers (woo hoo - global) which represent high level concepts, some of which (depending on random output token selection) may manifest in the output, and others just representing internal abstractions and output paths not taken.

Anthropics PR spin: brainz.. brains.. it's alive! it's conscious!

Comment Re:Lithography (Score 1) 28

I guess you could frame it like that, but knowing how it works doesn't help if you don't have the know-how to build it.

How do you increase your EUV power from 100W to 1000W (this took ASML years to figure out)?

How do you make mirrors smooth to within a single atom deviation?

How do you make chemical etch pure to the parts-per-trillion level?

No doubt the Chinese will figure these things out by themselves if they have to, but I'm sure there are a few secrets to be had that would speed that up!

Comment Re:Lithography (Score 1) 28

EUV would be nice to have, but at the end of the day it all comes down to cost. Without EUV your chips will be larger and slower, so you'll need more of them (more expensive) to build a cluster of the same power.

The cost of serving (not price to you) an AI model is mostly hardware depreciation cost, not operating cost, with the accelerator chip cost being a large part of that.

NVIDIA's H100 chip costs $25-40K

Huawei's comparable Ascend 950PR costs $7-16K

As you can see, lack of EUV is not stopping the Chinese from being price competitive.

Comment Re:Lithography (Score 1) 28

That's irrelevant. Sure China has been blocked from buying ASML EUV machines, but they are doing just fine with DUV and companies like SMIC (Chinese semicondustor fab cf TSMC) have pushed it to ~7nm node size.

The Chinese are making their own AI accelerators, such as Huawei's Ascend series, which DeepSeek are using, but just like OpenAI making their own chips to avoid the NVIDIA tax, DeepSeek now want to make their own presumably at least in part to avoid the Huawei tax (as well as perhaps to gain even greater efficiency).

Comment Re:Wait a minute (Score 1) 69

See the thing to remember is those people is that what they accuse the other side of doing, just blind ideology.

It's just a variation of the Goebbel's playbook, which the Trump administration loves to follow - "accuse the other side of the thing you yourself are guilty of".

- Try to rig the upcoming election while yelling loudly about how the other party consistenly cheats - and without evidence, of course.
- Make up stories about how crooked the Dems are, while actively grifting yourself.

Regardless, it's nice to see Congress occasionally showing signs of having a spine, finally. It'd be great if they'd also figure out that the revenge dismantlement of NCAR is also going to cost money and lives.

I'm not even sure if it's that deliberate, or it's just the fact that Trump is thinking about rigging the election... so he talks about rigging the election.

But it's hilarious how consistent the pattern is. Normally with something like that there's just a few occasional examples. But with Trump if he says "Democrats are kicking puppies!" chances are that we're about to find out that Trump kicked a puppy.

Comment Re:The SpaceX Valuation is Insane (Score 5, Insightful) 67

SpaceX is worth more than Microsoft or Amazon at this point. It boggles the mind how much people are betting on the future just because Musk is a genius. If he gets sick the stocks craters 80% easily and this $60B is more like $12B.

He's not a genius, I sincerely think he's average to slightly below average intelligence for a software dev. Just look how clueless he really is when he pretends to be a technical guru in front of actual experts.

That doesn't mean he doesn't have some exceptional skills, but IQ isn't one of them.

First, he's hard working, at least in spurts (during critical deadlines), and he's willing to make and implement big decisions quickly. Just look at DOGE, Republicans have been trying to lay waste to the US government for decades, but Musk is the only one to actually do it. It was a complete disaster, but it wasn't ethics or common sense that stopped the previous attempts, that's a legit talent for Musk.

Second, CEOs aren't allowed to lie, but Musk has figured out that you can get around that by building a cult of personality and then making ridiculously optimistic predictions and then sell minor advancements as progress. The result is he has a core group of retail investors that buy his stocks based on vibes and refuse to sell once in. Since these retail investors prevent the stock from going down too much institutional investors also jump in on the ride. It's basically tulip bulbs.

Comment Re:The speed of light (Score 1) 102

We don't understand dark matter, don't understand black holes due to not understanding physics in that realm (no unified theory), don't know how to interpret quantum theory. We know that entangled quantum particles act in synchrony over arbitrary distance without any signal between them being transmitted at all ... there's a lot we don't know (including not knowing the full extent of what it is that we don't know).

Even if did know it all, what if the thing traveling has a lifetime of millions or years, or in an AI, maybe traveling at near light speed?

Science simply is not in the business of saying what is impossible - it is in the business of predicting what happens in an experiment where we have an adequate theory. When the prediction is wrong you revise the theory.

Comment Re:For real or for the marketing? (Score 1) 102

Obviously he wouldn't know unless has had personally seen them.

One of the alien rumors is that Nixon wanted to impress his buddy Jacky Gleason and showed him some proof of aliens, and presumably if that did happen then the military would have learnt their lesson (as the public did) about the untrustworthiness of politicians and presidents, and kept them out of the loop in the future. I would not be surprised if the president is kept out of the loop on the most secretive black projects. You'd have to be an idiot to tell Trump something and expect him not to leak it.

Slashdot Top Deals

"The great question... which I have not been able to answer... is, `What does woman want?'" -- Sigmund Freud

Working...