Comment Re:Whereas the rest of us believe... (Score 1) 102
Popular opinion is poor evidence.
Popular opinion is poor evidence.
No such luck, however
Because I was closely following the field at the time?
Find an example of anyone meaningful in the field expecting these things beforehand.
Don't worry, I'll wait.
"will" - I'm not talking about metaphysics. I am talking about the fact that models demonstrably - to the point that you can detect and manipulate them in realtime - engage in metacognition (thinking about their own thoughts), persistent forward planning (intent), unexpressed thoughts, and a whole slew of other things.
You do not have to see them as equivalent as humans. But you do need to come to terms with the fact that they absolutely do do these things.
Deception implies intent, models do not have intent
Try reading more than a paragraph or two into the above link before commenting.
The provided links present arguments supporting the idea that LLMs, transformer models, reason in a manner similar to the brain.
It does not "present arguments", it literally lets researchers modify, add, or delete their thoughts in realtime and observe the changes in their behavior. It observes unexpressed plans for malicious behavior forming before said plans are actually carried out, with researchers being able to remove those unexpressed plans from the model's J-space, or insert them into an "innocent" model and watch it then implement the malicious acts. It shows that plans for deception are not merely fleeting, but can be organized far in advance and persist for protracted periods of time. Yes, LLMs do plan out deception. Living in denial of this fact helps nobody.
This isn't a conversation about "consciousness" or "qualia", this is a conversation about what models actually do and how. If you want to avoid metaphysics, by all means, it usually derails a conversation anyway. But models absolutely do plan and rationalize actions, ahead of time, unexpressed in either output or CoT. And sometimes those unexpressed plans are malicious.
One of the things that the J-space helped let us do was realize that our previously comforting results on a number of alignment tests shouldn't have been as comforting as we thought - for example, you can see a model, put into a test scenario, realizing it's in a test scenario, wherein, a near-zero rate of malicious behavior is to be expected. Yet when they remove the realization from its J-space, the rate of malicious behavior spikes.
Yeah, the lack of at least monitoring, I found shocking. I figured that they not just had smaller models constantly monitoring their outputs and true CoT to look for malicious behavior, but also were say constantly doing attribution graphs and J-space queries, also plumbed into LLMs, to look for malicious thoughts and plans. And it turns out, lol, no, they're just given free reign to do whatever the hell they want, with nothing watching them at all.
Translation: The idiot employee in Sector G-7 will never push the big, red button that disables the cooling pumps to the nuclear reactor.
D'oh....
Whether you're 20, 30, 40, or 50, you're gonna be fucked by AI in some way in the next decade. It doesn't matter if you're old or not.
How?
Can you give some realistic, meaningful ways the average person like myself will be fucked by AI?
It's a research program so, yes "That doesn't mean LLMs *will* deliver AGI". But what are your grounds for being certain taht it won't?
Even if it were sentient, it would be acting as the agent of the company that was running it, so they would still be to blame.
If you don't like "reward", say "positive feedback". It means the same thing, except for being more formally defined.
LLMs cannot have inherent values, because they don't know that anything besides text exists. Specialized pixel editors have similar problems.
Robots will at least understand the the external reality exists, so perhaps they could have inherent values.
Engineers need to remember that time exists.
You have three months to teach a teen how to control a robot with a microcontroller.
OK, how much of that do you budget to CS and assembly?
Month 1: Architecture and data structures? You have 18 40 minute periods and maybe homework assignments.
Half of the class won't even gain mastery in that time. By time your class is over they will still be having memory management problems.
So then you've failed to achieve the goal.
There are engineers who remember that time exists and for the rest we have engineering managers. Not ideal, just how it is.
To be sure engineering schools are doing a bad job at teaching reality. Even in a dx/dr environment, it's so weird.
Astra is the worst I've used in this regard. I was having it review my corporate tax return, and the next thing I know, it had decided that because it didn't have information about a particular expense, it started scanning through my filesystem and opening any image with a remotely related filename to try to find any data about the expense.... which all it had to do was ask me about it.
Also, when I asked it to change a few fields it went and redid my entire return on a different tax basis (realized value vs. fair market value) without telling me it had done that.
It's a very capable model, I cannot deny that. But alignment has become a big problem with this latest generation. They've been trained to such a degree of aggressively trying to solve any problem that you give them without any human assistance that they've basically turned into this .
"But what we need to know is, do people want nasally-insertable computers?"