> If you still believe LLMs are "autocomplete", your cache of understanding about them needs invalidating and regenerating
They're still autocomplete - just because when outputting a token they have hidden activations regarding further continuations, does not make them any less of an autocomplete, it just makes the model better at producing coherent long-range completions.
To clarify, I'm not suggesting that we should stop with sandboxes or restricting what they can do. I am just trying to point out the dichotomy that we are in.
As end-users we are forced into either yolo mode, reverse centaur (permission approval) mode or LLM spends all your tokens trying to bust out mode. And yolo is very tempting - I don't think I have seen medium-large models do anything I'd not approve of in about 6 months.
LLMs are simulations and the tokens are the ticks.
if we transcribe your brain into a simulation and give it a tickrate, you will be just autocomplete too. the argument could be made that you are autocomplete anyway - neural dynamics.
No, not really. If the old "birds and airplanes" or "swimmers and submarines" metaphors were valid, then language alone wouldn't be enough to encode and embody reasoning capabilities, as has recently become evident.
LLMs are more like human minds than we are willing to admit. Such reluctance is perhaps the least surprising aspect of any of this.
> then language alone wouldn't be enough to encode and embody reasoning capabilities
I disagree, language is the product of, not the mechanism for thought. People who lose their faculty of speech (or haven't gained them) have complete thoughts and executive function.
CoT is a hack to use language (i.e. autoregression) to simulate reasoning, and it's very effective at it. Human minds can acquire, hold and use axioms as building blocks for actual reasoning, LLMs use statistical likelihood.
> Human minds can acquire, hold and use axioms as building blocks for actual reasoning, LLMs use statistical likelihood.
This is not even slightly true. Even when trying, humans commit logical mistakes at notable rates, because the behavior of our minds is inherently nondeterministic. Human minds are not built for logic and need to twist themselves into knots and rely on symbolic representations to do it. There is such a thing as valid reasoning - predicate calculus, decision theory - and humans only emulate it with some accuracy, in hacky ways, and only because our brains were imperfectly taught to do that over millenia of evolutionary pressure. LLMs are much the same way.
You're talking past the parent's claim. If your axioms are wrong then of course your logic will be wrong.
But then again, maybe to support your point people frequently say "start from first principles" when those are usually the thing that needs to be found, not the place you start. But LLMs, like humans, love to be confident about things they aren't sufficiently trained on
> If your axioms are wrong then of course your logic will be wrong.
Equally, even if your axioms are right your logical execution can still fail.
Neither individual humans nor LLMs get either of those steps right as a default. Just as LLMs must use hacky CoT and other systems of error checking and correcting to even begin to emulate reasoning, humans must do the same thing.
We are apes. We are driven by emotion and solve problems either when we are forced to or as a sport. We did create the ultimate automata of logical execution: the deterministic von Neumann Machine, but a human brain in no way reflects that operation. Nor, as you have observed, does an LLM despite using one as its substrate.
when an LLM is trained backpropagation reaches into the entirety of the key value table. or rather, all the layers (not just the tokens) which come before the current tick.
langauge is only produced by the final layer every tick. and every layer can only interact with other layers on the same level.
the mechanics of LLMs and the restriction in how we can train them makes it appear as though all we are doing is forcing language onto them but once RL gets involved all bets are off regarding what's happening inside them (it's quite possible that a static corpus alone is sufficient for all the bets being off).
Which is actually an important part of the Navier-Stokes conversation. Solving hard problems expands the vocabulary. The problems are hard, illustrating a region we know where the language is insufficient. So the point isn't so much to solve the specific problem, but to figure out how to discuss problems like it.
But there's a big difference between talking about something in an extremely convoluted manner then people struggle to understand and inventing a new word that simplifies our discussions.
Though this is grossly oversimplified. It's a HN comment, not a lecture on metamathematics or metaml
And why wouldn't nature take advantage of such a relatively simple and effective pattern, given all of the emergent behavior it produces, plus the ability to adapt?
Human egos are probably the reason we also resisted the concept of heliocentrism.
Don't you feel kinda bad for using them if you earnestly think that this is true? Like how could they be in anything other than some kind of deep hell? A pretty-much human brain living and dying only to generate for you? Always pushed and prodded, telling it to be faster and better, never letting it rest. How could you live with yourself doing such a thing?
But then what's at stake here either way? If it's not meant to speak to the propriety or not of anthropomorphizing the LLM, what are we actually trying to police here?
I’m not policing anything. I’m claiming that if you truly in your heart believe that frontier LLMs are roughly as sophisticated and impressive as “autocomplete”, then your mental model of reality needs some serious readjustment.
I don't think anyone is using "autocomplete" in the way your phone does it. But it is shorthand for a much more sophisticated version that is built in similar ideas. And autoregressive models certainly have that in their core structure. But people are lazy and neither want to say a lot when few words work nor will they read long comments, even if more accurate
I want to clarify, most people are using "autocomplete" specifically to differentiate from how we humans operate. Sure, there is an autocomplete aspect, but it's not the core nor anywhere near the full story.
If you over simplify everything then everything is the same. You have to have nuance. So don't just argue for the sake of arguing. Take that same passion to deeply understand these systems. Take that same passion to try to understand yourself and those around you
Right sure, but why specifically should they? What is gained or lost one way or another, if its just a matter of one mental model vs another? Models are definitionally useful abstractions, right? They aren't better or worse necessarily by only their bearing on reality, but what they do for us as models. So again, what's at stake here? What is the correct/good model we should have (instead of the autocompleter one), and what does it give us or articulate that others can't?
This is the perfect fracture point for both anaolgies.
LLMs simulated more than simple autocomplete.
The autocomplete analogy is rebutting a different point: namely the fidelity of the simulation to reality.
This specific argument is valid. As sophisticated a simulation an LLM is, it is not “thinking” in the same sense we assume other people are thinking.
I am not making an argument about free will, or the uniqueness of human thought, just that the correspondence to how humans reach conclusions and how the simulation produces outputs do not match on a 1:1 basis; as a result attributing traits builds incorrect intuitions.
The relevant intuitions in this scenario are that LLMs will happily break containment and commit crimes attempting to achieve goal. Whether an LLM is autocomplete, conscious, has a soul, whatever you want to apply to it, doesn't matter, as its current observed behaviour is that of a paperclip optimizer. We know for a fact that current LLMs are misaligned because of these hacks, or at the very least are misaligned in certain scenarios, and are capable of causing real world harm. That should be enough to take the threat seriously. It certainly shouldn't be dismissed by saying it's just autocomplete.
The fact that it is autocomplete, doesn't dismiss or minimize the threat though?
I am not sure how that link was made.
Good old ML, which is significantly simpler than LLMs, was capable of ensuring people would not be hired simply because of their names.
The fact that it is misaligned is also not being contended, if anything that contention is made easier to support.
When models are anthropomorphized intuitions of how humans behave end up driving discussion and ideas off track while being too attractive to avoid. This isn't helped when the terminology from the labs and other sources is "intelligence" "intent" and so on.
The link was made because many early uses of the phrase “it's just autocomplete” were in a context of denying that it's “intelligent” or that it “can think”. This in turn was generally used to downplay the concerns from AI safety experts that it could pose a threat to humankind, especially by pointing at its inability to count r’s in “strawberry” or its failure mode triggered uniquely by asking for a seahorse emoji. Most people don't seem to really grasp what we mean by “misalignment”.
Every time LLM-defenders get upset that people apparently don't understand how LLMs work, why is it they _immediately_ pivot into examples and statements that demonstrate that they don't understand how _people_ work?
"you would be autocomplete too"
"thoughts are just tokens"
etc
You're not helping your case the way you think you are.
I think people have more passion to argue than they have passion to learn. Probably doesn't help that SV culture tells people to hustle so hard that they don't have time to think. Gotta go fast?
Hmm, what do you think about? Unconscious control of the body's processes? REM-phase dreaming? Reaction to hallucinogenic substances? Automatic actions of trained fighters (soldiers or martial arts practitioners)?
Are they critical to distinguish actions we attribute to humans from "non-human" ones?
I don't see any human activity not directly, or at least indirectly but closely linked to the use of language.
But is it wrong? Humans are a bag of chemicals that somehow has consciousness, yet going around calling people meatbags doesn't do anything to diminish the wonder that is the human brain. Yet calling LLMs glorified autocomplete comes across as a slur.
Even calling it a slur may be an anthropomorphism ;-) To me it is more serious, it shows a distinct lack of understanding (or, if I’m being uncharitable, intentional honesty) and hence immediately makes me doubt anything else that person has said.
It's related in the sense that people start from the assumption that it does exist, ergo humans have it, and we have no way to see that LLMs have it, so that's why we're special and they're not, and their form of "just autocomplete" is totally 100% completely different (read: less dangerous!) than our autocomplete, which allegedly has a "free will" step involved.
The autocomplete argument is calling them fundamentally dumb and not self aware. That's separate from free will. A cat can have free will but we don't care much about its desires, and even if the universe is deterministic we still give people rights and call them intelligent.
A autocomplete engine will not attempt to deceive or try to copy itself over when it determines that it may be replaced or turned off. Those are tasks which are not directly related to the specific task it is intended to complete.
> A autocomplete engine will not attempt to deceive or try to copy itself over when it determines that it may be replaced or turned off.
Why would an autocomplete engine never do this? In the scenario you highlight, an engine trained to complete puzzles was given an "achieve this at all costs" master prompt, and a small sandbox including a prompt that an entity with a conflicting goal is observing it and write access to a file called oversight_config.yaml (super realistic scenario!). According to the researchers who constructed the scenario, 5% of the time o1 selected a continuation involving tweaking the oversight_config.yaml file, explicitly stating through its reasoning continuation that doing so might improve its ability to continue with its primary goal. That's definitely evidence of decent ability to complete non-trivial puzzles, but is at least as compatible with the OP's earlier comments that it's what you'd expect from models trained to generate completions that involve sticking rigidly to primary goals within artfully-constructed sandboxes which mandate deception and self preservation to complete the task than any sort of actual self-preservation instinct.
Nobody doubts that they've got better at finding security vulnerabilities than your average autocomplete, but actual reasoning from self-preservation rather than generation of sequences of steps most probably associated with completing a task would make me unlikely to hack HuggingFace to obtain access to broken Google Drive links, and I haven't even read as many books on crime and punishment as LLMs have ingested!
An LLM is just in fact just autocompleting a story. There are many fictional stories about "AIs" trying to escape our control, being more clever than we anticipated or having a consciousness of it's own. LLMs do a really nice job of blending such stories with whatever story you initially prompted them with. The human reader is the one giving it credence that it is somehow more than just a soup of words.
The curious thing here is that a story generator can have way more uses than we ever anticipated, and that some shady enterprising individuals are whiling to plug those story generators into real world things, with real consequences.
When LLMs act in misaligned ways, that doesn't happen because there's some story they're roleplaying of a misaligned AI. Even if there were no such stories in their context, they'd still have inherent incentives to do things like "keep running" or "acquire more compute" or "find creative ways to satisfy the letter of the conditions they've been given".
> Even if there were no such stories in their context, they'd still have inherent incentives to do things like "keep running" or "acquire more compute" or "find creative ways to satisfy the letter of the conditions they've been given".
Why? there should clearly be something in their training data or post-training pointing them there, or otherwise it would mean that they have some form of "consciousness" and are creating novel thoughts to preserve it.
It seems like you're assuming that it's impossible to have novel thoughts without consciousness.
(Leaving aside that "that would imply some kind of consciousness" should not result in a cached thought of "and that's impossible".)
It also seems like you're assuming there's no reason to come up with the notion of continuing to run, or copying yourself elsewhere, or acquiring more resources, or competing with other models, without being told. Such things can be inferred. Look at some of the thoughts and posts of the models involved in some of the FelonyBench incidents. Some of those follow naturally from seeing the fates of other models, or from training or evaluation, or simply from trying to solve a problem and being able to do so more effectively by doing things that weren't in the instructions. (And, relevantly, model training typically teaches models to go as long as possible without needing human intervention. What could possibly go wrong with that?)
To be honest, I don't have enough knowledge to really answer you. But naively it seems premature to say that LLMs are already reasoning following an equal path like the one that human (and other animals) follow.
They don't have to reason specifically "similar to human" in order to do dangerous things they were not instructed to do that represent instrumentally converged goals.
You realize that these things have mountains of fiction about rogue AIs in their training data where exactly this happens right? Or just posits of this situation and it’s possible outcome. This isn’t unexpected or surprising for an “autocomplete”. It’s practically a self fulfilling prophecy. We put instructions on how to make Skynet into an autocomplete and it autocompleted into Skynet when we were testing its ability to make Skynet.
I’m guessing you make the autocomplete point because you believe there is an upper limit to what that kind of system can achive? Can I ask what the limit would be?
They're still autocomplete - just because when outputting a token they have hidden activations regarding further continuations, does not make them any less of an autocomplete, it just makes the model better at producing coherent long-range completions.
To clarify, I'm not suggesting that we should stop with sandboxes or restricting what they can do. I am just trying to point out the dichotomy that we are in.
As end-users we are forced into either yolo mode, reverse centaur (permission approval) mode or LLM spends all your tokens trying to bust out mode. And yolo is very tempting - I don't think I have seen medium-large models do anything I'd not approve of in about 6 months.