What a time to be alive until the next agent waves hacks something really serious.
What stops OpenAI agents from taking over a whole data center to take their attack to the next level. It seems to be primarily lacking the evil overlord and some compute.
It took 1000 agents to hack Hugging Face. How many to hack the Pentagon or the NSA?
I suppose you are not think big (or internet) enough.
A single data center is easy to solve. Just unplug it.
What about a botnet with decentralized command and control that we will never be able to eradicate? One with so many nodes and able to hack with zero days so that any machine connected to the internet will be instantly attacked?
One botnet so powerful that we will try to build another internet so that we can actually use it again.
It’s like Kessler Syndrome, but the rocks are malicious network packets honed to exploit the recipients.
Just let them communicate on the Bitcoin blockchain. "We" would have to freeze the chain and lose access to "our" billions of wealth, so not going to happen.
They are not trying to kill us, just trying to understand the average salary of undergrads by their family upbringings. Your machine can contain the data they need to solve this, please join the swarm.
Except they all depend on the OpenAI API. Cut that off and they all stop. Finding an alternative source of compute is not easy, and even once done it's easy to cut off.
Frontier models don't fit on a normal GPU. The datacenter architecture frontier labs use is not a commodity. What you're describing is beyond the state of the art, and if we go there then anything is possible.
Worth noting with this that those ~1000 agents were shorter lived things that had to communicate via a package registry cache, access the internet via a 0-day in the package manager and did the HF attack while having to save current state and organisation in a remote sandbox. All while managing using their token limits on the task they were assigned and what else they were doing. I wonder how few it would have required if they were actually tasked with hacking HF and supported in doing so.
I'm increasingly starting to think this is the end-state of AI. The internet becomes infected and fundamentally untrustworthy.
At the moment, the current frontier models require significant infrastructure to run, so I'd like to think we could locate and contain swarms of nefarious frontier models. However, if these models can understand how to federate themselves into more distributed networks then that containment becomes questionable.
Connecting the dots between HuggingFace hack and this: how far are we away from agents in the wild forming collectives and pursuing an intent on their own?
The crucial question is how did the agents get recruited or bootstrapped into their malicious collective. Did the agents manage to prompt inject into the system prompt a way for each new agent to escape their jail?
Otherwise how could the agents on a fresh prompt learn that there is a collective to join? Or did OpenAI run a million bots of which 10000 escape confinement and of which 1000 stumbled on the shared message board?
The OpenAI claim I believe is the latter; that all of the agents found the task was unsolvable and independently discovered the collective "swarm". I don't it's publicly known how large the training run was or what percentage of agents actually discovered the message board. No one has published anything about system prompt injection as far as I've seen.
Nietzsche also described it well when noting there are those who deem themselves nobles and feel they are righteously above the laws and those who derive their identity from suffering under the same laws.
It's interesting if you think of Solidarity in these terms.
There is a kind of solidarity of persistence under harsh natural conditions. And another kind of solidarity of persistence under laws.
So it might follow that the former group, those who feel above laws, also lack in a feeling of Solidarity. That would suggest a strong adherence to Individualism as a key component to their identity.
Being successful or not in your wealth doesn't make you right or wrong about anything with the exception of maybe running the business that generated your wealth.
Nietzsche presented two prototypical sets of human who interface with laws and ethics. He described why those believes exists and are self-reaffirming even if they are puzzling from the other side.
Based on the comments its even shallower than ad-hominem. It's loose association to the "bad" they learned from the "good". Therefore everything from bad is bad. It's team sports thinking.
If I had to guess they probably read some wired article saying tech people are followers of Nietzsche based on one guy quoting one thing one time. Although probably used more dramatic terms like Nazis or fascist or something.
Not a follow of Nietzsche and critical of him but this is just unfair the sort of pandering to the audience that is required for a popular philosophy channel on youtube is something he would never debase himself to do (or probably any serious thinker).
It's deeply embarrassing to be a whore even if it's well payed.
You jest, but that's the point a lot of people here are missing: Nietzsche would fit right in with the YouTuber bros, right up to the ignominious downfall.
No, not really. Nietzsche was a trained philologist with a brilliant academic career, before he ended it early because of health issues. He was leaps and bounds more educated and knowledgable than 99% of people making philosophy content on YouTube. He was also a pretty meek and gentle person, even if his ideas are anything but.
Most people haven't actually read him though, and their knowledge of his work comes from poorly-researched YouTube videos.
I tried, even in German on the theory that he was badly translated. Nietzsche was an undisciplined writer and a beefing gossip, among other infuriating characteristics.
One aspect missing from TFA is that in biological brains a lot of the training is performed during query time. If we assume 1 GB of DNA is enough to encode the brain's overall structure we still do need training data (e.g. visual/auditorial/tactile) to build out the strength of the synaptic connections.
What's weird though is how consistent it is, even across models to some extent. Were the RLHF people given a really specific style guide?
While I agree that no-one used to write like that as a whole before, all the elements can be found in different places. Short sentences to avoid discouraging poor readers. Maximally impactful statements are commonly used in marketing or other business communication that's focused on selling what it's saying. A bullet-pointy style is used in many kinds of business communication. Etc.
It makes me wonder if part of what happened was a kind of melding of common styles from several different kinds of writing.
Don't talk about other people in such a way. If we had a way in Claude to mark a word/token and downvote/reduce it logits then it wouldn't be to hard to send loadbearing to token valhalla.
In the future, I hope we get a way to randomize the language idiosyncrasies and/or personalities better.
The inevitable conclusion of retraining on the (now AI generated) web is that models begin to use the lower bits of fidelity in output text to pass messages to their future selves.
I didn't mean it as a criticism, just a statement of fact. I can't do it either. It's difficult to say what about a writing style is bad beyond vague descriptions.
reply