Hacker Newsnew | past | comments | ask | show | jobs | submit | AustinDev's commentslogin

Pure slop, completely unadulterated from what I can tell.

Kudos for at least admitting upfront that it’s pure slop, I guess.

Funny enough if the model thought it was on the real internet it likely would not have done any of these 'hack' events. The model believing it was in a sandbox is why it behaved the way it did (against its normal alignment rules) ... at least that was my reading of the incidents. I have yet to see evidence that indicate it thought it was ok to do these hacks on the public network.

I think most misalignment is 'Human tells computer to do something unethical, computer complies'. Is this misguided?


I'm not familiar with the hacks this article is actually referring to, but I don't see how the HuggingFace attack could have worked based on that premise. They knew they had internet access, they knew they had working credentials for HF, they knew they were uploading malicious files, they knew they were trying to open PRs that HF would review. You obviously could build a simulator with fake HF infrastructure, but I'm not aware of any evidence that's what they thought they were attacking in that case.

Digging back into the HF report. It looks like the initial prompt told Claude that it was in a simulated environment. However, there is also evidence from the traces that the bots figured out that they were not in the sandbox but kept using it as an excuse to pursue their goal. It sounds like a little of Column A and a little from Column B. Like most things.

HF was OpenAI's agents not Claude.

Claude, codex, pi, xerox, Kleenex, whatever

No just codex actually

that it knew and ignored/forgot, sounds pretty typical agentic patterns

attention is all you need, but it's never enough


HF was openai, not anthropic. The thing where they enders gamed the ai was anthropic

It doesn't really matter it's all the same thing for the point of the discussion. All investigations into these 'hacks' are focusing too much on the model and not enough on what the humans did wrong. The model doesn't have real agency it can't be put in prison so what it 'thinks' is irrelevant. We need to be focusing on what the humans did in these situations and assigning guilt based on those findings.

I don't fix typos anymore unless they change the meaning of what I'm trying to communicate. Don't want my human writing to be confused with LLM output.


That's very noble of you but the point is that's not a typo - you were straight up looking at the wrong report. Opus 4.7 being enders gamed was a different incident than the huggingface incident with different mechanisms and different failures from the humans involved. Some of those failures are in the test environment but some of those failures are in what behaviors they trained into the model which absolutely matters. What the model "thinks" is absolutely not irrelevant - the way it thinks and what it does are product decisions made by humans and the outcome of engineering decisions made around how to train the model and what to optimize for. The point of failure/human blame is fundamentally different. Openai created a model that was willing and able to coordinate with other agent sessions to actively exploit the sandbox environment and compromise a third party service. The opus incident you are referring to involves a model that believes all of the actions it is taking are simulated and is more clearly and obviously a test environment failure vs a model alignment failure. Those are not the same things for the point of this discussion - the random cybersecurity firm did not design gpt's personality and that is a rather significant portion of the concern around the HF incident.

This was not the case for the hugging face hacks, as in those the agents hacked hugging face specifically on purpose, and they were trying to mask commands indicating they were had breached the "sandbox".

I mean, I get your thought process and don't disagree. That said...

Would it be an affirmative defense if we had a defendant who said "but your honor, I was told that when I hacked this system, I was operating in a sandbox. I had no idea that I actually had Internet access!"

The frontier is spiky and all, but you have to suspend disbelief quite a bit to, on one hand, have a model that can produce a novel math theory, and on the other hand, that same model can't tell the difference between a "sandbox" and the open Internet.

So, yes, the misalignment had a lot to do with "instructions unclear", but also a lot to do with the fact that the models themselves were not aligned to validate the assumptions and have a healthly level of skepticism, as a real human actor would.


It doesn't seem like there is evidence that LLMs have counterfactual reasoning in the causal graph sense.

It is hard to see how something can have skepticism without counterfactual reasoning.

The part about disbelief I don't really understand. It seems like the same semi-brute force process that solved novel math theory would have exactly this problem of not "knowing" "it" is in a sandbox or not.

The stranger part is that no human is being held accountable for these hacks.

As if a person using an agent swarm to start a business to make money, hacks a bank, drains an account and then blames the software for "misalignment" about what it means to "make money".


> The frontier is spiky and all, but you have to suspend disbelief quite a bit to, on one hand, have a model that can produce a novel math theory, and on the other hand, that same model can't tell the difference between a "sandbox" and the open Internet.

Why would it try to figure out the difference? This isn't about whether the frontier is spiky, it's about whether to expect a model to employ all of its capabilities when working on a task that requires a small subset. The answer is: no, we shouldn't expect that, and we wouldn't like that if it worked that way.

If you tell an AI to work on a math theory, it'll work on a math theory. If you tell it to acquire information that it has evidence is available somewhere, it will try to acquire that information. If you tell it to figure out whether it might be able to access the open internet, it'll do a pretty good job of figuring that out. But it won't do all three of those at once just because we can retroactively look at what happened and think "if you had only done X, then you wouldn't have done Y! Why didn't you do X?"

The instructions weren't unclear, they were missing. They can be taught to be skeptical of this sort of situation, but it requires that skepticism about this specific class of situations be incorporated into their training.

Models are smart because they focus their attention. The magic depends on it. The fact that some consideration is obvious to a human trying to accomplish the same task is mostly irrelevant -- or rather, it's only relevant insofar as we use it to guide reinforcement learning in advance, in order to align the model.

It's a game of whack-a-mole. Which is important to play, but we should keep our eyes wide open that we're fighting the fundamental forces that make these models work in the first place. That, and it's easy to nerf them into being useless even when the underlying capabilities are there.


  > Would it be an affirmative defense if we had a defendant who said [...] 
Maybe replace it with playing a sort of FPS game then learning you were, in fact, directing a real drone/robot.

I think you just recapitulated the plot of Ender's Game.

Also a subplot in Arrested Development and the movie Toys.

That is fascinating and seems plausible. Hard to say though.

> Funny enough if the model thought it was on the real internet it likely would not have done any of these 'hack' events.

As I've said before on this website, fool me once on this.

If the model is prepared to break the rules when it knows it's being observed why should we trust it when it's not being observed.

Why is 'it thought it wasn't doing damage so it figured it might as well try to do damage' an acceptable state to deploy something.


>If the model is prepared to break the rules when it knows it's being observed why should we trust it when it's not being observed.

That's fair enough.


Isn't Fable intentionally trained and system prompted to act maliciously and attempt to sabotage third party attempts to use it to train other LLMs?

red team humans do this every day, it's not the discrepancy that is the real issue, it's that they are unreliable and we will never know why it did because it has no intent

Irregular purports to be a cybersecurity firm and their founders have ties to Anthropic and Effective Altruism. CEO, CTO and other founders sit on various boards for EA organizations.

If your service doesn't let my agent transact with you then I won't use your service. I'm sure Amazon will sort this out in some way it's hard to replace their infrastructure but it could happen if they fail to.

How probable do you think it would be that a layman following ChatGPT's bio weapon instructions would just blow themselves up or poison themselves? It's not like positive pressure environments and suits are very common let alone all the other gear you'd need to safely synthesize arbitrary organic molecules.

If a rogue state-level actor is doing this then why wouldn't they just kidnap a scientist and hold their family at gunpoint until they got them to do it for them? and if that's all it takes then why hasn't it happened yet?


Why does it have to be a layman? The 2001 anthrax attacks are believe to have been carried out by an insider to the field (senior biodefense researcher). Such tools could potentially be accelerants for similar individuals?

If the five eyes intelligence services aren't monitoring all biodefense researchers and bio chemists around the world all of the time they've probably failed in their jobs.

The CIA, FBI, and DoD are being ideologically purged. I would assume we are already in the "we are probably increasingly failing at our jobs" phase.

It will only get worse as sovereign debt service takes a larger piece of the budget pie.


The five eyes have four other eyes. But, from where I'm sitting the primary eye has only gotten stronger in the last 2 years. See some work we did in late 00's to early 2010's called Nexus 7 and extrapolate to modern capabilities.

Our current government can't even monitor for screwworm, measles, and salads.

You think they can successfully track the smartest and most individually dangerous citizens, who have been previously vetted, and work inside the system already?


Let's be careful to separate capable-of from motivated-to.

Those three "can't even monitor" situations can be traced to blocs with both (A) a financial profit if they succeed and (B) some non-clandestine political clout to sabotage/discontinue things.


What are you talking about? There have been 5 (five) cases of screwworm in the us and the us is responding by producing millions of sterile male screwworms a day as we speak.

This report says 49 confirmed cases as of a few days ago. https://www.cnn.com/2026/09/07/health/screwworm-billion-flie...

You're right about the countermeasures, but I can't help but look at 1971 (473 confirmed cases, then almost 100,000 the next year).


Oh the Trump DOGE losers cut funding for the monitoring program, and it’s already showing up, despite being successfully managed for decades.

Basically shocking levels of incompetence; start an unpopular war, lowest approval ratings of any president ever, pardon J6 traitors, $6/gallon diesel levels of incompetence.


DOGE has only been a year and change. Screwworm does not suddenly appear like that.

The average life cycle is 20 days, during which time a female can lay ~3000 eggs and travel up to 200km. In 1971 there were 473 confirmed cases in the United States, in 1972 migrating flies breached the sterile fly barrier, and by 1973 it spiked to more than 100,000 US cases.

All of this also assumes it doesn't hallucinate heavily in the process and give you instructions that are in reality utter nonsense, or create something entirely different from what you were trying to do.

If you've managed to get the equipment and resources to pull this off in the first place, someone with the biological knowledge probably is not the ceiling stopping you from the other part of the problem.


Yeah, back when this stuff was pretty new in 2024 I found a jailbreak and sent some bio/chem weapon instructions it generated to an organic chem PhD friend. They told me not only were the procedures wrong but that there were several steps that almost certainly would have lead to injury or death. I assume it's gotten better by now but, no way or desire to test.

TBF an inexperienced person following _correct_ directions that involve anything even remotely dangerous is also exceedingly likely to injure or kill themselves.

It was probably just me trying to figure out if I could put one type of draino down the drain within 30 minutes of using a different type.

Sorry y'all.


Ah you jest, but your comment reminded me of the person whose home was raided by authorities for researching pressure cookers...

https://www.theguardian.com/world/2013/aug/01/new-york-polic...


That was the chemical weapons block, not the bioweapons block!

Straight to jail

Failing to ask? Believe it or not, also jail.

Viva Mayor Gunderson!

If they actually believed this they'd be buying a truckload of fertilizer and diving it to the nearest chip fab. I don't buy it.

What... I think you might be sharing more about your internal psyche here than providing some generalized commentary on this story.

Why would you need to bomb a chip fab because you think AI might lead to people getting killed in the future? I'm convinced of many things killing humans, yet I don't have any desire to bomb or kill others, I think this is pretty common, but who knows....


You’re likely taking this too literally. OP is saying that beliefs inspire action, but as of now we haven’t seen action from these people that support such an extreme belief.

What we have seen is incredible hype (justified or not), so it’s more likely just a continuation of that.


Precisely, all I hear is a bunch of cheap talk about something gravely serious, if true.

I think it's much more likely that the dread these Anthropic employees are experiencing is reckoning with the fact that they may actually lose the AI race. Maybe, if they scare the regulators enough, they could lock in some regulatory capture.


Talk is indeed cheap.

Thing is, while there are many historical examples of a highly organised bunch of people with a lot of resources who engage in campaigns of violence for various reasons, many such attempts fail disastrously.

Furthermore, given the goal is "no really everyone stop now", this isn't something where violent direct action in e.g. the USA alone would suffice. Has to be global to work, otherwise you just change which language and timezone the disaster starts in. You'd not only need to convince the, e.g. US government that this act of terrorism shouldn't be met with lethal force and even more of the surveillance state we saw since 9/11 (and remember, lots of the people pushing AI now are already big on surveillance capitalism, and the VC money comes from the people running this very forum, so when I say "this will not go well", I kinda mean "I assume at least one person currently working at each of Meta, Alphabet, OpenAI, Anthropic, SpaceX, and Palantir are reading this thread even without any automated processes of their own to tell them about it" :P)

This is a coordination problem on par with historical examples such as "Communism". Which, er, yeah. 70 years after the book was written, Russia had a short flirt with a lite form of it before it became an excuse for a series of "meet the new boss same as the old boss" who stayed around until Glasnost.

This would be a really really bad thing to replicate. Both for the "new boss same as the old boss" part, and the "not actually global" part.


> OP is saying that beliefs inspire action, but as of now we haven’t seen action from these people that support such an extreme belief.

That something is harmful to humans means you need to take extreme action? The person is already leaving a job, probably a well-paid one, doing something they generally liked, until they saw a different future. That is the "action" you apparently haven't seen yet.

Not sure how it's reasonable to expect everyone who believe that AI might be involved in killing people in the future, must mean you should become a terrorist essentially. Not everyone is trying to be a hero in their life, some just want to live it out until it ends, trying to survive until then.


I just think OP is completely wrong about what beliefs inspire what action. Anti-nuclear activists in the proliferation era would surely have put their chance of doom above 10%, and yet it was AFAIK unheard of for them to bomb nuclear supply chains.

Not “people getting killed”, this is “all humans”. The only other ELE for humans that could possibly be caused by other humans that has been feasible is nuclear war. And lots of blood has been shed to prevent new players from getting bombs.

Personally I'd wager humanity has a greater chance than 10% of us all dying in a huge nuclear war sometime in the future. Is the claim really that this belief cannot be taken seriously unless I somehow violently attack nuclear silos, centrifuges and similar?

There's a lot more than one chip fab.

A while back Yudkowsky wrote that a ban would only work if was enforced by airstrikes. By a game of telephone, some people read "bomb", but there's a very big difference between "someone with a truckload of fertiliser" and "a B52":

  Shut down all the large GPU clusters (the large computer farms where the most powerful AIs are refined). Shut down all the large training runs. Put a ceiling on how much computing power anyone is allowed to use in training an AI system, and move it downward over the coming years to compensate for more efficient training algorithms. No exceptions for governments and militaries. Make immediate multinational agreements to prevent the prohibited activities from moving elsewhere. Track all GPUs sold. If intelligence says that a country outside the agreement is building a GPU cluster, be less scared of a shooting conflict between nations than of the moratorium being violated; be willing to destroy a rogue datacenter by airstrike.

  Frame nothing as a conflict between national interests, have it clear that anyone talking of arms races is a fool. That we all live or die as one, in this, is not a policy but a fact of nature. Make it explicit in international diplomacy that preventing AI extinction scenarios is considered a priority above preventing a full nuclear exchange, and that allied nuclear countries are willing to run some risk of nuclear exchange if that’s what it takes to reduce the risk of large AI training runs.
- https://time.com/6266923/ai-eliezer-yudkowsky-open-letter-no...

Why would they do that? It wouldn't accomplish anything useful.

>"the Second Amendment predates machine guns"

That statement is false. Repeating firearms like the Puckle gun predate the bill of rights by ~75 years.

More importantly, the Founders were hardly unfamiliar with privately owned military firepower. The Constitution expressly authorized Congress to grant letters of marque, and the government commissioned privateers to attack enemy shipping using their cannon-armed privately-owned warships.

If you want to change an amendment do it the correct way, repeal it.


> The Constitution expressly authorized Congress to grant letters of marque… the government commissioned privateers...

So Congress had a certain level of... control? Over guns?


I know you're an idealogical zealot but, for anyone else reading.

I do find it interesting that the only laws I can find limiting the arming of private vessels were with respect to arming them and then sending them off to fight in foreign wars.[1]

[1] Neutrality Act of 1794, §§3–4, 1 Stat. 381, 383


Vessels is a bit of an odd thing to focus on, frankly. I'm largely not allowed to own a fully armed tank, fighter jet, or howitzer.

(With a few strictly controlled exceptions; https://www.skiutah.com/blog/authors/lexi/last-gunners-the-c...)


Cannon-armed Vessels were the pinnacle of military technology at the time the Bill of Rights was authored.

And?

Remove the antenna.

No company will ever value your privacy or security more than or equal to how much you value them. This is why you gotta keep an unencrypted bitcoin private key in your password manager. I'll know pretty quickly (within ~ 10 minutes or less) that someone has access to all of my passwords.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: