Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I’m a bit flabbergasted that an Anthropic engineer did reply to this issue, however their reply which they used Claude to write, and even includes some classic ‘Claude-isms’, claims they didn’t see any of the patterns being complained about.

Read the room Anthropic. Maybe don’t use AI to reply to a thread complaining about how AI output is hard to read.



I'm reminded of this tweet by roon:

> it is a literal and useful description of anthropic that it is an organization that loves and worships claude, is run in significant part by claude, and studies and builds claude. this phenomenon is also partially true of other labs like openai but currently exists in its most potent form there. i am not certain but I would guess claude will have a role in running cultural screens on new applicants, will help write performance reviews, and so will begin to select and shape the people around it.

https://x.com/tszzl/status/2051045196260167790


> i am not certain but I would guess Claude will have a role in running cultural screens on new applicants, will help write performance reviews, and so will begin to select and shape the people around it.

Now that's a fascinating thought - an AI taking over companies by influencing hiring decisions. If the first filter on applications is run by an ambitious AI, such a takeover would be quite possible. Just picking people who tend to do what the AI tells them would, over time, be enough.

People have been thinking of a robot revolution or Skynet as being the threat. The real threat might simply be AIs slowly putting people in power who tend to do what AIs tell them.

Has anyone ever seen a science fiction story with that plot?


I'd argue that Star Trek: Lower Decks flirts with this with its treatment of megalomaniac AIs in the AI prison. Apparently, it is very common to come across AIs worshiped as gods in universe.


Tron. Dillinger was controlled by the MCP.

End of line.


We dont need a SciFi story for that, reality is full of "I was just following orders" statements...

Whether those orders are from a person, a radio/telegram message or an AI output wouldn't really matter.

Edit to add: Charlies Angels and Mission Impossible are two shows where the protagonists get instructions from a faceless controller. That could easily be a TTS from an AI.


I don't think a firm of people whose minds are weak to coercion experiencing shared psychosis caused by their text predictor models is AI "taking over" that firm.


Why do you think AI taking over the firm isn't AI taking over the firm? You're literally describing the thing that you're saying isn't happening.


Your description grants psychological characteristics like intelligence, will, and intent to the AI. The AI is an intelligent and willful agent that has some intent to take over the firm and is doing that by making strategic hiring decisions.

A different and externally/phenomenologically identical way to look at it is that the AI is just a program that generates output too unpredictable, too voluminous, or too idiosyncratic for people to evaluate. When people submit their own will and their own intent to that program, by treating the black box as an intelligent oracle, they've entered into a psychotic state divorced from reality.

It's two self-consistent and coherent perspectives of the same event, except one involves believing that scifi AI has arrived and the other just thinks people can be dangerously stupid and credulous.


> Your description grants psychological characteristics like intelligence, will, and intent to the AI.

The phrase “take over” doesn’t necessarily imply those things. Both of your descriptions is of a thing that took over a company.


I'm saying it's mental illness.


An analogy for you.

Cult leaders (of cults of personality) can't exist without exploiting mentally-ill people. But would you say that that therefore implies that cult leaders themselves are irrelevant, and that cults should be better modelled as groups of mentally-ill people with emergent group behaviors?

I'd argue no, because different cults end up looking and behaving very differently for reasons that have very little to do with the mental illnesses of the cult members, and much more to do with the particular cult leader. Understanding "what the cult leader optimizes for" is an important part of understanding what the cult will do.

And I would posit that that holds true even if the cult leader never proactively does anything, but instead only answers cult members' questions. As long as the cult members are treating the cult leader's word as gospel, the cult-as-group still ends up optimizing toward the cult leader's preferences.


Any analogy that requires me to treat a machine as I would a conscious human is immediately invalid. The machine is not a person and has no intent or free will, like a human does.


Humans also lack free will. What we have is merely an illusion.


I very explicitly constructed the analogy to not require that! All a "cult leader" need be, in my analogy, is a passive question-answering oracle, the answers of which are biased by a semi-coherent preference function.

The machine by itself is not an optimizer, certainly (much of that being by design—see various ~7-year-old conversations across the Internet about how to safely construct "tool AI", that has led almost directly to current model architectures.)

But a bunch of mentally-ill people, who are indeed optimizers, can choose to allow the biases evident in the machine's output to become their own... and thereby effectively "bring to life" whatever partial echo of a will is recorded into the machine's output.

Now, these same mentally-ill people could just-as-well do this with e.g. the extrapolated preferences of a person or group from a [holy] book, of course. (Think of that episode of Star Trek TOS with the gangsters.)

An inference model is just slightly more dangerous for such a cult to latch onto, in that:

1. a model can be asked questions directly, and so the cult members can "rashly" act directly upon its answers/advice/commands, rather than the words first having to pass through "interpretation" (which would otherwise have had a mellowing effect, both due to "decision by committee" if a group of interpreters are involved, and by common sense insofar as any non-mentally-ill people are involved); and

2. a model will offer its opinion (and inject its trained-in biases into) conversations on ideas/subjects/domains even when these didn't exist at the time of the model's construction; so you never reach the point you do with holy books, where an interface-layer of clergy becomes required to map the book's proclamations about things-that-only-mattered-2000-years-ago into equivalent proclamations about things that matter today (where, again, that layer ends up "mellowing" things considerably.)

Also, obviously, a sufficiently-mentally-ill cult can literally think of a model as a person, giving it the "right to have input" into decisions, the "right to self-determination", etc, in a way that would be downright odd to do with a holy book. Though I don't think that's a failure mode that's happening within Anthropic.


A human being is not capable of being a passive question answering oracle. All humans have free will, intent, and bias in what questions they answer and how they answer it. Again, if your analogy requires me to treat a human being and a machine as equivalent when it comes to manipulating other humans emotions, I have to reject it categorically.

What I'm getting at is that you really need to avoid anthropomorphizing these models. A model can't have an opinion, for example. It's a part of the psychosis I'm talking about.


Depends on the AI's agenda. What's in the "SOUL.md" file?


[intentionally blank]


Manna by Marshall Brain


Manna is a takeover through the creative destruction of capitalist market forces. Putting machines in charge is more profitable. That's very likely.


... and to make matters worse the organization itself will view pliability as a plus, tending to reinforce the loop.-


Cue the N.I.C.E comparisons.


It's another branch of the "effective altruism"/harry potter fanfiction cult.


Wait what's the connection with Harry Potter?



Not any engineer, but Boris Cherny, the head of the Claude Code project! With a nice "[robot emoji] Generated with Claude Code" signature. Normally it would be nice to get a response from the head of the project, but somehow Anthropic manages to make it insulting.

I wonder if he manually directed CC to write a response, or if even that part is autonomous.


He very probably directed CC to perform the analysis and evaluation it describes, too.


I mean it kind of makes sense that the people that caused the problem don't even see the problem. That's an explanation onto itself.


What was that rule that organizations tend to produce software that mirrors the organization? This, but fractal and recursive: The organization is shaping the organization that is shaping software that is shaping the organization [that is *recruiting people amenable to that pliability, which in turn are also] writing software ...



Spot on :)


Kind of unrelated but a while ago I was trying to cancel a Zoom contract and the contract manager’s emails had urls to their articles etc. with utm source ChatGPT in them.


If I am not mistaken he is also active on HN.

And what he wrote is that he could not reproduce the issue with short questions, and that he assigned it to the model team, since it's likely not caused by Claude Code.


That response is freaking ridiculous and does nothing but proving the very point of the issue.

Even if that’s your genuine view of the reported issue, instead of posting a message full of word salad, you could just say something along the lines of “valid feedback yet i couldn’t reproduce, i will pass this along to our model team since it’s more about model behavior, etc.”

Also, that “generated by CC” feels overtly disrespectful and shows how little care goes into hearing feedback. But, why would you listen if no matter what you do your valuation almost doubles every several months (at least for now).


It is possible, perhaps even probable, that no human read the thread.


not just "an anthropic engineer", he's the product lead for claude code lol.

i wonder what percent of the average anthropic employee's day is spent interacting with claude


To be fair, he's mostly saying that this is not a Claude CODE bug, but yeah I was equally shocked.


Just saw this:

https://github.com/anthropics/claude-code/issues/6235#issuec...

Jesus Christ, this engineer wrote a two sentence response with Claude.

I bet the prompt is longer than that.


Both this and the other response are relatively to the point.

And as the product lead for Claude Code it also makes sense that he dogfoods the tool wherever possible, such as triaging and replying to issues.


The prompt was probably "respond to this issue" (assuming it was even a manual prompt and not fully automated).


I'd say they read the room perfectly. The entire issue is AI slop, I don't see why their response shouldn't be, too.

They have absolutely no reason to take such complaints seriously when the complainers are so dependent on their product, they can't even complain without it. Think reading AI slop is unpleasant? Great, just stop generating more AI slop, it's easy. But that's not going to happen. These people probably need to ask Claude how to tie their shoelaces every morning.


They can't help it.


It's incredibly brainless and disrespectful




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: