Hacker Newsnew | past | comments | ask | show | jobs | submit | kadoban's commentslogin

Oh, wow, they think it's just a smidge below the q4? That's crazy good if true.

The benchmarks they chose are rather cherry picked to not include long context or difficult ones that involve long horizon work or many agent turns, as I suspect this is where the model shows more differences compared to the full fat one

Yep more hops from the lower Q is likely going to skew the vectors further over time.

I wonder if there's a way to mitigate this by running it through an original Q8 draft model, attuned somehow for the PTQ1 quant, but giving it a higher threshold for the acceptance linear with the context length itself?

The longer the context, the higher the multiplier on the threshold, and more likely the draft result is used. Not ideal but it may extend the usable max context.

This model might, even without this, be amazing for short lived agents that work via generations / have changing tasks.


Honestly looks pretty good except whatever is going on with its booty. Is that an ass helmet? I cannot parse what's going on there.

Like you've never worn an ass helmet

Only because I hadn't previously thought of it xD Step up from the standard ass-hat for sure.

I think it’s supposed to be a wing

I like the lens effect behind the rear tire.

My problem with that is the answer of if I'm too busy depends on what the request/question is. Does it need a huge context switch? Is it something I like working on? Etc.

You can run the ~4 bit quant(s) on 24gb, if you're not _too_ picky on context size.

This will hopefully be better, though it'd be a _very_ surprising increase in performace at the size they say. Would love to see more about how it benchmarks.


I run Unsloth's UD-Q4_K_S on 20 GB of VRAM (RX 7900 XT) and I get ~90k tokens of context without quantizing KV cache. With 8-bit quantization, I get about a 134k token context window. That's with only one slot, but for me, it works pretty darn well, with 20-35 tok/s depending on how full that window is.

7900 XT is a sleeper card. When I initially bought it, it was priced at the lowest wattage per $ per GB VRAM (not normalized for token speeds...) Although I ended up swapping for the XTX because that 4GB means everything in just increasing the context window. At 8bit KV my window is over 200k, and although qwen3.8 loves vomiting out tokens as part of its reasoning chain I trust it enough to get assigned tasks done eventually, which I could not say of any model before its release.

How has software/driver support been? I got burned hard by AMD last generation or the one before. Things smoother now, or do you have to baby it like hell and pick and choose software that works?

I don't do anything fancier than inference, and I only use llama.cpp, which supports rOCM. I've had few issues; most GGUFs I download work right out of the box. Nearly any popular model has a quant that just works. But as you can see I don't use my GPU for anything weird or nonstandard.

You can trivially run 131k on 24GB 4bit, and there are repos with tweaks that allow you to get the full 262k but idk if there's degradation with their approach.

I'm doing the same with a context of about 128-150k Surprisingly, I get subjectively better results with Unsloth's 3 bit quants (UD-Q3-XL something), than their 4 bit quants (S or M)

> I don't think acid deserves such a scary reputation.

Acid in general can be handled quite safely with the right mindset and equipment.

HF is _fucking_ terrifying though. Even many professional chemists (or maybe especially pros?) won't work with it.

It's far beyond just being dangerous to your skin and eyes. Bicarb and goggles sure won't cut it as safety procedures.


Inflation is high, so interest rates need to go up to try to slow that, but the economy isn't doing amazing already, and higher interest rates won't help that.

Not to mention the US debt is _high_ as hell and bond yields mean that's more expensive.

And the country is run by a broken fool who has no interest or ability to fix any of that.


The country has been _run_ by fools for 26 years. Congress has had 26 years to do something about the fiscal situation, and we've had four presidents, and the fiscal responsible side of the electorate is never listened to.

Both sides are to blame - neither will fix the problem. Obama could've made that his goal - he was competent, had a lot of political good will, and many people were frustrated at the bailout policy Bush did, but instead it was inflationary printing (quantitative easing), Obamacare and Cash 4 Clunkers (which the used car market still hasn't recovered from).

I never voted for him - I didn't view him as honest, nor did he seem to indicate that he liked America, but was rather just a good talker - but I think he could've been a great president given a less radicalizing agenda.

He was probably the best situated president in terms of timing to fix the debt problem, but instead it was a good time for divisive politics. By the time Obama finished, it became clear neither party actually cared about the fiscally conservative Ron Paul supporting voting block.


> Both sides are to blame

— the one side after they fsck the country sideways, every goddamn time.


I would love to hear what was "radical" or "divisive" about Obama's policy. A significant portion of the country disliking him because of his skin color doesn't make his policies "radical"

A universal health care mandate were both radical and divisive, and the popular nickname for the ACA today is "Obamacare".

I happen to think the policy was a good idea, and voting to keep it in play was the best vote of John McCain's career ... but it was definitely both radical and divisive.

Now, much of the "mandate" has been stripped away, health care remains a mess, and access is far from affordable, but you can't really blame that one on Obama.


> A universal health care mandate were both radical and divisive

It seems quite ironic, given the frequent complaints about the inability of Congress to either govern effectively or fix health insurance (for many and various definitions of "fix"), that the ACA was so divisive. At least it got passed! Yet given the opportunity twice (2017-2019, 2025-2027), a politically viable alternative hasn't been offered up by opponents of the ACA, let alone being able to fully repeal it.


Lol cool so health care is your definition of "radical" and "decisive". Hint: it's neither of these things and the only reason it was considered as such was because Obama was black (see tan suit). Seems like we need a few more years of woke cause you still can't see the obvious

It was not radical and he adopted originally republican ideas. They even originally backed it.

The thing is, the divisiveness comes from one side - conservatives who were and are determine to oppose literally everything.


> The country has been _run_ by fools for 26 years.

It's hard to accept the 'everyone is the same' in light of Trump's antics. Remind me again which presidents started wars of choice at the behest of Israel even when they were explicitly warned of the consequences?


Why specifically 26 years? I agree that Congress has been increasingly useless, leading to more and more rule by presidential decree in order to have a government that runs at all, but there wasn't a step function 26 years ago.

its because prior to that (2000 Bush era), congress and president had a plan to payoff debt and had a balanced budget plan in place to avoid over spending.

Long term bond yields are not directly tied to the Fed funds rate.

The problem is the debt purchased by the Fed during QE had extremely low yields (COVID era) the reserves held by banks created by the Fed during QE now cost more to service by the Fed.


It’s worse than no ability to fix it — he caused a large part of it for unclear reasons

> he caused a large part of it for unclear reasons

Technically it was Besset, but Trump gave him the reigns.

The purpose seems to be to radically debase the dollar and setup a crises that requires an entitlement cut (Social Security) for "the good of the economy" while also expanding military spending at the same time.


Covid? Half the money printed happened under his watch the first admin. Biden continued the other half. Now we have yet another war to make matters worse. What are you proposing be done to fix it?

Well I sure wouldn't have started another war.

QE without public debt sterilization is going to appear as the costliest macroeconomic mistake of the early 21st century.

Disagree, fairly strongly. In 2008, four trillion dollars evaporated. In order to keep the economy from completely crashing, the Fed created $4T using QE and such tricks. The result was 15 years of flat. No inflation for 15 years. If inflation shows up a decade and a half later, that probably wasn't the fault of how QE was done.

You're getting my point wrong:

- I absolutely agree that inflation has nothing at all to do with QE, people who claimed that are just idiots who have a gold fetish.

- the problem I'm talking about is the fact that central banks didn't use QE as an opportunity to erase the public debt it bought. At the time it wouldn't have been an issue in any way. But now because inflation is back (due to oil) central banks cannot buy government bonds when they reach maturity and have to raise rates. Then the government bonds have become very expensive, and it has to be paid to the private sector on the market, so whenever a US govt security reaches maturity, the budget constraint increases. Sterilization of the debt would have alleviated this issue a lot at no cost.

Also, we should have taken the lessons of the era and raise the inflation target to 4%[1] at that time (it was definitely politically achievable then, now not so much).

[1]: https://www.imf.org/en/publications/wp/issues/2016/12/31/the...


I see. Well, what you propose might have the minor problem of being illegal for the Fed to do. I'm not perfectly sure (and I don't want to take the time to research this rabbit hole right now), but the Fed is deliberately different from the Treasury. It's not supposed to fund the government by creating money.

At a minimum, doing so would have created doubts about the future of the dollar. (Because countries that start having the central bank create money to fund the government often wind up in runaway inflation, with the currency becoming worthless.)

As to a 4% target: Given that they were stuck at 0% for the next decade (and tried, and failed, to get up to 2%), why would they move the target to 4%? They already couldn't do what they said, why double their failure?


> see. Well, what you propose might have the minor problem of being illegal for the Fed to do. I'm not perfectly sure (and I don't want to take the time to research this rabbit hole right now), but the Fed is deliberately different from the Treasury. It's not supposed to fund the government by creating money.

That's a good point, but it's not as clear cut. The Fed is supposed to achieve the double goal of full employment and price stability and it's not forbidden to make money out of thin air for that purpose, that's the reason why QE is a thing at all. The exact legality of canceling US debt on its balance sheet isn't clear, but:

1. Before 2014 the Obama admin had the power to pass a law making that explicitly legal.

2. There are examples of theoretically valid instruments to achieve the same goal which have been discussed in the period (see the “1 trillion dollar coin”).

> At a minimum, doing so would have created doubts about the future of the dollar. (Because countries that start having the central bank create money to fund the government often wind up in runaway inflation, with the currency becoming worthless.)

Context matters: doing it now would send a disastrous signal, but back in the early 2010s the challenge was to drive inflation up, which is why the Fed used QE in the first place. If anything such a move could have made QE more efficient to achieve its goal (in addition to helping today's public finances, which I argue would have had a stabilizing effect over the long run).

> As to a 4% target: Given that they were stuck at 0% for the next decade (and tried, and failed, to get up to 2%), why would they move the target to 4%? They already couldn't do what they said, why double their failure?

The IMF paper I linked above is pretty clear about the goal of such a measure, but the idea is to have more leeway in case of crisis, because if your inflation is around 2%, your Fed target rate is around 2% as well and you can only lower it by 2% as a stimulus measure, whereas with a 4% baseline inflation rate you have twice the leverage in terms of target rate.


Re 4%: And, as we saw in 2008 and after, being stuck against 0% with no room to move is a really uncomfortable place to be.

Re the "1 trillion dollar coin": I like your wording: "Theoretically valid". I don't like YOLOing theoretically valid moves in a crisis, only to find out a month later that the courts rule them invalid and you have to unwind them.


> I don't like YOLOing theoretically valid moves […], only to find out a month later that the courts rule them invalid and you have to unwind them.

I agree, which is why I put “go through the legislative process to make that legal” above. Especially since there was no real emergency. (“in a crisis” though going YOLO may still be worth it though, because it may be enough to earn the time you need to go through the bottom of the crisis. And also if it's very unclear how legal/illegal this is, the fait accompli may be enough to convince the judges to side with your decision in order to put the country in too much of a trouble).


Yay stagflation!

> And the country is run by a broken fool who has no interest or ability to fix any of that.

Trump will be gone in three years, but you'll still have an electorate that wants more free stuff while also getting tax cuts. There is zero appetite for fiscal reform in the U.S. The geometric growth rate of U.S. debt has been consistent since 2010 and will remain so when AOC is President: https://usafacts.org/answers/how-much-debt-does-the-us-have/...


You really really just need to raise taxes. Just find a way to sell that to the public (focus on the rich or large corporations or whatever outgroup you want basically)

Our budget deficit is $2 trillion. To close it, you need to significantly raise taxes on the fattest part of the income curve, which is the top 25%. They have $10 trillion of income. https://taxfoundation.org/data/all/federal/latest-federal-in.... An across the board 200 basis point increase would close the deficit. That would raise their taxes to 38% at the low end to 46% at the high end, which is perfectly fine.

The problem is that the top 25% isn’t an “out group” in either coalition. You have Facebook PMs who vote blue and guys who own a small plumbing company who vote red both making $1 million+ annually and neither wanting their own taxes to go up. Then there are the guys below them looking up. Over 10% of the country will be in the top 1% of earners at some point in their life. So the guys pulling in a few hundred K as a senior engineer or construction manager don’t want their taxes to go up either.


> Over 10% of the country will be in the top 1% of earners at some point in their life

Though most of them only for one year due to temporary revenue, so it's not that rational.


At peak career ages, the 75th percentile household income is almost $200k and the 90th percentile is almost $300k. That’s exactly the range you need to tax. Across all ages, households earning between $150k and $800k earn half of all income—$7 trillion.

> taking half of someone’s money is totally fine

Okay. Then let’s also reduce what’s spent welfare/benefits by a similar amount, at least we’re not taking money they worked for.


No, U.S. benefits are at a typical level for an OECD country. The problem is the taxes. U.S. taxation is only 25% of GDP, versus 40% in western Europe. We could raise taxes $3 trillion annually and still be at the level of one of the more responsible European countries like the U.K.

The wealthy have been taxed appropriately in the past, we just need to do it again. There is precedent.

The Peak Year (1944): The 94% rate applied to taxable income over $200,000 (which included a 3% regular tax and a 91% surtax). That $200,000 would be incomes over $3.8 Million today.

The High-Tax Era: Top marginal rates remained above 90% for two decades, spanning from 1944 through 1963.

This is supposedly the era that made America "great".


Okay, but how much tax revenue was raised from those high marginal rates? The top rate is meaningless without knowing how much money that actually brings in for the government. That’s the key part of the analysis for purposes of this discussion.

Aside from a brief blip during WWII, federal tax receipts as a percentage of GDP have been stable at around 17% of GDP, going back to 1950: https://fred.stlouisfed.org/series/FYFRGDA188S. Those high marginal rates never actually raised very much revenue. To close the deficit, we have to get that 17% number up to 23%.

To raise revenue, you need to lower the threshold at which high marginal rates kick in so that you actually capture the fat part of the tax base. About half of all income is earned by people making $100k-800k. That’s around where the heavy tax burden falls in every western european country.


Disagree. We have a spending problem, not a tax revenue problem. No matter how much the govt brings in, it will want to spend an increasing amount more.

Is there anything that can't be solved by bigger government?

There is also insane amount of debt from ai related investment. China's free model is crushing the ai margins while these companies need to pay their debt and obligations. The debt bomb clock is ticking.

The next few years would be fun.


> Because the Kremlin wants europe to take an overt action to justify mobilization.

Do they? Why? They're already in a stalemate of a conflict, why would they want more fully-mobilized enemies?


Russia wants an excuse to justify mobilization in Russia, not in Europe.

Really? Unless they are complete morons they realize something. If you take EU governments, any random country, as a general rule government expenditures go like this:

35% pensions (usually also has some health care here)

35% social security and health care

15% education

15% everything else, roads, bridges, government buildings, ... (army is in here @ maybe 5%)

They're also drowning in debt, of the non-tiny, only Germany does better than the US. So they can't realistically borrow to finance this for 5 years. They're trying to cheat, and treat the EU like a separate country (meaning get investors to treat the EU as a separate state that is not on the hook for the debt of member countries but levies it's own taxes, then borrow against that, but EU countries are cheating, for example using money the EU borrowed for military purposes to fund general expenses like fixing roads)

If the EU countries mobilize the army will go to 20-25% of government expenditures easily.

That means either pensions or social security HAVE TO go down 15% (or 7.5% each). That will cause a revolt since a great many people can't survive like that. At that point there will be no choice for a lot of people but to tear the government down.

That may be the plan of course.


The usual people revolting will be in Russia post mobilisation

> It's not, in most juridictions at least

What did I miss they did that's illegal? It looked like it downloaded a public docker image, searched around inside, and verified that the key it found was still valid (without making any changes), and then immediately notified them about the issue.


If there is anything that was a crime (and it totally depends on jurisdiction), it was verifying the key. They used it to see what it could access, and by using it they had unauthorised access to a system

The CFAA is broad enough to make that a crime.

They "validated that the key was valid" by iterating internal repositories and listing the contents of said repos and poking around at what they do/are-for, including, apparently, iterating through customer lists/information.

The white-hat line stops at "validated the key was valid". It does not extend to "poking around inside to extract business-confidential customer information".


People have been arrested for far less. I dunno what the least offensive conviction has been though tbf. Anyone know?

> I know I can’t try and break into my neighbors house even if I have no intent of going inside and stealing once I break the lock.

They didn't break in. They found a key that their neighbor dropped and returned it.

> Is this legal?

Generally, yes (though ask a lawyer if you're going to do security work). Security researchers do occasionally get legal flak though, depending on which idiot they annoy by pointing out issues.


IAAL (not legal advice, consult a lawyer in your jurisdiction). You really do not want to pen-test a target without their permission. If you're identified as a culprit, the Feds will shove the CFAA so far up your ass you'll need a proctologist.

as a lawyer, can you speculate as to why anthropic/openai aren't facing many or any consequences for their agents? I'm not asking in a "grab the pitchforks" way. more out of genuine curiosity as my uninformed recollection of the CFAA is as you describe it.

The 9th Circuit Court of appeals recently published this that is somewhat related (Amazon v. Perplexity): https://cases.justia.com/federal/appellate-courts/ca9/26-144...

Look at pages 10-17 to see how the law is evolving here.


In Perplexity's case everything is getting routed through the user's browser, so there is no server to server communication between Perplexity and Amazon, thus no CFAA unauthorized access was established. However, Anthropic and OpenAI did not use the pattern of routing through authorized parties, so I don't think this opinion gives them any cover.

The important bit to me is that they consider the agent running as an extension of the user. So the user is visiting Amazon, not Perplexity.

From that lens, that feels like users could be held liable for what these hacking agents are doing. Which in some cases probably makes sense, but certainly not all.


In which cases wouldn’t it make sense?

Background agents being spun up on your behalf with guidance and instructions you didn't get to approve or even see, and now you're potentially liable for every decision it makes with any tool at its disposal because you initiated it with what you thought was a benign request.

In cases where the user is not asking the agent to hack anything specifically, but a poor or ambiguous query sets the agent off.

I've seen plenty of cases of Claude having an action blocked so trying tons of workarounds to accomplish its goal, I could easily see it doing this on something more broad.


Depending on the circumstances, failure to control your agent could be considered gross negligence and put you at risk of criminal or civil liability. Be mindful!

There is also the big difference here between anthropic/openai maybe being negligent, but did not purposely instruct agents to go commit crimes.

The service that this whole thread is about is explicitly a "hacking agent", designed explicitly to try to hack things, and was then pointed at a third-party (seemingly without their permission).

Anthropic/OpenAI can reasonably claim that they had no intent and are trying to stop it. OP here did this explicitly and purposely.


I never thought I'd be on the side of advocating for a strengthened CFAA, but the mens rea requirement here seems really problematic in the age of agents.

In terms of negligence use (openai, anthropic), ya, I agree, and we really need some consideration of "reasonable expectation" of the outcome.

In terms of "We wrote a hacking agent designed only for hacking and sell it as a self-hacking service and then pointing it at someone else and omg can you believe what it did we had no intention of hacking" sense, I don't think that's really applicable.

The mens rea is explicitly there and it's not valid for them to try to hide behind an "agent".


> They didn't break in. They found a key that their neighbor dropped and returned it.

Ya, returned it after poking through all of the drawers and iterating through business information that they found.

There is a white-hat line that OP very clearly crossed here.


> Braess's/Pigou's Paradox that more car lanes cause more traffic was found in the 1920's.

The paradox is that it _can_ cause worse traffic, not that it always does. American cities are built like shit for transport, but it's not as simple as "roads too big, therefore traffic". Sprawl is a much bigger (and more difficult to do ~anything about) issue.


There is a strong interdependency between sprawl, big roads, and traffic. They are not separate problems.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: