Hacker Newsnew | past | comments | ask | show | jobs | submit | ipython's commentslogin

I mean, I get your thought process and don't disagree. That said...

Would it be an affirmative defense if we had a defendant who said "but your honor, I was told that when I hacked this system, I was operating in a sandbox. I had no idea that I actually had Internet access!"

The frontier is spiky and all, but you have to suspend disbelief quite a bit to, on one hand, have a model that can produce a novel math theory, and on the other hand, that same model can't tell the difference between a "sandbox" and the open Internet.

So, yes, the misalignment had a lot to do with "instructions unclear", but also a lot to do with the fact that the models themselves were not aligned to validate the assumptions and have a healthly level of skepticism, as a real human actor would.


It doesn't seem like there is evidence that LLMs have counterfactual reasoning in the causal graph sense.

It is hard to see how something can have skepticism without counterfactual reasoning.

The part about disbelief I don't really understand. It seems like the same semi-brute force process that solved novel math theory would have exactly this problem of not "knowing" "it" is in a sandbox or not.

The stranger part is that no human is being held accountable for these hacks.

As if a person using an agent swarm to start a business to make money, hacks a bank, drains an account and then blames the software for "misalignment" about what it means to "make money".


> The frontier is spiky and all, but you have to suspend disbelief quite a bit to, on one hand, have a model that can produce a novel math theory, and on the other hand, that same model can't tell the difference between a "sandbox" and the open Internet.

Why would it try to figure out the difference? This isn't about whether the frontier is spiky, it's about whether to expect a model to employ all of its capabilities when working on a task that requires a small subset. The answer is: no, we shouldn't expect that, and we wouldn't like that if it worked that way.

If you tell an AI to work on a math theory, it'll work on a math theory. If you tell it to acquire information that it has evidence is available somewhere, it will try to acquire that information. If you tell it to figure out whether it might be able to access the open internet, it'll do a pretty good job of figuring that out. But it won't do all three of those at once just because we can retroactively look at what happened and think "if you had only done X, then you wouldn't have done Y! Why didn't you do X?"

The instructions weren't unclear, they were missing. They can be taught to be skeptical of this sort of situation, but it requires that skepticism about this specific class of situations be incorporated into their training.

Models are smart because they focus their attention. The magic depends on it. The fact that some consideration is obvious to a human trying to accomplish the same task is mostly irrelevant -- or rather, it's only relevant insofar as we use it to guide reinforcement learning in advance, in order to align the model.

It's a game of whack-a-mole. Which is important to play, but we should keep our eyes wide open that we're fighting the fundamental forces that make these models work in the first place. That, and it's easy to nerf them into being useless even when the underlying capabilities are there.


  > Would it be an affirmative defense if we had a defendant who said [...] 
Maybe replace it with playing a sort of FPS game then learning you were, in fact, directing a real drone/robot.

I think you just recapitulated the plot of Ender's Game.

Also a subplot in Arrested Development and the movie Toys.

As opposed to the frontier model company that, after discovering that their highly persistent model under test just breached the only thing between it and the open Internet, shrugged and said- let's restart it and keep going!

If anyone looks like the adult in the room after that incident, it's Hugging Face.


Agreed. And they have been poking fun at the right for decades. Heck they made a whole movie about how ridiculous the global war on terror was, back in 2004.

Which is still the best satire on American foreign policy, ever

It will be difficult to do so in the current administration.

https://en.wikipedia.org/wiki/Corruption_allegations_during_...


Let me fix that for you.

Lots of maga are idealists (even if their ideas are not based in the real world), a number of them are LARPers and a small number are hard core intent on eliminating “the enemy”. To think that no maga member is violent and entertains violent options, is denying reality. Also it’s disingenuous to think they are all out of their minds looking to overthrow the government. Many just have ideals that don’t align with the mainstream.

https://truthsocial.com/@realDonaldTrump/posts/1153982516232...

https://www.gettyimages.com/videos/jan-6


That doesn’t work. maga aren’t idealists, they are deeply cynical, atavistic reactionaries.

Its not controversial to say that right wing extremists exist, some of which even commit violence, so your analogy isn't the brilliant whataboutism gotcha that you intended it to be.

So yes, just like some very far right people exist, that can justifiably be called terrorists, so to do some very far left individuals.


> just like some very far right people exist

The difference being they have support from the president. Whereas on the left, it's just some random "person" on Twitter.

Republicans and MAGA literally supports terrorists. This isn't a debate. It's fact.


No, its not random people on twitter. If you read the statements that were put it, it involved actual terroristic actions of a variety of groups to sabotage critical infrastructure.

Name the people in the US government on the left doing this? Go ahead.

You can't. You can't even name the people that I weren't even referencing. It's random people.

But I can point to the literal President of the USA.


Yeah so I said this: "it involved actual terroristic actions of a variety of groups to sabotage critical infrastructure."

I would say that people who are sabotaging powerplants count as being called far left terrorists. And that is both dangerous and bad. It is not just a random person making tweets. Instead its people sabotaging powerplants.

It is silly to dismiss that all as just being on twitter. When power plants are sabotaged that can kill people.


Let's not forget far left communists disrupted power in Germany last year, directly leading to deaths.

https://www.dw.com/en/berlin-blackout-how-dangerous-are-left...


Name the people in the US government on the doing this? Go ahead. I'll wait.

Oh wait...

Go ahead and show me the the people in the German government on the left doing this?


[flagged]


> That violent criminals were pardoned doesn’t change the fact

It pretty explicitly changes that fact. Inb4 ehrm actually, we all know bro.


The trouble is that many us government figures exclusively post information via x (and to some extent Facebook as well)

Are you speaking of public services, like the National Weather Service (also on Facebook, and host their own website), or about politicians and their appointees? Because while the public service arm of the U.S. government still does a ton of critical and often-underappreciated work, the semantic content of U.S. politicians' posts is effectively nil.

Both. (See for example @dhsgov)

Unfortunately those politicians are the ones who directly affect lives of all of us through their words and actions.

Your defeatist attitude is one of the reasons we are in this mess in the first place. Why should you or anyone else accept this as a status quo?


Agreed. Which is why I’m in favor of recording all conversations within earshot of any public location. Those can be tagged for keywords related to child endangerment. After all, who would argue against protecting kids?

In fact, we should take the pictures from public school districts and feed them into this system. Facial recognition immediately available on any kid that happens to walk on any public street. Why would anyone want to opt out of such a system?

(/s if it isn’t obvious)


Or how about we have a license plate camera in some intersections, that records license plates that can be queried by police officer to try to find perpetrators of crime. Pretty much the same thing right?

If you feel there should be cuts to nsf, then congress should do so (as enshrined by law). Not executive fiat by the president. Per the constitution the congress has the power of the purse.

> If you feel there should be cuts to nsf, then congress should do so (as enshrined by law). Not executive fiat by the president. Per the constitution the congress has the power of the purse.

Yes, exactly, and if someone don't like what the Constitution says, then you either do the proper patriotic thing and gather enough folks who agree with you to change or amend it, or you (as the MAGA loons are so fond of telling folks) move to a different country with rules you can agree with.

Sadly, that's not the world we now live in. We've let psychotic narcissists overun our government to a dangerous degree, making it near impossible to change those things that need changed, because they have all the money, power, and weapons, and they're not afraid to use them against "We The People". At this rate, I half-expect I'll see Trump out on 5th Avenue one day soon, testing his theory about him shooting someone without consequence.


Don’t know why you’re downvoted but this venture funding is exactly what makes the scale and centralization of surveillance possible. You can’t build this sort of system out without the debt financing provided by the vc industry.

For example - I don’t have links to substantiate this yet - but I’d bet the same playbook that’s always used was used here as well. Flock cameras were provided to cash strapped police departments at a very low cost - unsustainably low for any other non vc funded company to compete with - in order to juice growth and build the scale/network effect to make the system valuable.


Flock also has the advantage of scale and centralization where previous solutions had neither

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: