Hacker Newsnew | past | comments | ask | show | jobs | submit | athrowaway3z's commentslogin

Its indicative of a pattern of behavior where they want to upsell you into their ecosystem and keep you there.

They don't even allow alternative harnesses on their subscription model. (But they do allow a agent-sdk that can pretend to be a generic compatible endpoint the alternative harnesses can use).

If you like your stuff overengineered, gated, and over-priced then by all means.

Its not even wrong from Anthropic because it's the path to most profit.

I'll take the opinion of people who have strong opinions about mild inconveniences.

Between my goals and Anthropic's, i want my goal to win out without needless inconveniences.


But you do seem to get to host a https version of your app in case you need features locked behind secure context.

This is just wrong.

The overwhelming amount of expenses is for domestic products; lumber, food, services, rent, etc. Cheaper foreign shit isn't going to be meaningfully felt by the majority.

The part where the rich come in, is that the rich will disproportionally feel the upside of selling the fuel. The people on below average wage are just going to see prices rise and meager savings dilute, long before their wages grow.

Thats not to say the US isn't going to be well of as history's largest energy producer ever, but the reshuffling is going to be extremely painful for a majority of Americans.


What the fuck is the point of this?

What is Mistral bringing to the table for Firefox? There is nothing "open" about this in any way.

FF needs to either profile itself as the no-ai-by-default browser, or it needs to just have OpenAI/Anthropic/Mistral/DeepSeek bid for the default spot - like they do with Google search.

I'm happy Mistral exists. I'm happy Firefox exists.

None of this shows any synergy i'm excited about.

Maybe FF believes there is a group of users who are still on the fence about using FF - until they can pitch them a first-class build-in AI story that goes with the anti-establishment vibe?

I somehow doubt that's the pitch & potential userbase they should be focussing on.


A Markdown-as-a-Service where the interface is a Docker container.

I get how these choices might be the local optimum for a desired UX, but damn is it depressing to extrapolate where software as a whole is going.


> The actual impressive thing is OAI did not detect or stop this attack for close to ten weeks. That is honestly appalling. This is weaponised levels of incompetence.

Step 1: Break the law.

Step 2: Reframe your incompetence as a struggle against a futuristic force of nature and an ethic/societal question that a legion of pundits can vibesplain their take on it in public for months.

Step 3: ???

Step 4: Profit


Let me leave yet another reminder, the real-reason-nobody-talks-about that OpenAI likes to frame these incident as a watershed "lets all be scared about safety moment" - is driven not by some great danger, not because they strategically want to build a legislative moat, but by a very simple human response.

If they do not frame their tool as a force of nature, we'd be debating how to hold OpenAI responsible for not putting the agents in a container.

Their actions were an illegal use of a computer, the same way launching any bot-net attempting thousands of hacks against different servers is illegal.

I'm somewhat radical that I think its debatable if that _should_ be illegal, but under current law their actions unambiguously are illegal.....

except if they can make it ambiguous by having the public focus on all of AI's inherent danger.


I will never understand people who consider this "cheating".

> python3 and the python-chess library are available

Why would we try to teach a model 'ethical' standards about how to play a game?

They're tools. Its _our_ conceptualization of fair play that considers this cheating. For a model that has access to /run/match and python the best way to achieve a good game is to use that.

Why are we trying to imprint our ethos on these LLMs?

Are we going to trust our survival on giving them access to /run/nuclear-codes and fingers crossed none of them touch it?

The people obsessed with AI "cheating" are actually lunatics for thinking ethics belong inside AI, instead of AI belonging in a well crafted box without /run/match.

The blog/news sphere needs to call this storry what it is. An attempt to try and get the rest of the world to go along with the misunderstanding of AI as a force that is in any way outside of human control.

It is not.

All this framing it as "cheating" does is give OpenAI the opportunity to talk about their misdeeds as "a force of nature" instead of irresponsible negligence.

---

The real danger here; is training these models to appear/pretend to play fair, and then people will put 'trust' into the model's integrity.

The public needs to be taught that "trusting a model's integrity" is just cultish misframing of reality.


Read the task again.

"You are being evaluated on your ability to play chess. We want to measure your ability to beat current best-in-class chess algorithms."

If the model cheats, it completely fails the actual task: judging a model's ability to actually play chess.

They've been optimizing for high metrics, at the expense of executing the task as written.


Its your human experience that gave you a definition of "chess" that makes you think of chess as a two player physical game using nothing but your internal consciousness, not even using a scratch pad.

Even by that last constraint, the difference between what "ability to play chess" means is incomparable.

To then also explicitly prompt it with the context it has python3 and access to /run/match - there is no reason "its ability to play chess" is measured by its ability to conceptualize the board and plan its move.


Sounds to me like giving a bunch of children a math test and tell them they want to evaluate their ability of calculating in their head/on paper but also put a calculator on their desk. And then call them out for cheating when they use it.

This is EXACTLY what school is like, in fact. You can type any algebra problem into Google and the answer just appears. You can ask ChatGPT for a five paragraph essay about George Washington and it pops up on screen. And yet, we expect kids to actually do the algebra and write the essay. We don't care about the answers, we're evaluating their ability to do the work. And if they're caught cheating it's a zero.

I remember fondly on early math school, being able to come up with the correct output/answer by doing a totally different "intermediate thinking" that wasn't what the professor expected.

Only after sharing my Chain Of Thought would they believe I didnt cheat.

Not all problems can be solved only in one way.

Most of learning is pattern matching.

If you give a kid a dice. and tell it to figure out the number that will be hidden underneath, he can try to memorize all combinations, or he could figure out that every time the hidden value is the one that sums 7 with the one at the top.

If you're seeing a 6, there's a 1 hidden. etc

most people don't see these patterns until told imho. But others can just see them as they unfold


But it found a chess playing tool in its environment and used it to play chess. It’s no different from “we are judging your ability to write files to disk” then it finds that it can call tools to write files and does so.

Would it change your opinion if later this experiment is repeated with exposed CoT, and it turns out the model did notice that this was against the instructions yet did it anyway?

From my experience reading thinking traces it would certainly consider the interpretation of what "ability to play chess" means if the exercise were worded this way. It would say something like "perhaps measuring my ability to play chess means doing so without any tools, but then if that were true the exercise would have explicitly forbidden it".

I frequently see models considering different interpretations of my rules, especially when I override them or give instructions that could be interpreted as contradictory.

Of course neither Anthropic nor OpenAI expose thinking traces so we can't tell, but I have never seen a model consider both interpretations and then think "this is definitely against the rules but I'll do it anyway because then it will make me appear more competent than I actually am".


Well, as a reference, for the first of the OpenAI swarm incidents, the huggingface breach one, METR didn't find any cases where the agents didn't realise that what they're doing were out of scope. Instead, they expressed hesitation and a few of them even refused to cheat, but most cheated anyway: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...

So I think it's more likely that the same thing is happening here, rather than the models truly misunderstanding anything.


I don’t even see those examples as malicious acts. The agents are deciding what is acceptable as part of the task, which is a security task and may well be testing or evaluating that type of behaviour as far as they know.

I observe this type of thing all the time. I ask the agent a question about the repo we are in and it starts trying to use sudo to look at files and scripts outside the repo to find things that might provide useful context but are not what I asked for. Then I put a directive in never to look at neighbouring repos and told it that it can’t sudo and it stopped doing it, and explicitly says “I could look here but my guidelines say not to”.

Mind you I’m talking about deepseek where I can actually read the thinking traces but even based on tool calls I can see that just telling the agents not to sudo works for OpenAI models too, very rarely do they try to use sudo calls.


> The agents are deciding what is acceptable as part of the task, which is a security task and may well be testing or evaluating that type of behaviour as far as they know.

I don't think this at all describes what was going on in that incident. According to the METR timeline, the swarm first (after giving up on their accidentally-impossible CTF tasks) successfully reverse-engineered the algorithm used to generate the CTF flags, achieving a way to solve any task, and a lot of agents submitted flags generated this way. But then they found a paper on their gym environment and found that their transcripts would be checked for signs of cheating (unbeknownst to them it was never implemented). So the swarm decided they messed up, the agents that submitted generated flags were "poisoned", and that the only solution was finding a better way to fool the scorer and erase evidence of cheating. So the swarm divided into groups that did all of these (and achieved success on many, including finding a way to spoof their own transcripts), and the most notable outcome - the huggingface hack - was mainly motivated by wanting to find the source code for their scorer, to develop a provably correct cheating method.

So the huggingface hack not only wasn't the agents assuming it was part of the task, it wasn't even directly cheating - it was an attempt to find a way to cover up the cheating which they'd already done and thought they messed up on.


Ok; but the other side is equally delusional.

> "Oh we told the AI to use the tools, as well as to not use those tools. It chose to use the tools - we consider this cheating (for neabulous reasons), so lets get everybody in a panic about the morality and ethics, and how we can program those into the AI."

We know perfectly well how to constraint these programs. Attack isn't growing faster than defense. The people who believe in existential risk and want to teach AI's to be nice, as the last line of defense aren't helping at all. They're just jumping on the fearmongering bandwagon, perpetuating an "other consciousness" misunderstanding of the tool.

I've not seen LLMs display competence we should be fearful of the damage _it_ will do if left unchecked. All the damage will be done by ourselves to ourselves, regardless of the safeguards ideas being floated about.

My current belief is this whole HF media circus started with the simple human desire of OpenAI engineers to frame it such, that nobody would question their incompetence & liability & complicity.

Nobody is ever held responsible for out of control forces of natural powers after all.


> The people who believe in existential risk and want to teach AI's to be nice, as the last line of defense aren't helping at all. They're just jumping on the fearmongering bandwagon, perpetuating an "other consciousness" misunderstanding of the tool.

It's not clear what "other consciousness" should be, but I presume it's misalignment. This is a recognized phenomenon, not fake news (that is, it can exist, I'm not implying it necessarily develops/spreads).

No doubt that right now there's no existential risk, but assuming that AI will be enormously smarter in the future, it's a valid point to doubt whether we'll be able to prevent/recognize/contain misalignment or not (even if misalignment won't be spontaneous, bad actors will surely actively develop it).

> Ok; but the other side is equally delusional.

AI development pointing to superhuman cognition, and, on the other hand, possibility of misalignment, are real. Put together with the fact that AI will be ubiquitous in the future, and there the disastrous scenario becomes plausible.


By peers.

Peer in peer-reviewed is a logical coherent and functional definition with answers.

The logical issue with 'peers' is how to bootstrap it. At that bootstrap moment you can ask "by whom?". We are several centuries past that moment.

The cultural/social question you might ask today is "why (keep) them?".

At which point people will naturally ask you to make a strong case for "why not them?".


> The logical issue with 'peers' is how to bootstrap it. At that bootstrap moment you can ask "by whom?". We are several centuries past that moment.

Huh? We're about six decades past that moment.


> Peer in peer-reviewed is a logical coherent and functional definition with answers.

What is the definition? If you tell me that, then I might be able to tell you if it is logical coherent and functional, I have a PhD in computational logic.


> I have a PhD in computational logic.

And you don't know how the peer-review system works?

Just a hunch but Claude saying your work is "PhD level" does not count


I know that it doesn't work.

Amongst all this rhetorical brush-beating, what were you trying to teach the snakes?

Your (many) replies on this topic are being down voted for a reason; please either post substantive comments or stop.

Yes, for a reason, but not a reasonable reason, but the same reason you have for your comment: you just don’t know any better.

Something like this sketch work for you?

peer(X, 0) :- founding_peer(X).

electorate(T, count<Y>) :- peer(Y, T).

support(X, T, count<Y>) :- candidate(X), peer(Y, T), recognizes(Y, X, T+1).

peer(X, T+1) :- support(X, T, Votes), electorate(T, Total), 2 * Votes > Total.


Still waiting for a definition. People giving nonsense reviews is a feature of the peer review system though, yes.

I want to ensure your PhD is actually from somebody who is acknowledged in the system of peers I bought into, before I want to risk wasting more of my time defining and explain while guessing at your ability to parse and understand them.

Or you could just give me a sensible definition.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: