> I do like Claude Code a lot when it comes to pure coding use cases
The harness is great, but opus is such an arrogant little prick that spews out unintelligible word salad. Opus 5 is so bad at communication it amazes me that somebody green-lit it. It's absolutely awful.
The fact that this isn't an acknowledged regression (and Fable 5.1 isn't much better) leads me to believe that people at Anthropic actually like Opus 5's output.
I wouldn't call copy & paste code from a webui of chatgpt either traditional or boring. I'd call it tedious, error prone and guaranteed to get poor results. There is much better tooling and harnesses to leverage now.
Because not enough people do this earnestly, and many more do it maliciously (bot endpoints that lie, or provide significantly less information than people endpoints) or put it behind a business contract (yes, APIs), so the bots or agents can't trust it in general.
Also let's not forget that innocent sites suffering from floods of scrapers are actually the minority here - this is just a special case; the main reason for the tension is simply that most websites and businesses on-line rely on users wasting their time, and cannot abide any form of end-user automation. Their business plans hinge on their ability to force themselves on you.
Not to mention that it solves none of the rate issues. If the scrapers are hitting your site 10,000 times a day, adding markdown isn’t going to change that at all.
You’d think that, I thought that… but then I realized I’m just kidding myself thinking its output makes sense. It doesn’t. It doesn’t. Sometimes it might as well just speak tongues.
In other words, it ain’t you. It’s the model. It’s just genuinely bad.
Then you switch to ChatGPTs lineup and realize how things can actually be better. It took about a week to really get the feel for how to use their models… then I basically switched. I’ll check in every now and then when they actually make a deal about how opus “now makes sense”.
But honestly I’m half convinced Anthropic actually prefers the output of opus 5. I dunno why, but how else could you explain how such a thing got shipped? I mean somebody in the pipeline had to say “dude this model doesn’t make sense, you think we should fix it?” Right? Like it’s a pretty massive drop in quality for such a major brand in this space, you know? How did it make it out the door?!?
yeah, I just switched to the GPT models and it's a breath of fresh air.
as for the other guy, the claude talk is definitely not less ambiguous, it often is incredibly ambiguous and hard to parse, I have no clue why it produces such output, if not to fingerprint it?
it's really weird man. when Opus 5 came out, I was really confused. I saw a bunch of hype about how it's better than fable, but I just felt frustrated with it, although at times it'd do fine, but especially in Claude Code it'd just delve into the whole "load bearing" type of lingo real fast and I'd get a headache.
I don't think it's worth using even if it scores 2 points higher in some bs benchmark
it's definitely surprising how the magic and smoothness of 4.6 and such is no longer there with the >5 models
It’s the same thing with Covid. People thought 1 in 10 people would be dead. And believed it for years, despite it being orders and orders of magnitude wrong.
Humans love a good end-of-world story. Always did, always will. It’s when they want the world to conform to their ungrounded irrational bat-shit crazy nonsense that I get a worried. Especially when people in power actually take their nonsense seriously.
Anthropic needs to shut down. They are promoting insanity from every pore for some crazy reason.
The best message these people can get from the public is to have their IPO fail miserably. Nobody should buy one share.
Either they are crazy, in which case you should not give them your money.
Or, they are truly building a doomsday machine, in which case you should not give them your money.
Ergo: Don't buy their IPO.
I've been trying to steel-man the argument and find a way in which AI could kill all of humanity. Maybe I am too stupid to come up with a clever enough idea. You would have to assume that eight billion people are NPCs in the game of life and simply fall over without any thought, effort, action or desire for self preservation.
Sorry but… uh… bro. I always wondered if opus 5 was a regression or intentional. I’m seriously thinking it’s actually derived from how people at Anthropic talk.
The last people on earth who should be regulating this are governments and tech oligarchs. I find that vastly more scary (and plausible) whatever unintelligible nonsense opus and fable spew out these days. Seriously. The way those models talk and behave do more to show the limits of AI than anything else.
"Plus rigorously ensuring backwards compatibility for a project that is 2 hours old and has zero users."
That is exactly how the slop accretes and you get a pile of crap. Claude somehow assumes that said 2 hour old userless app is some dusty enterprise app with millions of users and billions of dollars at stake for a 1 second outage.
I have to constantly have these things "take a deep breath, step back and look at the entire thing and do this change holistically. please restate what i'm asking you to do and why it's important"
Which, again, if you attempt to use these LLMs in an actual enterprise project with the assorted legacy mess in it, that does have to have backwards compatiblity, and interacts with 'weird' tech (well, weird to full stack Node/React devs), well it's an exercise in frustration.
> please restate what i'm asking you to do and why it's important"
This doesn't actually do anything though, right? There's no understanding, so the machine will just reiterate the original token query back to you. The 'why it's important' part will just generate some patronizing boilerplate as a raison d'etre.
For one thing, it helps clarify my own thinking and uncover any gaps. For two, it lets me peek into its own understanding of the problem and ensure we are both aligned. For three, the garbage I spew at it is usually hastily typed crap from a mobile phone. So it helps clean it up and unpacks things.
It’s one one the main ways I know what’s cooking under the hood.
Ps: I used to have to say “in your own words” to make it work right but now days that isn’t required. Other times I’ll add “make sure to ground yourself before doing so” or something to encourage it to load the right shit into its context be it web, files, service calls, whatever. I’ll also some times add “why it’s important (or isn’t, push back if I’m talking nonsense)”
"They botched it." <-- sure, but worse.... they shipped it anyway. And that is the part that gets me. It's a vastly worse product than it was on like 4.6. I suppose you can (and should) use opus 4.6 -- they do make it available still. Then just treat 5 as something you avoid until they push out a new version.
The harness is great, but opus is such an arrogant little prick that spews out unintelligible word salad. Opus 5 is so bad at communication it amazes me that somebody green-lit it. It's absolutely awful.
The fact that this isn't an acknowledged regression (and Fable 5.1 isn't much better) leads me to believe that people at Anthropic actually like Opus 5's output.
reply