Hacker Newsnew | past | comments | ask | show | jobs | submit | abtinf's commentslogin

The most surprising thing about this to me is how low the number is.

Astra is just so good. And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.

I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).

Edit to address questions below:

ChatGPT supports oauth login.

Exe.dev has it built in. IIRC, pi also has it built in via /login.


> And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

This is news to me. Excited to try it out! Thanks.


news to me as well. i thought you were forced to use Codex if you wanted their subscription. I completely ignored it because of that. How do we do it?

With the pi harness it just opens the browser (or gives you a link if you're on headless) and you sign on as usual.

How about cheaper? Astra is $10 in $50 out, Opus is $4 in $20 out. Even on a subscription you'll get considerably more usage out of Opus.

Per the link someone else posted, the actual difference in $/task is not nearly so stark:

https://artificialanalysis.ai/models/releases/claude-opus-5-...

  Opus 5.5 Medium = $1.34
  GPT-6-Astra High = $1.76
And that assumes Opus 5.5 Medium is actually equivalent to Astra High in all real-world usage/personal work loads, which isn't guaranteed as benchmarks saturate. The High vs. High comparison (probably not equivalent, but for reference):

  Opus 5.5 High = $1.82
  GPT-6-Astra High = $1.76
If Opus 5.5 Medium isn't equal/better for what you're working on vs. Astra High across the board, the price difference would narrow a bit more each time you had to switch to High.

So, if you're happy with Codex already it's not like Opus is now 1/2 the price and you'd be leaving a crazy amount of money/tokens on the table. Plus you have way more flexibility on the low end of the intelligence curve with GPT 5.6 Luna: Haiku (and Sonnet) can't touch that price/value ratio.


Once you get to roughly the 51+ Intelligence Index range, Opus 5.5 appears to define essentially the entire cost/performance frontier, from ~$1.34/task through ~$6/task.

Directly below this in the Cost per Intelligence Index Task table, the most efficient by far is Opus 5.5 Low.


Also Opus 5.5 is often made useless due to its [cyber] guardrails (even worse than Astra), they're even worse than Astra's

In the end what matters is how much you pay for the task you want completed. And Astra will usually do that using less token and offer a better quality solution so in the end it might be cheaper.

This. I was fed up today with constantly correcting Opus for a specific task. So I finally decided to try Astra. It handled all of my prompts in one go.

A cheaper price has no value if I can’t use the thing I’m paying for.

The Claude lock-in simply disqualifies anthropic entirely (for my use).


> Even on a subscription you'll get considerably more usage out of Opus.

That's an incredibly bold assumption.


It's not, I have a subscription to both and Astra burns usage like crazy.

yeah, Astra burned through 70% of my weekly usage in ~5hrs on a $100 plan. even fable doesn't run out that quickly for me. it's great, but it's on the same tier as fable for me - use it sparingly, only when really necessary.

What sort of things are you building? What do you like about exe.dev? Just curious what teh overhead and $20/month subscription is enabling for you.

>I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.

This is my experience with Claude code on my local machine. I suppose maybe you are doing something that naturally has system side effects? Obviously sandboxes have advantages sometimes but I havent seen a need for what I'm building.


For me, the essential difference between exe and local harnesses is that they have really nice built in methods to take care of stuff like auth and inter-vm interactions, which makes it easy to build live internet-connected services that I can share with other people. There is no deploy step, which is a surprising amount of time savings. They also have nice things like email receive/send, the ability to issue phone notifications through their app, and just a ton of little things where it feels right.

FWIW, the $20/month subscription also includes $20/month of LLM credits. That’s obviously not sustainable, but it should make it easier to try out the service. I would stick with them even if they dropped it.

Here is an invite link for a 30 day trial (that benefits me too if you were to become a paying member):

https://exe.dev/i/rlDF6GI5PGBZV4P

Or

ssh rlDF6GI5PGBZV4P@exe.dev

Edit to add:

Shelley is a fantastic agent and their batteries included vm image makes the most of it. It includes things like a browser for Shelley to check its own work. And Shelley has root and full access to them, so it can solve any problem and do pretty much anything you need.


The astroturfing and guerilla marketing on HN is getting insane.

ya i've been a gpt hater for a while. almost exclusively used claude up until astra. astra feels like it blows everything out of the water. its fast, correct, organized, and less verbose.

Yes. Also, you get image generation included with the ChatGPT subscription, which is very nice for certain kinds of development.

> And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

Can you give more details here? This sounds intriguing.


Anthropic is absurdly vague about 3rd party harnesses for subscriptions, if you try to use anything besides Claude Code, you are likely at risk of getting banned, you can "do it", but are at their mercy if they decide to ban you. OpenAI gives their blessing to using oauth on any harness, you can make your own or use any of the popular public ones like opencode, pi, whatever exe.dev is that this guy mentioned.

So in simple terms, OpenAI doesn't restrict you to Codex, and gives their blessing to try whatever you want with their models(besides serving others with your subscription usage, that is still afaik against tos).


What worked well for me was a custom version of Open Web Ui with some customization to spawn an exe.dev instance for each new chat. I can just work on my phone, deploy stuff for development purposes on an easy to share way etc.

If you read the page, Opus is now significantly better than Astra while also being cheaper and having more performance headroom available.

I read the page. It seems like a marginal improvement.

Better than Astra looks insane. I saw a Higgsfield video yesterday where they gave the same prompt to Astra and Opus 5.5 to create a samurai video game, and the difference was huge.

Let's wait for independent benchmarks at least

the benchmarks provided are already from independent organizations:

Terminal-Bench 4.0 - Stanford & Laude Institute (with funding from all of the AI companies)

FrontierCode v1.1 - Cognition

CursorBench - Cursor (now SolarBoringSpaceXAI I believe)

GDPVal-AA - Artificial Analysis

AutomationBench - Zapier

Humanity's Last Exam - CAIS and Scale AI

Terminal-Bench-Science - Stanford, Laude, Ai2, Allen Institute

OSWOrld - XLANG Lab @ the University of Hong Kong

Chartography - Surge AI



Where the US sees itself penalizing China with an export restriction, China sees the US gifting it with zero-political-cost “protective” import tariff.

You can’t really hurt a country that has a culture with a positive attitude toward growth.


Good. Images consume even fewer tokens than text.


Amazon’s chatbot processed a price adjustment the other day without forcing me to call or chat with a human agent.


It is not a “clarification”. You are wrong and are spreading misinformation. In Farsi, the word for tomato is gojeh farangi.


But they accepted the transit fees, which means they endorsed the action.

The guild would not allow any action to jeopardize the flow of spice. Thus any attack on Dune is implicitly sanctioned, notwithstanding their spice trade with the Fremen.


This is why I love Hacker News. Come for the AI system outage root cause, get a treatise on the power of the spacing guild in Dune.


> capital offense

No need. They simply cut off the offending faction from all space travel.


> actual harness makers

I agree IBM Bob is not the revolution, but…

We are still in the punch card era of LLMs. Maybe even the Altair era.

None of this stuff has really been figured out yet.


Yes my first instinct is to roll my eyes at these efforts but we'll probably all get something better in the end from all the competition.


> How are cached tokens priced?

> There is no additional fee for using prompt caching. Input tokens, whether served from the cache or processed fresh, are billed at the standard input token rate for the respective model.

Well, talk about flipping the narrative.


heh

Is there a speed increase or is that purely marketing spin on “we might cache on our end but no discount for you”?


Pure marketing.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: