Astra is just so good. And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.
I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.
I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).
Edit to address questions below:
ChatGPT supports oauth login.
Exe.dev has it built in. IIRC, pi also has it built in via /login.
And that assumes Opus 5.5 Medium is actually equivalent to Astra High in all real-world usage/personal work loads, which isn't guaranteed as benchmarks saturate. The High vs. High comparison (probably not equivalent, but for reference):
Opus 5.5 High = $1.82
GPT-6-Astra High = $1.76
If Opus 5.5 Medium isn't equal/better for what you're working on vs. Astra High across the board, the price difference would narrow a bit more each time you had to switch to High.
So, if you're happy with Codex already it's not like Opus is now 1/2 the price and you'd be leaving a crazy amount of money/tokens on the table. Plus you have way more flexibility on the low end of the intelligence curve with GPT 5.6 Luna: Haiku (and Sonnet) can't touch that price/value ratio.
Once you get to roughly the 51+ Intelligence Index range, Opus 5.5 appears to define essentially the entire cost/performance frontier, from ~$1.34/task through ~$6/task.
Directly below this in the Cost per Intelligence Index Task table, the most efficient by far is Opus 5.5 Low.
In the end what matters is how much you pay for the task you want completed. And Astra will usually do that using less token and offer a better quality solution so in the end it might be cheaper.
This. I was fed up today with constantly correcting Opus for a specific task. So I finally decided to try Astra. It handled all of my prompts in one go.
yeah, Astra burned through 70% of my weekly usage in ~5hrs on a $100 plan. even fable doesn't run out that quickly for me. it's great, but it's on the same tier as fable for me - use it sparingly, only when really necessary.
What sort of things are you building? What do you like about exe.dev? Just curious what teh overhead and $20/month subscription is enabling for you.
>I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.
This is my experience with Claude code on my local machine. I suppose maybe you are doing something that naturally has system side effects? Obviously sandboxes have advantages sometimes but I havent seen a need for what I'm building.
For me, the essential difference between exe and local harnesses is that they have really nice built in methods to take care of stuff like auth and inter-vm interactions, which makes it easy to build live internet-connected services that I can share with other people. There is no deploy step, which is a surprising amount of time savings. They also have nice things like email receive/send, the ability to issue phone notifications through their app, and just a ton of little things where it feels right.
FWIW, the $20/month subscription also includes $20/month of LLM credits. That’s obviously not sustainable, but it should make it easier to try out the service. I would stick with them even if they dropped it.
Here is an invite link for a 30 day trial (that benefits me too if you were to become a paying member):
Shelley is a fantastic agent and their batteries included vm image makes the most of it. It includes things like a browser for Shelley to check its own work. And Shelley has root and full access to them, so it can solve any problem and do pretty much anything you need.
ya i've been a gpt hater for a while. almost exclusively used claude up until astra. astra feels like it blows everything out of the water. its fast, correct, organized, and less verbose.
Anthropic is absurdly vague about 3rd party harnesses for subscriptions, if you try to use anything besides Claude Code, you are likely at risk of getting banned, you can "do it", but are at their mercy if they decide to ban you. OpenAI gives their blessing to using oauth on any harness, you can make your own or use any of the popular public ones like opencode, pi, whatever exe.dev is that this guy mentioned.
So in simple terms, OpenAI doesn't restrict you to Codex, and gives their blessing to try whatever you want with their models(besides serving others with your subscription usage, that is still afaik against tos).
What worked well for me was a custom version of Open Web Ui with some customization to spawn an exe.dev instance for each new chat. I can just work on my phone, deploy stuff for development purposes on an easy to share way etc.
Better than Astra looks insane. I saw a Higgsfield video yesterday where they gave the same prompt to Astra and Opus 5.5 to create a samurai video game, and the difference was huge.
Where the US sees itself penalizing China with an export restriction, China sees the US gifting it with zero-political-cost “protective” import tariff.
You can’t really hurt a country that has a culture with a positive attitude toward growth.
But they accepted the transit fees, which means they endorsed the action.
The guild would not allow any action to jeopardize the flow of spice. Thus any attack on Dune is implicitly sanctioned, notwithstanding their spice trade with the Fremen.
> There is no additional fee for using prompt caching. Input tokens, whether served from the cache or processed fresh, are billed at the standard input token rate for the respective model.
reply