Hacker Newsnew | past | comments | ask | show | jobs | submit | palguna26's commentslogin

At this point, i firmly believe all companies are building benchmaxed models, which perform well on older benchmarks but struggle on new one, terminal bench is the best example.

I just used it yesterday, and overall it does a very clean job, i was given the architect title, which is really close to how i actually use coding agents, it also gives u feedbacks on where you were weak as well, so you should definitely give it a try as its free.


I personally use Codex(plus) a lot, with sol+ponytail, u get concise to the point answers rather than long explanation that openai models are known for, and till date im really satisfied with its coding performance. I also use opencode to try out new openweight models as well, currently im testing out k3.


I believe customer support should actually be done by humans, as customers feel undervalued when they talk to scrappy voice agents, atleast make an effort to use good ones, so it tries to speak like humans


Im building termyte, the runtime safety layer for ai agents.


I just recently shifted to codex since i got frustrated with the token usage which did not allow me to get my work done. Here's my honest opinion i think with the right configs codex does a very good job, like cc is better in terms of quality of code generated but codex seems to understand everything way better, and you can improve your code quality if just read through it once.


I just shipped a causal memory system for AI agents and am now working on the mcp for claude code. It's open source u can check it out on: https://github.com/CausalOS/causalos-python


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: