Hacker Newsnew | past | comments | ask | show | jobs | submit | bdlowery's commentslogin

The fact that gemini 3.8 flash is so high up there just tells you this is an awful benchmark.

Try and use gemini 3.8 yourself for any real world work and you'll see it's terrible. It'll just go in circles reading the same file 20 times for no reason making hundreds of tool calls for a simple change.

EDIT: I was using gemini cli... it's not a harness issue lol


There's no such things as gemini cli these days. It's called "agy" (short for antigravity). And if you don't know what that is, you're probably 3-6 months behind already.

PS: Just Googled it to confirm: Gemini CLI was deprecated on May 19th, 2026. The correct harness is called agy or antigravity for Gemini 3.8 Flash.

https://developers.googleblog.com/an-important-update-transi...


The benchmark page itself asserts that it used Gemini CLI as a harness. I came to the comments just because I noticed the error. For my part — using agy — I found Gemini 3.8 Flash mid.

Hard disagree. I use 3.8 flash in Antigravity a lot, and thoroughly prefer it to most Pro-class models. It's really fast, and I've had it make crazy progress on compiler-like problems that previous models including Opus simply failed at. On ultra plan you can have it going for hours, and make incremental progress with good prompting for review interrupts. It solved a problem I couldn't solve for weeks in under 6 hours. 10k LOC total. The harness and test suite is key.

This is my exact experience with the model - https://x.com/ThePrimeagen/status/2095565354726502683

And it just BURNS tokens like crazy.


Why so brazenly confident? Isn’t it possible that the benchmark is correct, and your experience is correct too, but you haven’t tried all the thousand different modalities of work that programming encompasses and so maybe you don’t actually have standing to judge?

very outdated experience from me: when I first tried gemini something, in an existing rust codebase, it looked around for files that would indicate if its go, javascript, java or c++ project, then declared I must have asked it build a new app in javascript and proceeded to circle around to figure out how it can install node and npm on my machine.

So I totally believe that Gemini is just bad. Which is surprising because Gemma is very good for some tasks, but I never ever had any success with Gemini, be it in cli or chat thing or anything else that has gemini branding.


You sure it was 3.8 Flash? It hasn't been called Gemini cli in a WHILE...

havent tried that model, but it sounds like a potential harness issue. Have you tried it in different harnesses?

i appreciate the feedback, the benchmark is primarily long horizon real world engineering tasks on big private codebases.

so 10 minutes is long horizon now?

benchmaxxed model


Good.


i've found this site easier to use and have better results than alternativeto. https://openalternative.co/


I’d have given them bonus points if they’d used “alternativetoalternative.to”.


> I don't want a full screen pop up telling me to sign in or a persistent bottom bar telling me to open the tweet in the app or a reply section mostly hidden unless I log in.

make an x account then?


I really don't want to support a man who has killed so many people and advocates so aggressively for white supremacy and the elimination of trans people.


no


you seem to be uninformed. This should help: https://www.youtube.com/watch?v=H_c6MWk7PQc


A link to a YouTube video without context e.g. summary of findings, primary author/creators, and primary citations, methodology, and so on is a next to useless for making a point. For all we know you’re linking to a crackpot or an industry sock puppet and I don’t care to “watch” any of what can potentially be a dubious video or worse a malicious video. It is your job as the linker to convince me the video is worth even one iota of my time.

On the other hand recent papers highlight the validity of concern/suspicion:

> This systematic review demonstrates that the environmental footprint of artificial intelligence is a structural and increasingly consequential challenge, shaped by interdependent decisions across algorithms, software pipelines, hardware infrastructures, and deployment contexts. The synthesized evidence shows that energy consumption and carbon emissions associated with AI systems are highly variable, context-dependent, and often underestimated — Beyond Efficiency: A Systematic Review of Energy Consumption and Carbon Footprint Across the AI Lifecycle (published in “Sustainability” an international, peer-reviewed, open-access journal) https://www.mdpi.com/2071-1050/18/3/1359


"I know nothing about the subject, but I would _guess_ that the power consumption is bigger concern than the reported and studied water usage."

That's the whole 21 minute 59 second video in a nutshell.

A loud and hectic quick cut rambling video essay. Millions of views, naturally.


> raised heat is a kind of pollution

> water is not unlimited

> you don't want to release a bunch of acid into a river

> I know next to nothing about this

> I know basically nothing

> corn is one of the thirstiest major crops grown in the US

What is this rant supposed to inform? Whats wrong with OP being concerned about the costs of operating a DC?


Nobody is concerned about the costs and environmental impact of data centers.

When someone in a discussion about the benefits of AI goes "did you think about the environment?!", it's always performative.

The real motivation is disliking AI itself or doubt about the government's ability to offset the labor market impact. Discussions that start with feigned concerns being raised are nearly always going to be unproductive.


Wrong.

I don’t dislike AI but really of mine are negatively affected by climate change and AI isn’t helping what is easily observed when Google and MS scrapped their CO2 reduction targets.

So every time I use AI I think about the necessity and usefulness of what I‘m doing with AI and if the use outweighs the costs.

Since the rise of AI the environmental impact doesn’t seem to matter anymore.

I guess because it’s the shiny new toy of the hackernews audience.

Privacy also lost importance given the fact that the same people who refused to give information like their phone number to companies like Google and Meta now upload their whole life to their AIs to asks what should the eat, hyperbolically speaking


You're not helping your case by questioning the usefulness of software produced with AI in the same comment section, or complaining about other people supposedly compromising their own privacy in overusing AI.

The reason it doesn't matter is because the environmental impact is moderate, and the benefit obviously tremendous.


The environmental impact is anything but moderate and the benefits are not obviously tremendous at all. I'm happy using Claude Code as much as the next guy, but saying that the impact has been "tremendous" is vastly overstating the actual results.


I disagree, with the projected doubling by 2030 we're looking at 3% of global electricity consumption or 3.4 EJ, less than 1% of final energy consumption.

That is moderate. Energy-intensive industry is around 130 EJ, and global final energy consumption > 450 EJ.

Existing documented applications of today's AI have the potential to decrease energy consumption by >13 EJ/year by 2035.

Now that was about operational energy consumption. Someone might bring up manufacturing and construction.

From what I could find the climate impact of those are estimated somewhere between 10-35% of the total climate impact of data centers, so relatively small compared to the operational energy consumption.

It is very hard to justify more than moderate environmental impact here, in my opinion.

For the benefits of AI, my personal results have been great, so I am quite optimistic. And objectively, I find it hard to ignore recent results in mathematics and security research.


"Global data centers consumed around 415 terawatt-hours (TWh) of electricity—about 1.5% of the world's total electricity—with AI acting as a primary accelerator for new power demand."

https://www.iea.org/reports/energy-and-ai/energy-demand-from...

That is today where we already consume too much. If by 2030 AI's consumption doubles it gets worse. While training large models draws major initial power, everyday AI usage (inference) now drives roughly 80% to 90% of cumulative AI energy

"Existing documented applications of today's AI have the potential to decrease energy consumption by >13 EJ/year by 2035."

Seems like AI helps slowing down the rise of energy consumption.

We are at a point where we want less CO2 not moderataly more. In the end more is more.

If your doctor tells you to lose weight or you get sick it's not a success to gain weigth slower


Well, at least we could establish that the environmental impact is moderate rather than extreme.

And if the potential of >13 EJ/year is actually realized, it would seem like the net impact of the data centers is not just "moderately more CO2" but possibly "moderately less".

Please avoid low quality analogies on HN.


Please avoid acting as the judge for analogy quality.

Anything but a reduction is bad and AI is a setback for that.

Potential benefits are as long useless as they aren’t realized.

BTW the energy consumption reduction is achieved by what kind of AI? LLMs?


I'm pretty sure that most of people private data contains data about third parties too. It's more as their own privacy they compromise


What make you think I‘m talking about water consumption?

The construction of data centers needs resources and also creates more the CO2.

The energy for these data centers is often created through fossil fuels which also creates additional CO2

What do you think why Google and MS scrapped their CO2 reduction targets


GP wasnt mentioning water consumption.


use codeberg if you're ok with never using ai in any of your projects. Any AI use and your repo is banned. - https://blog.codeberg.org/protecting-our-floss-commons-from-...


if you use ai to program at all you can't use codeberg.


You've made the same exaggerated claim in two comments. Not only are you not banned from Codeberg if you have used LLMs to help you write code "in any of your projects", but you're not even prohibited from using LLMs to help you develop projects that you host on Codeberg.

What you can't do is "share projects that mostly consist of code written by 'generative AI'-tools". <https://codeberg.org/Codeberg/org/commit/71149c7fc95ccfeae36...>


I feel like you just directly contradicted yourself so there must be some nuance I am missing. Or is the key word "mostly"? Like I can use LLMs to help the work but still have to type most of the code myself?

I have not written more than maybe 10 lines of code in the last year so it seems I am prohibited from hosting on codeberg?

Or is the key word "share" and that's somehow different from "host"?


I suspect the fidelity of your comment to your actual level of confusion is low. In any case, if you're really this confused by the information provided earlier, then the chances aren't good that there's anything that anyone could say to get you unconfused.


Right... That's exactly the kind of nonsense that turns people away from ecosystems.

Good luck with that.


The scope of confusion expands unabated.

1. I am not Codeberg.

2. Turning away people doing unwanted things is the whole purpose of Codeberg's new policy. Why you present this as a undesirable side effect rather than exactly what the policy was designed to do is the only difficult thing to comprehend in this thread.


It's not exaggerated. If you use AI at all for any code in the project it cannot be hosted on codeberg. They made that pretty clear in their blog post.

You can read it yourself - https://blog.codeberg.org/protecting-our-floss-commons-from-...


> If you use AI at all for any code in the project it cannot be hosted on codeberg. They made that pretty clear in their blog post.

No, they didn't. That's not what the blog post says.

> You can read it yourself

I didn't fail to do the reading beforehand. I posted a direct link to the change in the TOS and which is currently linked at the top of all Codeberg pages. Can you read it yourself?


See also the codeberg blog post (https://blog.codeberg.org/protecting-our-floss-commons-from-...) which has more informal guidelines. That includes projects that are:

1. created by LLM agents -- any vibe-coded projects;

2. mainly written and maintained by LLMs -- this would cover the recent changes to the rsync project;

3. tied to the LLM ecosystem -- this covers pytorch, llama.cpp, cursor, SillyTavern, AI skills repositories, and a whole host of other projects.


Good point, I forgot they’ve recently banned it. Updated my comment.


I'm sure you'd like $100 then to give the founder of pangram an essay they classify as AI that is in fact not ai.

Free money, right?

https://x.com/max_spero_/status/2085041200923394459


pangram says this is AI generated and has a very, very, VERY low rate of saying something is AI generated wrong.

https://www.pangram.com/history/0b397596-fa85-4869-bae3-2b1d...


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: