Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

There are already some providers offering cheap LLM services that will give you a response within 24 hours instead of within seconds. That allows them to schedule tasks during low-request hours when they have spare capacity and use better batching. For some automated tasks this is perfectly acceptable. A bit of effort to accommodate, but easy to justify when it halves your inference costs


Any examples of such providers?


OpenAI [1] as well as Azure OpenAI, Anthropic [2], as well as Parasail [3] for all the "open source" models. There are others that I was thinking of, but those are the first I could find without my notes. Typically the batch API is 50% cheaper than live inference

1: https://platform.openai.com/docs/guides/batch

2: https://docs.anthropic.com/en/docs/build-with-claude/batch-p...

3: https://docs.parasail.io/parasail-docs/batch/batch-quickstar...


Certain tasks at OpenAI when I checked a few months ago. Embedding for one.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: