Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The big question for local LLMs is whether there is a 100 tok/s model which requires less than 16 GB of memory and is competitive on most tasks with the cloud models.

There is some signal that this is possible through both hardware innovation and training/data improvements.

Cloud models have their own constraints - I can’t have opus4.8 spend 4 hours on a deep research question I had in the shower without spending money. I can’t do real time video game upscaling and graphics work in the cloud period.

A laptop is about an order of magnitude cheaper than a cloud server thanks to economies of scale, uptime requirements, and other factors.



> The big question for local LLMs is whether there is a 100 tok/s model which requires less than 16 GB of memory and is competitive on most tasks with the cloud models.

Benchmarks maybe? Real world, no.

You just need the context otherwise. There's no way around it.


Context is more available locally. You can have the LLM operate for arbitrarily long periods, use your credentials to access services (if desired), store memory locally etc.

Whether such a model exists or not is a different question.


if you do the electricity math you'll see that you pay more on local models while getting less (local is more heavily quantized) compared with OpenRouter.

I'm not talking local Gemma/Qwen vs cloud Opus, but against OpenRouter same Gemma/Qwen

there are reasons to run local - privacy, availability, but cost is not one of them


I am allowed to plug in 800w of solar panels into a wall socket here in spain. That would more then cover my current computer with 16gb vram. Now if i went and built a LLM server, at full load i would probably be closer to 3600w (Dual Epyc CPUs that gives you 8 x16 PCI channels and up to 8 cards - Way overkill, i know). If i half that with 1 EPYC and 4 x16 PCI channels, and add the same amd 7800xt i currently have then i should in theory be able to run at around 1800w under full load. Now that could still be covered with a 2000w solar install (get a professional setup OR get a battery unit like a EcoFlow that can output 2000w and can input about the same amount of solar).

Now, this all brings the upfront costs way up, the solar panels are cheap, its all the rest around them that tends to cost money.


That's assuming consumption pricing remains as-is.

There has been a lot of market-subsidy in AI which is starting to fade away: e.g. the copilot quotas/pricing. When VC switches from investing to wanting a return, the price equation is likely to change.


There is no subsidy on most OpenRouter providers, they are profitable today.

You buy a big GPU, you serve LLMs, you print money.


And if you skip open router and go direct you save another 5%


> do the electricity math

Could you give an example with real figures?


obviously depends on your location and GPU

for me it would be about $2 per day in electricity to generate 8 mil tok of Gemma4-26B at 4 bit quantization. this is excluding how much the GPU cost (no amortization)

ignoring the fact that I could get more free tokens per day for this model from Google/OpenRouter, it would cost $4 per day on OpenRouter if paid, but they would run it at full 16 bit precission

this would be the most "profitable" model for me

for Gemma4-31B I can generate only 1 mil tok per day, and so I pay more to get less quality than OpenRouter (ignoring that this model is also free on Google)


This is very interesting to me. How do you come up with the kwh/token for your setup?


I know how many tokens per second my computer generates, and I have a wall power meter which measures how many watts my computer is using when generating.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: