Hacker News new | ask | show | jobs
by rldjbpin 28 days ago
no matter your luck with hardware or your sysadmin skills, doing local inference for just yourself and/or to emulate typical usage (e.g. your coding workflow and deep research, etc.) is just very inefficient in current model architecture.

to me, this is a "truck" approach to city driving as a single person who does not do furniture hauling every weekend. the sense of privacy and freedom is nice but online inference is more "economical" as multi-user load is more effectively served than going solo.

maybe new architectures would make it effective to do text inference locally [1], till then great on you if you can spend car money on your setup. hope it is a great learning experience as well.

[1] https://deepmind.google/models/gemma/diffusiongemma/