|
|
|
|
|
by rldjbpin
28 days ago
|
|
no matter your luck with hardware or your sysadmin skills, doing local inference for just yourself and/or to emulate typical usage (e.g. your coding workflow and deep research, etc.) is just very inefficient in current model architecture. to me, this is a "truck" approach to city driving as a single person who does not do furniture hauling every weekend. the sense of privacy and freedom is nice but online inference is more "economical" as multi-user load is more effectively served than going solo. maybe new architectures would make it effective to do text inference locally [1], till then great on you if you can spend car money on your setup. hope it is a great learning experience as well. [1] https://deepmind.google/models/gemma/diffusiongemma/ |
|