Sobes.tech
Senior

Prečo môže kvantizovaný LLM fungovať rýchlejšie a prečo niekedy kvantizácia naopak spomaľuje inferenciu?