Sobes.tech
Senior

Koje parametre treba koristiti za procenu memorije za inference LLM i koje arhitektonske optimizacije smanjuju KV-cache?