Sobes.tech
Senior

Jaké parametry je třeba použít k hodnocení paměti pro inference LLM a jaké architektonické optimalizace snižují KV-cache?