Sobes.tech
Senior

What parameters should be used to evaluate memory for LLM inference and what architectural optimizations reduce KV-cache?