Sobes.tech

Machine Learning / AI

What is gradient accumulation and why is it needed on limited memory?

190

What is Pix2Struct?

Middle — Senior
178

What is FlashAttention and what memory gains does it provide?

178

What is Donut and why is OCR-free approach used?

Middle — Senior
176

What is a router LLM and why are requests directed to different models?

169

What is weight-only quantization and activation quantization?

Senior
165

What is expert parallelism in MoE models?

Senior
153
/2