What is PPO and how is it more convenient than TRPO in practice?
Machine Learning / AI
What is RT-DETR and what is its idea?
What are A2C and A3C?
What is RankNet and its loss?
What is self-supervised pretraining for GNN?
What approaches are there to responsible scaling of LLM?
What are temporal knowledge graphs?
What is BLEU and in what tasks is it used? What are the problems with this metric?
What are recovery prompts in low-confidence?
What is DenseNet and how does it differ from ResNet?
What cost-aware HPO approaches exist?
What is the RAG evaluation framework (Ragas, TruLens)?
How to cache the most frequent queries in RAG?
What is coalesced memory access?
Why is bias=False used in Linear layers in modern LLMs?
How does sparse retrieval (BM25) differ from dense?
What is OCR-free document understanding (Donut, Pix2Struct)?
What are the features of product ranking in e-commerce (CTR, conversion, revenue)?
What is Score-based SDE and how are they related to DDPM?
What is diffusion distillation?