What does BERT add to the transformer architecture?
Machine Learning / AI
What is positional encoding in a transformer?
How do convolutional networks work with texts?
What is DenseNet and how does it differ from ResNet?
How does the Viterbi algorithm work?
What resources on models and articles do you monitor?
How to parse documents for RAG?
What is the exploration-exploitation tradeoff and how is it implemented (Thompson sampling, epsilon-greedy)?
What is semi-supervised learning and self-training?
What classification metrics should be used in case of severe class imbalance?
What features do voice dialogues have (latency, ASR errors)?
Как найти треугольник на изображении без deep learning?
How to measure the distance between two embeddings?
What is knowledge graph embedding (TransE, DistMult, ComplEx, RotatE)?
What is the difference between a smart pointer and a regular pointer in C++?
How is a transformer structured and what are its major blocks?
What is NGBoost and probabilistic predictions?