In which cases is semantic search insufficient for good RAG and what should it be supplemented with?
Machine Learning / AI
What are audio verdicts in VKontakte advertising moderation?
Can the training of gradient boosting be parallelized?
What is forced alignment and have you used it?
How to diagnose and fix speech rate and interval differences between phrases?
What is the difference between WAV and PCM?
If speech rate and intervals between phrases differ, how to diagnose and fix?
Do you have any additional questions for the interviewer?
What to do if there are no transcripts for audio segments?
What is G2P?
After fine-tuning, the model started swallowing words. How to diagnose and fix?
What to do after VAD segmentation?
Why are weak models usually used in gradient boosting?
Is there experience with distributed training and what exactly?
How to ensure reproducibility of ML experiments and the ability to restore the model if the artifact was deleted in S3?
Can boosting be built on linear models?
What needs to be saved in the checkpoint to continue training after 47 hours without losing state?
Какая асимптотическая сложность у self-attention и что такое KV-cache на инференсе?
What is normalized text and how to normalize regular text for TTS?
How to annotate unannotated audio if different ASR normalize text differently and add punctuation?