What is the HAVING operator in SQL? Where is it used?
Machine Learning / AI
Regression task: predicting stock prices in the 5th year based on 4 years of data using boosting. What mistake did we make?
What are MAPE and SMAPE metrics? Where might they be ineffective?
Tell us about the project for evaluating dispatcher performance: what models were used, how were they compared with the reference script?
Why choose CatBoost over XGBoost or LightGBM? What are the differences?
MVP task: binary classification with class imbalance, data distributed across 10 machines, with a text feature. What approach would you suggest?
Tell us about the project of clustering requests with annotation based on business importance.
What metrics are most prioritized when there is a large class imbalance (99 to 1)?
Are you familiar with HTTP, UDP, TLS protocols?
What is the main difference between Spark and Pandas? What is Hadoop?
Do you know what bagging is? Provide an example of an algorithm.
Что такое LoRA (Low-Rank Adaptation)?
Tell me more about the project on employee churn prediction: what did you do, what data did you use, why did you choose CatBoost?
Tell us about yourself, your experience, and your projects.
Tell us about the K-means algorithm and its advanced variants (K-means++).
What metrics are used to evaluate clustering quality? How to select the number of clusters?