How to technically implement a data storage and alert system for model monitoring?
Machine Learning / AI
How are random forest and gradient boosting similar and different? How are trees built in each?
Do you have ideas for further improving the demand forecasting system?
Tell how the clustering algorithm you used worked — how did it aggregate points in space?
Case: development of a pricing system for S7 business class. 300 flights a day, ~100 with business class, 4-6 seats. Two goals: maximum occupancy and maximum revenue. Prices updated daily and after each purchase. Currently, prices are set manually, based on occupancy. How would you approach this task?
How did you determine the quality of clustering? What methods did you use?
Regarding Spark — what specific tasks were parallelized on it?
What is a stationary time series, how do you check it, and why is it important?
What is the difference between a confidence interval and a prediction interval? What do they show?
How to build a system for tracking quality and decision-making on model replacement?
What models did you use for demand forecasting?
How to deal with low feedback volume on quiet routes (few flights, noisy data)?
Tell us about your experience in Data Science
Tell about interesting optimization cases in Spark related to features and complex joins
How to solve the problem of shifted historical data (prices were undervalued, the model poorly sees demand at new levels)?