Have you used XCom in Airflow?
Data Engineer
Tell about partitioning in Oracle: why it is used and what types exist?
What about CSV family formats — when can they be used?
Is it possible to view the table schema before executing a query in Hive?
What are views (VIEW) and materialized views in Oracle? How is data updated in a materialized view?
Does the execution plan from EXPLAIN always match what is actually executed?
Why are views and tables separated? What are views used for?
What are lazy operations in Spark and what are they used for?
What other file formats, besides Parquet, work well with Hadoop and Hive?
How to find out the actual execution plan of a query in Oracle?
Tell me about your previous project, what were your responsibilities?
Have you worked with Oracle Data Integrator (ODI)?
How does your experience match the responsibilities in the vacancy: supporting and analyzing loading processes, monitoring and identifying anomalies, testing and deploying improvements in production, second-line support, checking technical documentation?
What are the advantages and disadvantages of distributed database systems?
There was a slow analytical query in Hive. How do you understand what happened to it and how to fix it?
What is the difference between multithreading and multiprocessing? For which tasks should you use threads, and for which processes?
Have you worked with batch loading or streaming?
Can you recall your experience with query optimization? How did you act, analyze?
What architecture did you have on the project? DWH or something else?
Have you worked with window functions? What types of window functions do you know?