What is ACID? What does each letter stand for? How does it relate to transactions?
Data Engineer
Do you know about the Avro format?
Tell about relational databases in general — what are they?
When do we not need data storage?
Have you had experience with Docker images, setting up environments on a virtual machine?
What is duck typing?
What are the limitations on using hashes (hash tables)?
How is QA (quality assurance) organized in your process?
You mentioned Docker processes, can you tell me more about how it's set up in your environment?
What is the difference between refactoring and optimization?
Have you had experience working with Table Engine Join in ClickHouse?
What is a query plan? Why are they needed?
What are the types of physical connections (join methods)?
Tell us about your experience working with distributed databases (Greenplum, ClickHouse)?
What do you currently dislike about your current job?
Imagine that Data Vault is unavailable and there are two wide tables that need to be joined. How would you optimize this in Greenplum? And if it were in ClickHouse?
Is it possible to come up with something else instead of a regular CTE, considering that filtering still occurs in memory on the entire table?
Do you have experience working with engines like Trino, or with table formats (Iceberg, Parquet) in Spark?
What was the purpose of Data Vault for you in general?
In what environment was Spark run?