What can you tell about SOAP API?
Data Engineer
How were DAGs launched: by cron, datasets, triggers, or dependencies?
What data quality checks were performed in Data Vault?
How did you connect to OpenMetadata and what did you do with it?
Give an example of a typical DAG task in Airflow.
How to track in which DAG run a specific row appeared in Data Vault?
Have you administered Airflow or only worked as a developer?
Can you use random distribution in Greenplum and when is it appropriate?
How did you optimize tables and queries: partitions, indexes, selection logic?
Tell me about yourself: what have you been doing, what stack did you use, and what kind of project was it?
What SQL queries do you write: simple checks or complex scripts for data warehouses?
Have you used Bridge and PIT tables in Data Vault? Provide an example and explain their purpose.
Did you do additional logging besides Airflow logs?
When can you give feedback and do you have offers on hand?.
What additional entities of Data Vault, besides hub, link, and satellite, have you used, such as same-as links or transactional links?
Was S3 the only primary storage or were multiple used?
List the SQL window functions you know.
Did you use one S3 bucket or multiple? Based on what principle?
Have you established lineage from source to consumer?
What type of storage does Parquet use: row-based or columnar?