What order or type of compression is used in Parquet?
Data Engineer
Besides API, what other methods and protocols have you worked with for data sources?
Have you written ingestion processes for OpenMetadata that scan sources?
Why was Parquet used instead of ORC?
How did you work with 1C: did you retrieve data directly from the database, through a bus, or by another method?
What happens when data quality constraints are triggered: alert, crash, restart?
Why are 'folders'/prefixes in S3 needed, besides convenience?
What technical fields were used in satellites, links, and hubs Data Vault?
Have you worked with FTP to obtain files or data through a portal?
Have you stored tabular view results in S3 as Parquet or otherwise?
Have you configured tags in S3 on buckets or objects?
How do you determine that a record already exists and does not need to be recorded again?
What file formats have you worked with?