Sobes.tech

Data Engineer

Data Vault is a convenient thing, but there are quite a few objects. Can you explain the difference between hubs, links, and satellites?

189

What processes are triggered when deleting or changing records in ClickHouse?

187

You work with Data Science and big data processing — is there a replacement in this library written in Rust, multithreaded, unlike Pandas, very fast. Do you know such?

184

How do you break down a query into stages? What do we do for that?

184

Parquet, you uploaded data — what technologies were used to receive and upload the data, and what mechanisms were employed?

182

With dependencies between jobs and DAGs — did you use sensors to run several jobs one after another?

181

You uploaded data from Spark to S3 in Parquet format, and then from Parquet to Greenplum — how does the loading process occur?

181

Have you worked with Git? What did you store in Git?

178

Did you usually design the database schema manually in the database, or did you store scripts somewhere, or did you use Liquibase at all, what was the process?

177

How was data from an external table loaded into target tables?

176

Did you write Airflow DAGs and jobs manually or using templates?

176

Have you heard of Iceberg working with streaming? It's a trashy file storage.

175

If CTE is used only once in a query, does it make sense to use it?

171

This is real-time, how to implement some kind of connection between Kafka and ClickHouse Spark?

170

Have you worked directly with Spark? What is its disadvantage?

169

We will parse JSON — who will parse JSON? As far as I know, ClickHouse does not have functions for this.

166

We need a pipeline engine — a mechanism that allows very fast stream creation by providing contracts and parameters as input. How would you implement this at a high level?

165

Iceberg is an engine that allows you to turn S3 into a database, and it even complies with ACID. What is its downside?

163
/3