Sobes.tech

SRE / Platform Engineer

The problem has been confirmed, what to do next?

341

Is the incident resolved, the service restored — is that all or is there something else to do?

339

What is the difference between Prometheus and VictoriaMetrics? How did you configure scrape from multiple clusters?

319

How can you see if there have been any changes, who rolled out what?

275

How does a package go from an external client to a pod in another namespace?

198

What controls the number of pods?

194

What points should be included in the post-mortem?

190

What is the difference between Deployment, StatefulSet, and DaemonSet?

190

How did you set up monitoring for a new service? How do you store dashboards?

189

How did you ensure the fault tolerance of the Kafka platform?

189

What is the fundamental technical difference between FluxCD and ArgoCD? Pull vs Push model?

185

Is it necessary to notify users about the incident?

185

Tell us about yourself and your experience.

184

What happens if you manually delete StatefulSet and DaemonSet pods?

178

How was the integration with Vault organized? How were secrets rotated?

177

What are readiness, liveness, and startup probes?

173

QoS classes in Kubernetes: Guaranteed, Burstable, BestEffort

168

How to fix if the service keeps an old secret in memory?

168

Did you create dashboards in Grafana yourself? How did you maintain them?

167
/2