Provide examples of critical incidents you responded to within 30 minutes
SRE / Platform Engineer
Write a command or script that finds all files with the extension .log in the current directory containing the word error, and outputs only the names of these files.
Describe the process of an HTTP request to google.com step by step (from curl to content retrieval)
Is it necessary to notify users about the incident?
What methods of inter-process communication (IPC) do you know?
Have you ever deployed PostgreSQL from scratch?
What Ansible playbook examples have you created?
Tell about a practical case where you had to worry about the result?
What are the most unusual or specific metrics you have had to display on dashboards?
Have you worked with Rancher?
How is the deployment schedule organized: is there a calendar and a regulation or are releases made outside the queue?
How do you relate your experience to the job position tasks? What experience is lacking and what would you like to develop?
What is SystemD and why is it popular? What initialization systems were used before?
How did you ensure the fault tolerance of the Kafka platform?
What is the purpose of Headless Service in Kubernetes?
Did you configure static or dynamic routes on ELTEX ESR?
How did you access the router when the tunnel failed — via external IP?
What are the disadvantages of containerization technology?
What is TSDB (Time Series Database)?
What does the permissions 644 mean for a file?