Az állásról
The role focuses on supporting and optimizing infrastructure for machine learning workloads, including GPU/TPU clusters, containerized services, monitoring, and cost efficiency.
What to do
- Support GPU/TPU clusters
- Handle containerization and orchestration with Docker and Kubernetes
- Audit and optimize infrastructure and costs
- Set up monitoring and alerts