Over de vacature
The role focuses on supporting GPU/TPU infrastructure, improving reliability and efficiency, and setting up observability for machine learning systems.
What to do
- Support GPU/TPU clusters
- Work with containerization and orchestration using Docker and Kubernetes
- Audit and optimize infrastructure and costs
- Set up monitoring and alerts