Build and scale our AI infrastructure platform. Kubernetes, cloud-native architecture, and CI/CD expertise needed.
Responsibilities
- Design and maintain a scalable, multi-region Kubernetes infrastructure for AI workloads.
- Develop internal tooling to accelerate ML engineering workflows.
- Ensure high availability, security, and performance of our core platform.
- Manage cloud resources and optimize GPU utilization.
Requirements
- 4+ years of experience in Infrastructure, DevOps, or Platform Engineering.
- Extensive experience with Kubernetes, Terraform, and AWS/GCP.
- Strong programming skills in Go or Python.
- Experience with GPU infrastructure is a massive plus.