Design and deploy production ML systems. Experience with PyTorch, distributed training, and MLOps required.
பொறுப்புகள்
- Architect and scale distributed training pipelines for large language models.
- Optimize model inference for latency and throughput in production environments.
- Collaborate with researchers to transition experimental models into robust products.
- Establish MLOps best practices, monitoring, and CI/CD pipelines for ML models.
தேவைகள்
- 5+ years of software engineering experience with a focus on machine learning.
- Deep expertise in PyTorch, CUDA, and distributed training frameworks (DeepSpeed, Megatron).
- Strong proficiency in Python, C++, and systems programming.
- Experience deploying models at scale using Kubernetes and Triton Inference Server.