Responsibilities:
• Design, build, and maintain backend services for deploying and serving ML models in production.
• Develop high-performance APIs and services in Java, with supporting components in Python.
• Deploy and operate services on AWS using ECS + Fargate, SageMaker, or EC2 with Kubernetes.
• Enable both real-time (low-latency) and batch inference workflows.
• Ensure service reliability, scalability, and performance in production environments.
• Implement and maintain CI/CD pipelines and deployment automation.
• Monitor systems using observability tools and proactively resolve production issues.
• Collaborate with ML engineers to optimize model serving and integration workflows.
Requirements:
• Strong backend engineering fundamentals with proficiency in Java.
• Working knowledge of Python, especially in ML-related workflows.
• Hands-on experience with AWS services (ECS, EC2 DynamoDB, Redis, S3 SageMaker).
• Experience with containers and orchestration (Docker, ECS, Kubernetes).
• Familiarity with Terraform or other infrastructure-as-code tools.
• Experience building and operating production systems with CI/CD and monitoring.