
MLOps Engineer
Evlo AI
Job description
About The Role
The role owns the infrastructure, orchestration, and deployment pipelines that power large-scale machine learning and generative AI systems in production.
The team works closely with machine learning engineers and data scientists to ensure models scale reliably, maintain low latency, and remain observable under heavy enterprise workloads.
Key Responsibilities
-
Design and implement scalable MLOps infrastructure using Kubernetes, Docker, Terraform, and modern CI/CD pipelines
-
Build automated model training, validation, and deployment pipelines to streamline the path from research to production
-
Deploy and manage model serving endpoints using Triton, TorchServe, vLLM, or AWS SageMaker for low-latency inference
-
Configure automated monitoring systems for data drift, concept drift, system latency, and infrastructure resource utilization
-
Implement comprehensive logging, tracing, and observability frameworks for complex LLM and traditional ML workflows
-
Collaborate with security and engineering teams to ensure compliance, model governance, and robust access controls
What We Are Looking For
-
3โ7 years of experience in MLOps, DevOps, or machine learning infrastructure engineering
-
Strong proficiency in Python, containerization technologies (Docker), and orchestration platforms (Kubernetes)
-
Hands-on experience with cloud infrastructure providers such as AWS, GCP, or Azure
-
Familiarity with model serving frameworks, feature stores (Feast, Hopsworks), and vector databases
-
Solid understanding of CI/CD principles and infrastructure-as-code tools like Terraform or Ansible
-
Bonus: Experience managing LLM inference optimization, vLLM, Triton, or large-scale distributed training clusters