AI Infrastructure Engineer
dicedemo
Job description
Job Description
AI Infrastructure Engineer
Position Overview
We are seeking an AI Infrastructure Engineer to design, build, and scale the infrastructure that powers our artificial intelligence and machine learning workloads. This role sits at the intersection of AI/ML, cloud infrastructure, distributed systems, and DevOps/MLOps.
The ideal candidate has experience building highly available, scalable infrastructure for training, deploying, and operating machine learning and generative AI applications. You will partner closely with Machine Learning Engineers, Data Scientists, Software Engineers, and Platform Engineering teams to ensure AI workloads can run reliably, securely, and efficiently at scale.
Key Responsibilities
Design, build, and maintain scalable infrastructure for AI, machine learning, and Generative AI workloads
Build and manage cloud infrastructure across AWS, Azure, and/or Google Cloud Platform
Deploy and operate GPU-based compute environments for model training and inference
Design infrastructure supporting LLMs, model training, fine-tuning, inference, and AI applications
Build and manage containerized workloads using Docker and Kubernetes
Develop infrastructure-as-code using tools such as Terraform, CloudFormation, or Pulumi
Build CI/CD and MLOps pipelines supporting model development and deployment
Optimize GPU/CPU utilization, infrastructure performance, scalability, and cloud costs
Implement monitoring, logging, observability, and alerting for AI infrastructure and services
Support distributed training and high-performance computing environments
Build secure, highly available systems capable of supporting production AI workloads
Partner with ML Engineers and Data Scientists to move models from experimentation into production
Troubleshoot infrastructure, networking, performance, and deployment issues
Evaluate emerging AI infrastructure technologies and recommend improvements to the platform
Required Qualifications
3+ years of experience in Cloud Infrastructure, DevOps, Platform Engineering, SRE, MLOps, or AI/ML Infrastructure
Strong experience with at least one major cloud platform: AWS, Azure, or GCP
Experience with Kubernetes and Docker
Experience with Infrastructure-as-Code tools such as Terraform
Strong scripting/programming skills in Python, Bash, Go, or similar languages
Experience building CI/CD pipelines using tools such as GitHub Actions, GitLab CI, Jenkins, or similar
Knowledge of networking, Linux systems, distributed computing, and cloud architecture
Experience implementing monitoring and observability solutions
Understanding of machine learning development and deployment workflows
Preferred Qualifications
Experience managing GPU infrastructure, including NVIDIA GPUs and CUDA environments
Experience with AI/ML frameworks such as PyTorch, TensorFlow, JAX, or Hugging Face
Experience supporting LLM training, fine-tuning, RAG, or inference workloads
Experience with MLOps platforms such as MLflow, Kubeflow, SageMaker, Vertex AI, or Azure Machine Learning
Experience with distributed training technologies such as Ray, DeepSpeed, PyTorch Distributed, or Horovod
Familiarity with AI inference technologies such as vLLM, NVIDIA Triton, or TensorRT
Experience managing Kubernetes-based GPU clusters
Understanding of model serving, vector databases, and modern Generative AI architecture
Experience optimizing infrastructure for performance and cloud/GPU cost efficiency
What Success Looks Like
In this role, you will help create the infrastructure foundation that allows AI teams to experiment faster, train models efficiently, deploy AI applications reliably, and scale them into production. You will reduce friction between AI development and production while improving reliability, performance, security, and infrastructure cost.
,
Required Skills
3+ years of experience in Cloud Infrastructure, DevOps, Platform Engineering, SRE, MLOps, or AI/ML Infrastructure
Strong experience with at least one major cloud platform: AWS, Azure, or GCP
Experience with Kubernetes and Docker
Experience with Infrastructure-as-Code tools such as Terraform
Strong scripting/programming skills in Python, Bash, Go, or similar languages
Experience building CI/CD pipelines using tools such as GitHub Actions, GitLab CI, Jenkins, or similar
Knowledge of networking, Linux systems, distributed computing, and cloud architecture
Experience implementing monitoring and observability solutions
Understanding of machine learning development and deployment workflows
,
Desired Skills
Experience managing GPU infrastructure, including NVIDIA GPUs and CUDA environments
Experience with AI/ML frameworks such as PyTorch, TensorFlow, JAX, or Hugging Face
Experience supporting LLM training, fine-tuning, RAG, or inference workloads
Experience with MLOps platforms such as MLflow, Kubeflow, SageMaker, Vertex AI, or Azure Machine Learning
Experience with distributed training technologies such as Ray, DeepSpeed, PyTorch Distributed, or Horovod
Familiarity with AI inference technologies such as vLLM, NVIDIA Triton, or TensorRT
Experience managing Kubernetes-based GPU clusters
Understanding of model serving, vector databases, and modern Generative AI architecture
Experience optimizing infrastructure for performance and cloud/GPU cost efficiency
,
About dicedemo
New Company