D

AI Infrastructure Engineer

dicedemo

On-siteBoston, CTmid$130kPosted 7h ago

Job description

Job Description

AI Infrastructure Engineer

Position Overview

We are seeking an AI Infrastructure Engineer to design, build, and scale the infrastructure that powers our artificial intelligence and machine learning workloads. This role sits at the intersection of AI/ML, cloud infrastructure, distributed systems, and DevOps/MLOps.

The ideal candidate has experience building highly available, scalable infrastructure for training, deploying, and operating machine learning and generative AI applications. You will partner closely with Machine Learning Engineers, Data Scientists, Software Engineers, and Platform Engineering teams to ensure AI workloads can run reliably, securely, and efficiently at scale.

Key Responsibilities

Design, build, and maintain scalable infrastructure for AI, machine learning, and Generative AI workloads

Build and manage cloud infrastructure across AWS, Azure, and/or Google Cloud Platform

Deploy and operate GPU-based compute environments for model training and inference

Design infrastructure supporting LLMs, model training, fine-tuning, inference, and AI applications

Build and manage containerized workloads using Docker and Kubernetes

Develop infrastructure-as-code using tools such as Terraform, CloudFormation, or Pulumi

Build CI/CD and MLOps pipelines supporting model development and deployment

Optimize GPU/CPU utilization, infrastructure performance, scalability, and cloud costs

Implement monitoring, logging, observability, and alerting for AI infrastructure and services

Support distributed training and high-performance computing environments

Build secure, highly available systems capable of supporting production AI workloads

Partner with ML Engineers and Data Scientists to move models from experimentation into production

Troubleshoot infrastructure, networking, performance, and deployment issues

Evaluate emerging AI infrastructure technologies and recommend improvements to the platform

Required Qualifications

3+ years of experience in Cloud Infrastructure, DevOps, Platform Engineering, SRE, MLOps, or AI/ML Infrastructure

Strong experience with at least one major cloud platform: AWS, Azure, or GCP

Experience with Kubernetes and Docker

Experience with Infrastructure-as-Code tools such as Terraform

Strong scripting/programming skills in Python, Bash, Go, or similar languages

Experience building CI/CD pipelines using tools such as GitHub Actions, GitLab CI, Jenkins, or similar

Knowledge of networking, Linux systems, distributed computing, and cloud architecture

Experience implementing monitoring and observability solutions

Understanding of machine learning development and deployment workflows

Preferred Qualifications

Experience managing GPU infrastructure, including NVIDIA GPUs and CUDA environments

Experience with AI/ML frameworks such as PyTorch, TensorFlow, JAX, or Hugging Face

Experience supporting LLM training, fine-tuning, RAG, or inference workloads

Experience with MLOps platforms such as MLflow, Kubeflow, SageMaker, Vertex AI, or Azure Machine Learning

Experience with distributed training technologies such as Ray, DeepSpeed, PyTorch Distributed, or Horovod

Familiarity with AI inference technologies such as vLLM, NVIDIA Triton, or TensorRT

Experience managing Kubernetes-based GPU clusters

Understanding of model serving, vector databases, and modern Generative AI architecture

Experience optimizing infrastructure for performance and cloud/GPU cost efficiency

What Success Looks Like

In this role, you will help create the infrastructure foundation that allows AI teams to experiment faster, train models efficiently, deploy AI applications reliably, and scale them into production. You will reduce friction between AI development and production while improving reliability, performance, security, and infrastructure cost.

,

Required Skills

3+ years of experience in Cloud Infrastructure, DevOps, Platform Engineering, SRE, MLOps, or AI/ML Infrastructure

Strong experience with at least one major cloud platform: AWS, Azure, or GCP

Experience with Kubernetes and Docker

Experience with Infrastructure-as-Code tools such as Terraform

Strong scripting/programming skills in Python, Bash, Go, or similar languages

Experience building CI/CD pipelines using tools such as GitHub Actions, GitLab CI, Jenkins, or similar

Knowledge of networking, Linux systems, distributed computing, and cloud architecture

Experience implementing monitoring and observability solutions

Understanding of machine learning development and deployment workflows

,

Desired Skills

Experience managing GPU infrastructure, including NVIDIA GPUs and CUDA environments

Experience with AI/ML frameworks such as PyTorch, TensorFlow, JAX, or Hugging Face

Experience supporting LLM training, fine-tuning, RAG, or inference workloads

Experience with MLOps platforms such as MLflow, Kubeflow, SageMaker, Vertex AI, or Azure Machine Learning

Experience with distributed training technologies such as Ray, DeepSpeed, PyTorch Distributed, or Horovod

Familiarity with AI inference technologies such as vLLM, NVIDIA Triton, or TensorRT

Experience managing Kubernetes-based GPU clusters

Understanding of model serving, vector databases, and modern Generative AI architecture

Experience optimizing infrastructure for performance and cloud/GPU cost efficiency

,

About dicedemo

New Company