Kubernetes Cluster Engineer - HPC/AI Platforms
Fluid Numerics
Job description
Overview
We are seeking a skilled Kubernetes Cluster Engineer to support the design, deployment, and ongoing operations of large-scale, GPU-dense compute environments where Kubernetes and Slurm run side by side. This role centers on cloud-native platform engineering — cluster lifecycle, GitOps-driven configuration, operators, and multi-tenant access control — applied to current-generation, GPU-dense bare metal rather than typical microservice workloads. You will own the Kubernetes layer of clusters that schedule AI/ML training and inference jobs, working closely with researchers, vendors, and partners. Slurm expertise is not a prerequisite, but you should be comfortable working next to it and interested in learning it: our environments increasingly bridge the two schedulers through Slinky, and that seam is where much of the interesting work lives.
You will work alongside our team to support in-house, partner, and customer infrastructure.
Why This Role
You will personally operate the newest silicon. NVIDIA B300 and AMD MI-series clusters, on bare metal.
Breadth instead of a single monolith. Multiple clients, multiple GPU vendors and generations, and greenfield bring-ups where you make the architectural calls rather than inherit them. You will not be the third Kubernetes engineer maintaining someone else's cluster.
A small team and real autonomy. Team oriented project management. You will own systems end to end, talk directly to the researchers and vendors who depend on them, and see the consequences of your choices quickly.
A path into HPC depth. Kubernetes fluency is common; Kubernetes fluency plus real Slurm and HPC scheduling expertise is scarce and getting scarcer as AI infrastructure converges on both. We will actively work with you to develop that second half, because we need it and because it makes you harder to replace anywhere you go next.
Responsibilities
Kubernetes Platform Engineering
-
Build and operate production Kubernetes clusters on bare metal, including control plane lifecycle, upgrades, etcd operations, and disaster recovery.
-
Deploy and maintain GPU and high-performance networking stacks: GPU Operator, device plugins, Node Feature Discovery, RDMA/SR-IOV, Multus, and topology-aware scheduling.
-
Manage cluster configuration as code through GitOps, Helm, and Kustomize, with promotion paths across dev, staging, and production clusters.
-
Implement multi-tenancy and governance: namespaces, RBAC, resource quotas, network policy, and admission control.
-
Integrate parallel and object storage into the cluster via CSI drivers, and tune for AI/ML data-loading patterns.
-
Evaluate and deploy batch and queueing layers on Kubernetes and Kubernetes/Slurm bridging via Slinky.
Cluster Engineering & Deployment
-
Participate in the design and bring-up of bare metal HPC/AI/ML environments.
-
Integrate heterogeneous hardware platforms (multiple GPU vendors and generations) into cohesive scheduling environments.
-
Develop provisioning and imaging workflows (Ansible, MAAS, cloud-init, Terraform, CI/CD pipelines) for reproducible cluster build-out.
-
Coordinate communications between vendors, researchers, and other partners during cluster bring-up and operation.
Working Alongside Slurm
-
Operate and help configure the Slurm Workload Manager in mixed Kubernetes/Slurm environments — growing into partition design, QoS, preemption, and GRES GPU scheduling over time.
-
Support identity, accounting, and health-checking integrations that must behave consistently across both schedulers.
-
Contribute to scripts and plugins (prolog/epilog, node health checks, job cleanup) that extend scheduler functionality.
System Administration & Observability
-
Administer Linux systems at the host level: network configuration, storage integration, kernel and driver tuning for GPU and RDMA workloads.
-
Deploy and maintain observability stacks (Prometheus, Grafana, Datadog, DCGM/ROCm exporters) covering cluster health, GPU utilization, and job-level metrics.
-
Automate failure detection, node draining and remediation, and job cleanup to protect uptime on expensive hardware.
-
Manage security and access control across layers (Kubernetes RBAC, OIDC/SSO, LDAP/SSSD, VPN, PAM, SSH session auditing).
User & Stakeholder Support
-
Help cluster users build workflows that make efficient use of compute resources.
-
Containerize AI/ML and HPC applications (Docker/Podman, Enroot-Pyxis) and integrate GPU-aware runtimes into both Kubernetes pods and Slurm jobs.
-
Automate cost accounting and cluster usage reporting.
Qualifications
Required
-
Substantial hands-on experience operating Kubernetes in production, including bare metal clusters and their networking and storage integration.
-
Fluency with the cloud-native toolchain: Helm, GitOps (ArgoCD/Flux), CRDs and operators, RBAC, and admission control.
-
Strong Linux system administration, networking, and performance-tuning skills.
-
Proficiency with infrastructure automation (Ansible, Terraform, CI/CD pipelines) and version control workflows.
-
Demonstrated ability to operate GPU-accelerated infrastructure at scale.
-
Familiarity with Slurm and HPC/AI batch scheduling concepts, and genuine interest in deepening it.
Preferred
-
Experience with GPU-specific Kubernetes tooling: NVIDIA GPU and Network Operators, AMD GPU Operator, MIG or partitioning strategies, Dynamic Resource Allocation.
-
Exposure to Slinky, slurm-bridge, SUNK, or other Kubernetes/Slurm integration patterns.
-
Experience with parallel file systems (WEKA, Lustre, GPFS, BeeGFS) and high-performance interconnects (InfiniBand, RoCE, 100/200/400 GbE).
-
Writing Kubernetes controllers or operators, in Go or otherwise.
-
Familiarity with common AI/ML software dependencies (CUDA/ROCm, NCCL/RCCL, PyTorch, vLLM) and researcher workflows.
-
Multi-site, hybrid cloud, or federated cluster operations.
Job Type: Full-time
Pay: From $140,000.00 per year
Benefits:
-
401(k)
-
401(k) matching
-
Health insurance
-
Relocation assistance
Experience:
- Linux and HPC cluster system administration: 1 year (Required)
Language:
- English (Required)
Work Location: Hybrid remote in Hickory, NC 28602