Swiss Nexus logo

Senior GPU Infrastructure Engineer

Swiss Nexus

On-siteSan Francisco, CAsenior$180k–$250kPosted 6h ago

Job description

Swiss Nexus is partnering with a fast-growing technology company building next-generation infrastructure for large-scale AI workloads.

The team operates at the intersection of

distributed systems, GPU infrastructure and high-performance computing

, developing infrastructure designed to make large-scale compute more accessible, reliable and efficient.

As the platform continues to scale, we are looking for a

Senior GPU Infrastructure Engineer

to help build and optimize the systems powering demanding AI and machine-learning workloads.

The Role

You will work on the infrastructure layer responsible for orchestrating and operating GPU compute at scale.

This is a highly technical, hands-on engineering position for someone comfortable working across

GPU systems, distributed infrastructure, Kubernetes and performance optimization

.

You will have significant ownership over architecture and technical decisions and work closely with a small team of experienced engineers.

What You'll Do

  • Design, build and operate infrastructure supporting large-scale GPU workloads

  • Improve the reliability, scalability and performance of distributed compute systems

  • Build and optimize GPU orchestration and scheduling infrastructure

  • Work extensively with

    Kubernetes and containerized environments

  • Diagnose performance bottlenecks across compute, networking and storage

  • Improve observability, monitoring and infrastructure automation

  • Design systems capable of operating reliably across heterogeneous GPU environments

  • Collaborate with engineering teams working on AI/ML infrastructure and distributed systems

  • Contribute to architectural decisions as the platform scales

What We're Looking For

  • Strong experience in

    infrastructure, distributed systems, SRE, platform engineering or systems engineering

  • Deep practical knowledge of

    Kubernetes

  • Experience operating

    GPU infrastructure or large-scale compute workloads

  • Strong understanding of Linux systems, networking and containerization

  • Experience building highly available production infrastructure

  • Strong programming or scripting skills, ideally with

    Go, Python or similar languages

  • Comfortable debugging complex performance and reliability issues across multiple layers of a system

  • Ability to operate independently in a fast-moving engineering environment

Particularly Relevant Experience

We'd be especially interested in candidates who have worked with:

  • GPU clusters and accelerated computing

  • AI/ML infrastructure

  • Kubernetes scheduling and orchestration

  • Distributed compute platforms

  • High-performance networking

  • Cloud infrastructure at significant scale

  • Bare-metal or heterogeneous compute environments

  • Infrastructure supporting model training or inference

Experience within AI infrastructure, cloud compute, HPC or technically complex distributed platforms would be particularly relevant.

About Swiss Nexus

Swiss Nexus is a specialized recruitment and talent advisory firm connecting exceptional talent with companies building the next generation of technology.

With Swiss roots and a global network, we recruit across engineering, product, commercial and leadership functions for companies operating in blockchain, digital assets, fintech and emerging technology.