MetAntz logo

Senior Staff Security Engineer

MetAntz

On-sitePalo Alto, CAseniorPosted 5h ago

Job description

What You’ll Do (Roles & Responsibilities)

  • Architecture:

    Own the end-to-end security architecture for multi-tenant Kubernetes GPU clusters.

  • Isolation & Segmentation:

    Design tenant isolation, egress control, and network segmentation across compute, storage, and networking layers.

  • Workload Security:

    Define and implement runtime security and intrusion detection for untrusted AI workloads.

  • Security Primitives:

    Build security primitives (identity, secrets, encryption, policy enforcement) that platform teams build on.

  • Supply Chain Security:

    Secure the software supply chain, from CI/CD pipelines to container admission.

  • Threat Modeling:

    Lead threat modeling and security design reviews for new platform features.

  • Compliance:

    Drive compliance readiness (SOC 2, ISO 27001) without slowing engineering velocity.

  • Culture & Leadership:

    Act as a force multiplier – unblock teams, set standards, and raise the security bar across the org.

  • Incident Response:

    Lead security incident response and turn incidents into systemic improvements.

What We’re Looking For (Required Qualifications)

  • Experience:

    7+ years of experience in security engineering for cloud-native, infrastructure, or distributed systems.

  • Kubernetes Security:

    Deep, hands-on expertise in Kubernetes security (RBAC, PSA, network policies, admission controllers).

  • Cloud Security:

    Strong understanding of cloud security primitives in AWS, GCP, or Azure.

  • Runtime & Policy Tools:

    Experience building or operating runtime security and policy enforcement (Falco, Cilium, OPA, Calico, eBPF-based tools).

  • Networking:

    Solid grounding in network security and zero-trust architectures.

  • Secrets Management:

    Practical experience with secrets management and key systems (Vault, cloud KMS).

  • Automation & Scripting:

    Strong automation skills in Go, Python, or Bash.

  • Execution:

    Proven ability to operate autonomously, make architectural decisions, and deliver in ambiguous environments.

  • Collaboration:

    Experience partnering deeply with platform, infra, and SRE teams.

Bonus Experience (Nice to Have)

  • GPU isolation and virtualization security (MIG, SR-IOV)

  • InfiniBand, RDMA, or high-performance networking

  • HPC or large-scale multi-tenant compute platforms

  • Security for AI/ML systems or data-intensive workloads

  • Incident response leadership or red team experience

  • Security certifications (CKS, OSCP, CISSP)

  • Open-source contributions in security or cloud infrastructure

What Makes This Role Different

  • You’ll define the security model, not just implement tickets.

  • You’ll work on real isolation problems in GPU and high-speed networking environments.

  • You’ll influence architecture across the platform, not sit in a silo.

  • You’ll help build a security culture at a company where speed and safety both matter.