
Senior Platform Engineer (Compute Services)
CoreWeave
Job description
-
We are seeking a Senior Platform Engineer to join our Kubernetes Infrastructure team. This role involves administering our critical multi-tenant Kubernetes platforms and collaborating with development teams to establish proper deployment architectures. The ideal candidate will have a strong background in resilient kubernetes application architecture and deployment
-
Champion reliability initiatives for Kubernetes application deployments: Advocate for best practices to ensure high availability, scalability, and resilience of applications in Kubernetes, focusing on robust testing, secure pipelines, and efficient resource use
-
Administer multi-tenant Kubernetes platforms: Manage complex multi-tenant Kubernetes clusters, configuring access, quotas, and security for isolation and optimal resource allocation while upholding SLAs
-
Perform lifecycle and day 2 operations on clusters: Execute Kubernetes cluster lifecycle, including provisioning, patching, monitoring, backup, disaster recovery, and troubleshooting
-
Deep dive into reliability issues: Conduct in-depth analysis and root cause identification for complex reliability incidents in Kubernetes, utilizing advanced debugging and monitoring tools to propose preventative measures
-
Perform on-call duties: Respond to critical alerts and incidents outside business hours, providing timely resolution to minimize disruptions, collaborating with teams, and communicating clearly- Strong communication and collaboration
-
Proficient in Go
-
Strong Gitops/Devops with Argocd or similar helm chart management
-
CKA or similar certifications is highly desired
-
Excellent problem-solving, debugging, and analytical skills
-
Proven Docker and containerization experience
-
Strong Linux OS experience
-
Bachelor’s in CS, Engineering, or related field, or equivalent experience preferred
-
2+ years administering multi-tenant SAAS Kubernetes (EKS, AKS, GKS)
-
Master’s degree in Computer Science, Engineering, or a related field
-
Knowledge of network protocols and distributed consensus algorithms
-
Experience with performance profiling and optimization of distributed systems