
Senior Staff Software Engineer (Core Infrastructure)
Robinhood
Job description
-
As a Senior Staff Software Engineer on Core Infrastructure, you will own the architectural evolution of three deeply interconnected domains: service mesh and connectivity, compute platform, and infrastructure provisioning
-
Your decisions will directly shape engineering velocity, operational reliability, and Robinhood’s ability to expand to new regions and markets. This is not an operations role — it’s a once-in-a-platform-lifecycle opportunity to redesign the foundation before complexity becomes permanent!
-
Own service mesh architecture end-to-end — standardize and evolve Robinhood’s Istio and Envoy deployment, eliminate deprecated communication patterns, and build the networking foundation for multi-region operation
-
Drive Kubernetes compute platform evolution — define the long-term direction for cluster lifecycle, scheduling strategy, control plane architecture, and multi-region resiliency across Robinhood’s entire compute fleet
-
Redesign infrastructure provisioning — transform Terraform from a collection of ad-hoc configurations into a scalable, self-service platform capability that enables rapid region provisioning and developer autonomy
-
Operate with source-level depth — diagnose and resolve issues at the Kubernetes API machinery, etcd, Envoy xDS, and Terraform provider layer, and contribute upstream where it serves the platform
-
Define technical direction across teams — produce architecture proposals, lead design reviews, and drive cross-functional alignment on infrastructure decisions that affect every engineering organization at Robinhood- Deep, hands-on expertise with Kubernetes internals — API server, controller reconciliation, etcd watch propagation, scheduler mechanics, and admission control
-
Deep working knowledge of Istio and/or Envoy at the source code level — including the xDS API as a protocol; experience debugging at the control-plane/data-plane boundary
-
Terraform at architectural scale — module design patterns, provider internals, remote state management, and multi-region provisioning
-
Proven ability to drive cross-functional architectural change across large or complex engineering organizations — influencing without authority and shipping durable decisions
-
Production Go or Python for infrastructure tooling; strong fundamentals in writing well-structured, tested, operationally sound code
-
HTTP/2 and gRPC wire-level knowledge — connection management, flow control, stream multiplexing, and observability