
Site Reliability Engineer
AKKA
Job description
Description
Akka is the platform for building and running AI agents and distributed systems at scale. Our actor-model runtime powers agent and workflow systems for companies like Manulife, John Deere, Capital One, and CERN. We also run Akka Automated Operations, our managed platform for customer workloads on dedicated, BYOC, and BYOK8s deployments, backed by extensive availability guarantees and certifications.
We're hiring staff-level SREs to help run and evolve that infrastructure, working alongside the senior engineers already on the team. You'll contribute to architecture decisions for how we deploy, observe, and secure the platform, and help shape the standards the team builds against. It's not a team that watches someone else's dashboards, and it's not a role with a long runway: we expect you to be contributing meaningfully within your first few weeks.
The core of the platform is a set of Kubernetes operators written in Go. Every customer service, route, datastore, secret, and region flows through a reconciler, and when something goes wrong in production the fix usually starts with an operator's logs and a custom resource's status. Customer clusters on AWS, Azure, and GCP are provisioned with Crossplane compositions and delivered with Flux; with Terraform at the core. Each region runs managed Postgres that both our control plane and customer workloads depend on. Everything we run is defined and changed through code.
What we're looking for
-
7+ years in infrastructure or platform engineering, with production experience across AWS and Azure. GCP is a plus.
-
Experience operating Kubernetes controllers built on controller-runtime in production: CRDs, admission webhooks, finalizers, status conditions, and debugging a reconcile loop that isn't converging. You can read Go well enough to trace a reconciler to a root cause.
-
Infrastructure as code with real production work, Terraform, and Crossplane at the level of authoring Compositions and XRDs rather than only applying claims, delivered through Flux and Kustomize.
-
Operating managed Postgres (RDS, Cloud SQL, or Azure Flexible Server) in production, point-in-time restore, major-version upgrades, and moving a live database between instances inside a bounded outage window.
-
Observability at scale with the Prometheus operator, scrape and relabel configuration, cardinality and cost control, alert rules as code, and distributed tracing with OpenTelemetry. Running a long-term metrics store (Cortex, Mimir, Thanos) is a plus, not a requirement.
-
Securing Kubernetes clusters, service mesh (Linkerd or similar), OIDC/workload identity, mTLS, cert-manager for PKI including trust-anchor rotation, and secrets management via cloud KMS.
-
Production on-call experience, you've carried a pager and written up what happened afterward.
-
Skilled use of LLMs as a tool to sharpen your work, not to run on autopilot.
-
Strong written communication, we weigh this heavily in our process.
Nice to have
-
Sizing JVM services in containers, heap versus container limits, direct memory, GC behavior, and reading a heap dump. Our platform computes JVM flags per service, and getting it wrong shows up as OOMKilled.
-
Operating event-sourced systems, projection lag, offsets, replay semantics, and what a journal replay does to a read model. Our own control plane is event-sourced, and so are our customers' workloads.
-
Messaging or streaming systems (Kafka, Pub/Sub, or similar) at production scale.
-
Teleport or a similar access plane, managed as code.
-
Distributed, stateful, or actor-based systems (Akka, Erlang/OTP).
Benefits
-
Competitive salary with performance-based incentives.
-
Comprehensive health and wellness benefits.
-
Opportunities for professional development and continuous learning.
-
Flexible remote working environment.
-
Collaborative, inclusive, and innovative company culture.
-
A transparent, distributed work environment with a strong focus on work-life balance.
-
Challenging work that interacts with innovative applications used by millions.
-
A collaborative culture that attracts the "brightest minds" in the technology community.
About Akka
Our Vision:
Distributed systems that feel local.
Our Mission:
To make it simple to build, run, and evaluate agentic systems.
Our Values:
-
We’re Authentic: We value transparency and genuine communication, without politics or games. We're honest and assume good intentions, cultivating trust and accountability within our organization and in our interactions with the community. Authentic defines who we want to work with.
-
We’re Customer-focused: We value customer outcomes above all else. By prioritizing our customers' interests, and meeting them where they are today, we help ensure their success. We are dedicated to deeply understanding our customer’s needs, anticipating challenges, navigating time constraints, and striving to exceed expectations. Customer-focused defines how we decide where we spend our time.
-
We’re Nonconventional: We value fearless innovation by challenging the status quo and embracing alternative approaches. Continuous learning and a growth mindset aimed at improving ourselves, our company, and our products, drives us to push boundaries and explore new solutions. Guided by a bias for action, we leverage industry and customer insights to inspire fresh ideas, enabling optimal future offerings. Nonconventional defines how we approach and test a strategy.
-
We’re Persistent: We value excellence through continuous experimentation and courageous problem-solving. We recognize that achieving success often demands approaching challenges with tenacity and taking calculated risks to achieve leading-edge solutions. Persistent defines how we want to be perceived by others.