
Site Reliability Engineer (GCP)
Veson Nautical
Job description
-
As a Site Reliability Engineer at Veson Nautical, you will design, build, monitor, and support the cloud infrastructure that underpins our rapidly growing SaaS platform
-
His is a multi-cloud role spanning both AWS and Google Cloud Platform, with an immediate focus on growing our GCP footprint and the systems that connect our environments across regions, accounts, and clouds
-
The work is a mix of greenfield and stewardship. Youâll stand up infrastructure for new applications from scratch, and youâll take on the harder problem of making our existing estate more consistent, more observable, and easier to operate
-
Youâll join a global Site Reliability Engineering team with members in the United States and the United Kingdom
-
Our Stack:
-
Google Cloud Platform - primarily PaaS services (Bigtable, Cloud SQL, Dataflow, Datastore, GKE, GCS, KMS, Pub/Sub)
-
Amazon Web Services - multi-region, multi-account, with a broad range of managed services
-
Containers and orchestration - Kubernetes (GKE and EKS)
-
Infrastructure-as-Code - Terraform, Terragrunt, and Atlantis
-
CI/CD - GitLab Pipelines, ArgoCD, Octopus Deploy
-
Data - ElasticSearch hosted with Kubernetes Operator, PostgreSQL, SQL Server, BigQuery
-
Monitoring and Security - Splunk, Grafana / Grafana Tempo, OpenTelemetry, Cloud Armor Enterprise, OpsGenie, Renovate, Sentry
-
AI Tools - Claude, Amazon Bedrock, Gemini, Vertex AI
-
Design, implement, and operate scalable, reliable, and secure infrastructure across Google Cloud Platform and AWS
-
Lead greenfield infrastructure builds for new applications, and modernize existing infrastructure toward common patterns
-
Solve cross-region, cross-account, and cross-cloud problems â networking, identity, data movement, and the operational patterns that hold across environments
-
Drive automation of infrastructure provisioning and configuration management using Terraform and related IaC tooling
-
Establish and maintain comprehensive monitoring, alerting, and observability practices
-
Cross-train and mentor other engineers, with the explicit goal of broadening GCP and multi-cloud capability across the team
-
Partner closely with development teams to ensure the reliability, performance, and scalability of our platforms
-
Set technical direction through design reviews, architecture proposals, and clear written documentation
-
Participate in and improve incident response, and drive the follow-through that keeps the same incident from recurring
-
Improve the cost effectiveness of our cloud footprint through visibility, analysis, and sound architectural choices- 3+ years of hands-on experience with Google Cloud Platform services and architecture, including running production workloads at scale
-
Bachelorâs degree in Computer Science, Engineering, or a related field, or equivalent practical experience
-
Experience with Google Cloud networking, including VPC, Cloud Load Balancing, and Cloud CDN
-
Previous experience working on a large-scale Software-as-a-Service (SaaS) platform supporting thousands of global users in a 24x7x365 environment
-
Proficiency with Infrastructure-as-Code, preferably Terraform
-
Production Kubernetes experience, particularly Google Kubernetes Engine (GKE)
-
We are focused on building a diverse and inclusive workforce. If youâre excited about this role, but do not meet 100% of the qualifications listed above, we encourage you to apply
-
Strong programming skills in Python, Go, or TypeScript for automation and tool development
-
Google Cloud Professional certifications (Cloud Architect, Data Engineer, or DevOps Engineer)
-
Experience with GitLab CI or Octopus Deploy
-
Experience working on a geographically distributed team across time zones
-
Experience with Google Cloud data services including BigQuery and Dataflow
-
Working knowledge of Google Cloud security best practices and IAM implementation
-
Hands-on experience with both GCP and AWS, and a clear point of view on where multi-cloud helps and where it hurts