
Site Reliability Engineer
Cantor Fitzgerald
Job description
Job Description
At Lucera, we're seeking an experienced SRE/DevOps Engineer to join our engineering team. You'll play a crucial role in maintaining the reliability and performance of our global trading infrastructure and financial technology platforms. This role offers an opportunity to work in a high-stakes, performance-driven environment, where your expertise in automation, infrastructure management, and operational excellence will be paramount.
Responsibilities
-
Design, build, and maintain highly available, scalable, and resilient production infrastructure.
-
Implement and manage Infrastructure as Code (IaC) solutions for automated provisioning and configuration.
-
Enhance CI/CD pipelines to ensure efficient and reliable software delivery.
-
Manage containerized workloads across modern orchestration platforms.
-
Build and improve monitoring, alerting, and observability solutions for rapid issue resolution.
-
Automate operational workflows, deployment processes, and administrative tasks.
-
Troubleshoot complex production incidents across various systems.
-
Collaborate with development teams to enhance system reliability and performance.
-
Participate in incident management, root cause analysis, and continuous improvement initiatives.
-
Contribute to disaster recovery and platform resilience strategies.
Qualifications
-
5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Engineering.
-
Hands-on expertise with Infrastructure as Code technologies (Terraform, Ansible, etc.).
-
Experience with container orchestration platforms (Kubernetes, Nomad, OpenShift).
-
Proven track record in designing and maintaining CI/CD pipelines.
-
Strong Linux systems administration and scripting skills (Bash, Shell, Python).
-
Experience with observability platforms (Grafana, InfluxDB) and automation.
-
Ability to troubleshoot complex distributed systems and applications.
-
Familiarity with Git and modern source control workflows.
-
Understanding of high availability, fault tolerance, and disaster recovery principles.
-
Excellent communication and collaboration skills for effective team work.