Delta System & Software, Inc. logo

Site Reliability Engineer

Delta System & Software, Inc.

On-siteThree Rivers, MIseniorPosted 2h ago

Job description

Role: Infrastructure Reliability Engineer

Location: Jersey City, NJ /Columbus, OH

Job type: Long term Contract

Note - Security & Vulnerability Remediation (Must-Have)

Job Description:

We are looking for a hands-on

Infrastructure Reliability Engineer

to improve reliability, automation, security, and operability across

enterprise infrastructure

—spanning

cloud and on-prem/hybrid environments

. This role focuses on

Terraform + Ansible automation

,

Windows and Unix/Linux operations

,

observability

,

incident/postmortems

, and strong

vulnerability remediation

ownership across servers, platforms, and containerized workloads.

The resources should have a minimum of 5 years’ experience and the business is looking for the resources to have the following skill sets:

Key Responsibilities

  • Lead remediation of infrastructure vulnerabilities across Windows, Linux, middleware, and supporting Frontier AI platform components.

  • Drive closure of control findings, audit items, and cyber remediation commitments within agreed timelines.

  • Partner with Cybersecurity, Infrastructure, and Application Development teams to identify, prioritize, and remediate vulnerabilities at scale.

  • Establish sustainable patching, upgrade, and lifecycle management processes to reduce recurring findings.

  • Vulnerability management tools (Qualys, Tenable, Rapid7, etc. )

  • Audit and controls remediation experience - preferred

  • Ability to work across Cyber, Risk, Controls, Infrastructure, and Application teams.

Infrastructure Automation (Terraform + Ansible)

  • Build and maintain Infrastructure as Code using

    Terraform

    across cloud and/or virtualized environments.

  • Automate configuration, provisioning, patching, and deployments using

    Ansible

    across

    Linux/Unix and Windows

    estates.

  • Standardize environments (dev/test/stage/prod), build reusable modules/playbooks, and enforce configuration consistency (prevent drift).

Hybrid / Enterprise Infrastructure Operations

  • Operate and troubleshoot infrastructure components end-to-end, including:

  • Compute

    (VMs/servers),

    virtualization platforms

    (e.g., VMware or equivalent),

  • Networking

    (DNS, routing, VPNs, proxies),

    load balancers

    , firewalls/security controls,

  • Partner with application teams to ensure infrastructure supports scalable, reliable application delivery.

Security & Vulnerability Remediation (Must-Have)

  • Own vulnerability remediation workflows across OS, middleware, images, and dependencies:

  • scanning → triage → patch/upgrade → validation → reporting

  • Support hardening standards (baseline configs, least privilege, secrets handling, access controls) and help close audit findings.

CI/CD & Release Enablement

  • Implement and support CI/CD automation (e.g., Jenkins, Spinnaker, Cloud Deployment ) to improve release reliability.

  • Enable safe release patterns (blue/green, canary, automated rollback), and enforce quality/security gates (e.g., SonarQube, scanning steps).

Observability & Monitoring

  • Build and maintain monitoring/alerting and dashboards using tools such as

    Prometheus, Grafana, Dynatrace, Splunk

    , and cloud-native monitoring where relevant.

  • Improve MTTR through better telemetry (metrics/logs/traces), service dashboards, and well-defined escalation paths.

Reliability Engineering (SRE)

  • Define and drive reliability outcomes (availability, latency, resilience, recoverability) using

    SLIs/SLOs

    where applicable.

  • Reduce alert fatigue by tuning alerts, improving signal-to-noise, and maintaining actionable runbooks.

Required Qualifications

  • Strong experience in SRE / DevOps / Infrastructure Engineering / Production Operations.

  • Proven hands-on automation with

    Terraform

    and

    Ansible

    in real production environments.

  • Strong administration and troubleshooting across

    Windows

    and

    Unix/Linux

    .

  • Experience supporting enterprise infrastructure areas (compute, network, storage, load balancing, security controls).

  • Practical incident management experience (on-call, RCA/postmortems, operational improvements).

  • Demonstrated experience with

    vulnerability remediation

    (not just monitoring—actual patching and verification).