Thrive logo

Cloud Engineer - Linux and Automation

Thrive

RemotemidPosted 3h ago

Job description

Position Overview

The Cloud, Engineer - Linux and Automation is responsible for the administration, automation, and continuous improvement of Thrive's Linux-based infrastructure. This role serves as an escalation point for Tier 1 issues and works closely with senior engineers and architects to design and implement repeatable, code-driven infrastructure solutions. The engineer will maintain and develop automation frameworks using Ansible and Terraform, assist with managing Thriveโ€™s private cloud on HPE Morpheus Enterprise and VMware ecosystems, and ensure Ubuntu-based systems are secure, patched, and operating within established SLAs. The ideal candidate brings hands-on IaC experience, strong Linux troubleshooting skills, and a mindset oriented toward automation and operational excellence โ€“ all with a security first focus required to avoid unexpected downtime.

Responsibilities

  • Administer, monitor, and troubleshoot Linux systems (primarily Ubuntu) across physical and virtual environments.

  • Serve as a Tier 2 escalation point for infrastructure incidents; perform root cause analysis and implement remediation actions.

  • Manage and operate virtual machine workloads within HPE Morpheus Enterprise HVM and VMware ESXi, including VM provisioning, lifecycle management, and hypervisor-level troubleshooting.

  • Define and enforce standardized build documentation

  • Develop and maintain Ansible playbooks for configuration management, OS patching, application deployment, and compliance enforcement.

  • Write and manage Terraform configurations for infrastructure provisioning and lifecycle management.

  • Enforce infrastructure-as-code best practices including version control and peer review.

  • Maintain and update NetBox as the authoritative source of truth for IP address management (IPAM), DCIM asset records, and network topology documentation.

  • Integrate Netbox with IaC tools for automation of routine physical equipment commissioning

  • Participate in change management processes, ensuring all changes are documented, reviewed, and approved in alignment with Thrive's ITSM standards.

  • Perform routine system health checks, capacity reviews, and performance tuning for Linux servers.

  • Maintain and improve internal runbooks, technical documentation, and knowledge base articles.

  • Collaborate with networking, security, and cloud teams to resolve cross-functional infrastructure issues.

  • Participate in an on-call rotation to support critical infrastructure incidents outside of standard business hours.

  • Identify opportunities to improve operational efficiency through automation and standardization.

Requirements

  • 3-5+ years of hands-on Linux system administration experience with strong proficiency in Ubuntu Linux including systems, networking, storage management, and security hardening

  • 2+ years of experience with Infrastructure as Code tools, specifically Ansible (roles, playbooks, inventories) and Terraform (state management, CI/CD plan/apply workflows)

  • Hands-on familiarity with HPE Morpheus Enterprise or similar HVM/KVM platforms for VM provisioning, template management, and self-service automation; experience with VMware vCenter is a strong plus

  • Working knowledge of NetBox for IPAM, DCIM, and network source-of-truth management; experience with NetBox API integrations preferred

  • Proficiency in scripting languages for automation and tooling development

  • Solid understanding of networking fundamentals including TCP/IP, VLANs, DNS, DHCP, routing, and firewall rules as they apply to virtualized and cloud environments

  • Familiarity with Git-based version control and collaborative workflows (GitLab or GitHub)

  • Ability to diagnose and resolve complex system and infrastructure issues independently and under pressure

  • Strong documentation habits - able to produce accurate and maintainable runbooks, topology diagrams, and change records

  • Ability to work in a fast-paced environment with a diverse workload

  • Strong team player - collaborates effectively with peers and cross-functional teams to solve problems

  • Proactive, change-oriented mindset - actively seeks process improvements and drives automation-first approaches

  • Ability to communicate technical concepts clearly to both technical and non-technical stakeholders

  • Bachelor's degree in Computer Science, or a related discipline โ€” or equivalent combination of education and relevant work experience

  • Knowledge of ITIL and ITSM best practices

  • Preferred Certifications: Linux Professional Institute Certification (LPIC-1/LPIC-2) or Red Hat Certified System Administrator (RHCSA); Red Hat Enterprise Linux Automation with Ansible (RH294) or better; HashiCorp Terraform Associate (003 or 004);

Other:

  • Work Schedule: Standard business hours with participation in an on-call rotation.

  • Remote work eligible

  • Travel Requirements: Occasional travel may be required for datacenter activities. Less than 10%

  • Applicant selected will be subject to a criminal and credit background investigation and must meet eligibility requirements for access to restricted information. Candidate must be able to pass CJIS clearance for the state of Florida.