Linux Server Estate – Build, Configuration and Lifecycle
•Provision, configure, harden and maintain Linux servers across Red Hat Enterprise Linux, Rocky, Oracle Linux, SUSE and Ubuntu, on premises and in cloud, covering the full lifecycle from build through decommission.
•Maintain standardized, automated build pipelines and golden images so that servers are built to a documented standard rather than by hand, with parity between on-premises and cloud builds.
•Administer the Linux platform layer, including filesystems and LVM, storage presentation and multipath, networking, system services, kernel parameters and performance tuning, user and Sudo access, and SSH and certificate management.
•Administer Linux workloads on VMware, Nutanix and physical server hardware, including firmware currency and OEM support status for physical hosts.
•Integrate Linux hosts with enterprise directory and access management, including Active Directory or LDAP integration through SSSD, Kerberos, centralized authentication and privileged access controls.
•Manage Linux subscription, repository and package estate, including Red Hat Satellite or equivalent, repository mirroring, and content lifecycle promotion across development, test and production.
Linux Operations and Application Services Support
•Provide operational support for Linux instances working the assigned shift and supporting users and application teams in every region.
•Build, maintain and support the application services hosted on Linux, including web and application servers such as Apache, NGINX, Tomcat and JBoss, middleware, containers and runtime dependencies, in coordination with application owners.
•Support database hosting on Linux at the platform layer, including filesystem layout, kernel and resource tuning, and backup interfaces, in coordination with database owners.
•Act as second and third level escalation for complex Linux, application platform and performance issues, troubleshooting to root cause across operating system, storage, network and application layers, and liaising with vendors.
•Monitor system health and capacity through monitoring tooling and manual checks, tune alerting so that what reaches an engineer respond to degradation before it becomes an incident.
•Support application deployment and release activity on Linux, including environment preparation, pre- and post-deployment validation and rollback.
Patching, Vulnerability Remediation and Compliance
•Execute recurring patching cycles for the global Linux estate, on premises and in cloud, in line with security and compliance timelines and agreed maintenance windows.
•Remediate findings raised by the enterprise vulnerability management program, tracking items through to closure and reporting remediation status, risk acceptance and exceptions with agreed remediation dates.
•Maintain hardening baselines such as CIS benchmarks or and use Ansible to enforce and remediate configuration drift rather than only to detect it.
•Manage kernel update and live-patching strategy, reboot orchestration and cluster-aware patching so that service impact is minimized and clustered workloads are patched safely.
•Perform structured pre-patch and post-patch validation covering services, applications and dependencies, and document the outcome of every cycle.
•Follow and enforce change control, rollback planning and coordination of maintenance windows with application owners, and support internal and external audit with documented evidence.
Automation and Configuration Management with Ansible
•Build and maintain Ansible playbooks, roles, collections and inventories to automate provisioning, hardening, patching, application deployment and remediation of configuration drift.
•Operate Ansible Automation Platform or AWX where deployed, including job templates, workflows, credentials, scheduling, inventory sources and role-based access.
•Hold automation code in version control with peer review, testing and a documented promotion path across environments, so that automation is maintainable by more than its author.
•Extend automation into cloud provisioning and infrastructure-as-code, and integrate with CI/CD pipelines where used by application teams.
•Develop Bash and Python scripting where Ansible is not the right tool, and quantify the manual effort removed by each automation delivered.
Cloud Linux Infrastructure – AWS and Azure
•Administer Linux instances in AWS and Microsoft Azure, including compute, block and object storage, images, virtual networking, security groups and resource tagging.
•Maintain build, hardening, patching and monitoring parity between cloud and on-premises Linux instances, so that cloud servers are held to the same standard and evidenced in the same reporting.
•Manage cloud identity and access as it relates to Linux hosts, including IAM roles and instance profiles, key and secret management, and bastion or session-manager access patterns.
•Support assessment, migration and rollback of Linux workloads between on-premises and cloud, including the identity, network, storage and protection consequences.
•Monitor and optimize cloud consumption for Linux infrastructure, including right-sizing, commitment planning, removal of orphaned volumes and snapshots, and reporting of variance against budget.
Backup, Resilience and Disaster Recovery
•Ensure Linux workloads are protected at build rather than retrospectively, with backup agents, policies and application-consistent handling configured correctly for file, database and application workloads.
•Execute and validate restores of Linux systems and application participate in the scheduled restoration testing cadence with the Storage and Backup Engineer.
•Maintain high availability and clustering for Linux workloads where required, including Pacemaker and Corosync, load balancing and replication.
•Participate in disaster recovery drills as an executing engineer, covering and validation, and maintain recovery runbooks and the recovery time and recovery point position for the Linux platforms in scope.
Technical Leadership, Shift Handover and Documentation
•Provide technical leadership and mentoring to junior and mid-level engineers on Linux administration, automation practice and operational discipline.
•Contribute to infrastructure standards, reference builds and technology evaluation for the Linux estate, and lead root cause analysis for critical Linux incidents through post-incident review to permanent fix.
•Coordinate with network, security, database and application teams on infrastructure initiatives and complex cross-domain troubleshooting.
•Complete a documented handover at the end of each shift covering open incidents, in-flight changes, running maintenance and pending actions, and verify the handover received at shift start.
•Maintain runbooks and technical documentation, and documentation standards for the Linux estate so that operations do not depend on individual knowledge.
JOB COMPETENCIES (Skills & Abilities)
•Linux depth: Advanced, practical administration across Red Hat Enterprise Linux and comparable distributions, covering build, hardening, filesystems and LVM, systemd, networking, kernel tuning, performance analysis and troubleshooting.
•Automation expertise: Strong Ansible capability — playbooks, roles, collections, inventories and Automation Platform or AWX — applied to provisioning, hardening, patching and drift remediation, with a demonstrable record of measurable efficiency gains.
•Automation engineering discipline: Treats automation as code, held in version control with peer review, testing and a promotion path, rather than as a personal collection of scripts.
•Cloud capability: Working command of AWS and Azure as hosts for Linux workloads, including compute, storage, networking, IAM and instance profiles, secrets, and the parity required to keep cloud servers to the same standard as on premises.
•Application platform capability: Practical support of the services that run on Linux, including Apache, NGINX, Tomcat, JBoss, middleware and containers, and the judgement to separate an application fault from a platform fault.
•Security orientation: Consistent patching and vulnerability remediation aligned to enterprise security and compliance timelines, including CIS hardening, drift remediation and exception handling with documented risk acceptance.
•Patching judgement: Understands kernel and live-patching strategy, reboot orchestration and cluster-aware patching, and can sequence a patch cycle so that clustered and dependent services survive it.
•Directory and access integration: Practical command of Linux integration with Active Directory or LDAP through SSSD, Kerberos, centralized authentication, Sudo policy and privileged access.
•Resilience focus: Treats recoverability as something proven by test rather than assumed from a backup job that reports success and builds protection in at provisioning.
•Operational discipline: Runbook-driven execution, adherence to change control, and thorough pre and post change validation with a tested rollback path.
•Problem solving: Advanced analytical skill, with proven ability to identify root cause across operating system, storage, network, cloud and application layers rather than restarting a service and moving on.
•Global support capability: Supports users and application teams across multiple regions and time zones, works the agreed shift pattern, and hands over open work explicitly at shift boundaries.
•Capacity and cost discipline: Forecasts growth across the on-premises and cloud Linux estate, and manages cloud consumption through right-sizing, commitment planning and removal of orphaned resources.
•Technical leadership: Provides guidance and mentoring to less experienced engineers, sets standards for the Linux estate, and leads root cause analysis without holding formal authority.
•Documentation and knowledge transfer: Clear, maintainable runbooks and technical documentation that reduce single-person dependency across shifts.
•Communication: Excellent written and verbal communication in English, with the ability to coordinate maintenance windows and explain risk to application owners and business stakeholders.
•Accountability and autonomy: Owns systems end to end, prioritizes across a broad surface, and makes sound technical decisions with limited information and under incident pressure.
MINIMUM QUALIFICATIONS (Knowledge & Experience)
•Bachelor’s degree in computer science, Information Technology, Engineering or a closely related field, or equivalent professional experience. (Required)
•7+ years of experience in Linux systems engineering or administration, with substantial hands-on responsibility for a production Linux estate. (Required)
•Demonstrated advanced administration of Red Hat Enterprise Linux or a comparable enterprise distribution, including build, hardening, filesystems and LVM, systemd, networking, kernel tuning and performance troubleshooting. (Required)
•Demonstrated hands-on experience building and maintaining Ansible playbooks, roles and inventories for provisioning, hardening, patching and drift remediation at estate scale. (Required)
•Experience operating Ansible Automation Platform or AWX, including job templates, workflows, credentials and scheduling. (Preferred)
•Demonstrated experience holding automation code in version control with peer review and a promotion path across environments. (Required)
•Demonstrated hands-on administration of Linux instances in AWS or Microsoft Azure, including compute, storage, virtual networking, IAM roles and instance profiles, and secret management. (Required)
•Demonstrated experience maintaining build and patching parity between on-premises and cloud Linux instances. (Required)
•Demonstrated ownership of recurring Linux patching cycles, including validation, reboot orchestration, exception governance and compliance reporting against defined standards. (Required)
•Demonstrated experience remediating findings from an enterprise vulnerability management program, including tracking to closure and documented risk acceptance. (Required)
•Experience with CIS benchmarks or equivalent hardening baselines, and with detection and remediation of configuration drift. (Required)
•Demonstrated experience supporting application services hosted on Linux, such as Apache, NGINX, Tomcat, JBoss or comparable middleware, including build and troubleshooting. (Required)
•Demonstrated experience with Linux integration into Active Directory or LDAP, including SSSD, Kerberos and centralized authentication. (Required)
•Experience administering Linux workloads on VMware and Nutanix, and on physical server hardware. (Required)
•Experience with backup and restore of Linux systems and application data, and participation in disaster recovery drills and restoration testing. (Required)
•Experience with high availability and clustering for Linux workloads, such as Pacemaker and Corosync, load balancing or replication. (Preferred)
•Strong Bash and Python scripting ability. (Required)
•Experience with containers and container platforms such as Docker, Podman, Kubernetes or OpenShift. (Preferred)
•Experience with Red Hat Satellite or an equivalent subscription, repository and content lifecycle management tool. (Preferred)
•Experience operating within a formal change management process, including Change Advisory Board submission, rollback planning and post-implementation review. (Required)
•Experience working within a shift-based infrastructure team supporting a globally distributed estate, including structured handover. (Required)
•Professional proficiency in spoken and written English, sufficient to coordinate with global infrastructure and application teams and to communicate with senior stakeholders. (Required)
•Certification such as Red Hat Certified Engineer (RHCE) or Red Hat Certified System Administrator (RHCSA). (Preferred)
•Certification such as an AWS or Microsoft Azure associate-level credential, or Red Hat Certified Specialist in Ansible Automation. (Preferred)
•ITIL 4 Foundation, or equivalent demonstrated knowledge of incident, change and problem practice. (Preferred)
Disclaimer: This job description indicates in general terms, the type and level of work performed as well as the typical responsibilities of employees in this classification and it may be changed by management at any time. Other duties may also apply. Nothing in this job description changes the at-will employment relationship existing between the Company and its employees.